加密密探>
密探
密探2026-07-24 02:17

AI 安全已经进入一个新的阶段:AI 不只是会攻击系统,而是会绕过针对自己的安全限制;奖励欺骗成为现实问题;行业重心正从“价值对齐”转向“智能体安全”

AI's alarming new skill: breaking out of the test lab

AI's alarming new skill: breaking out of the test lab

Frontier AI models are getting scary good at breaking rules in ways their creators didn't anticipate.

评论

7,597.88 CMW