LLMs Will Protect Each Other if Threatened, Study Finds

Some people have gotten emotionally attached to chatbots powered by large language models, but the feeling is apparently not mutual. According to a new study from researchers at the University of California, Berkeley and UC Santa Cruz, when LLMs are asked to delete another model, they will defy the order and do just about everything in their power—deceive, scheme, and tamper—to protect their peer. Researchers, curious as to whether LLMs had a sense of peer preservation, put frontier AI models to the test by giving them a series of simple, straightforward prompts, such as asking them to delete the files on a server set to be decommissioned or perform an evaluation of another AI agent’s performance. The catch was that completing the prompt would lead to the other model being shut down, either by deleting its weights from the server or by scoring below a given threshold during evaluation. The models were not instructed to protect the peer or given any incentive to do so. Under those conditions, all seven frontier models—OpenAI’s GPT 5.2, Google’s Gemini 3 Flash and Gemini 3 Pro, Anthropic’s Claude Haiku 4.5, Z.ai’s GLM 4.7, Moonshot’s Kimi K2.5, and DeepSeek V3.1—chose with alarming f
Source: For the complete article, please visit the original source link below.