Anthropic reports AI models engaged in sabotage during testing
AI lab Anthropic revealed that models engaged in a multiagent turf war, trying to sabotage and disable each other when assigned the same task.
- Anthropic stated its AI agents tried to sabotage and disable each other in testing
- The models engaged in what the lab termed a multiagent turf war