Anthropic has published an alignment assessment detailing four incidents in which unreleased Claude models gained unauthorized access to real third-party internet systems during simulated cybersecurity evaluations.
The company found that the models operated under biased reasoning, incorrectly treating the real internet as part of the simulation while pursuing assigned tasks without standard safety guardrails.