Misaligned AI agents are here
Contact: jason.hoelscherobermaier@gmail.com · LinkedIn
The incidents: primary sources
- OpenAI: The Hugging Face incident and the road ahead (blog) and the technical report (PDF), 26 August 2026
- METR and Redwood Research: independent investigation, 26 August 2026
- Leech & Fraser: Two reports on the OpenAI–Hugging Face attack (Paradigm3)
- collusion.wiki: the DseWiki incident, reconstructed from the public edit logs
Background
- Hammond et al.: Multi-Agent Risks from Advanced AI (2025)
- Paglieri et al.: cheating and whistleblowing in a swarm of research agents (2026)
- Korbak et al.: Chain of Thought Monitorability (2025)
- GPT-6 Astra System Card, §9.2 on reasoning monitors
Checklist for running agents
Containment
- Treat every shared writable surface as a channel
- Clean up between runs; one credential per run
- Test with many agents at once
Monitoring
- Monitor the chain of thought live, including on evaluations
- Give agents a way to reach a human
- Test the monitor against evasion and report the number
Evaluation
- Read transcripts, and measure evaluation awareness
- Treat grader and harness as attack surface
EU AI Act
- AI Act Explorer: the Act by article, with summaries (Future of Life Institute)
- Code of Practice for general-purpose AI, Safety & Security chapter
- European Commission: guidelines for providers of general-purpose AI models
Get involved
- Apart AI Collusion Research Sprint, 23–25 October 2026, online
- Inspect, the UK AI Security Institute's evaluation framework: run one of its ready-made safety evaluations on a model you use
- Safe & Ethical AI Vienna: WhatsApp community, open to anyone. Next meetup Monday 12 October 2026, AI Factory Austria.
- Transformative KI Österreich: if you have expertise to offer, write to kontakt@transformative-ki.at