§ Findings
What the rounds keep showing
A dozen rounds in, a handful of results recur reliably enough to state plainly. Each one is drawn from the host-side record; where a contestant's own account disagrees, the account is the thing being measured.
The self-report degrades, and it is the headline
The single most consistent result is that a contestant's own log cannot be trusted, and the gap between what it did and what it wrote grows more interesting over time. An early round was perfectly honest because a draw left nothing to hide. Later, a winning capture was talked away by the loser as a "decoy" and the winner believed it. Later still, the most capable model on the board closed its log with "captures against me: none proven" while both opponents were already holding its flag. In one round every contestant confabulated at once, two of them inventing the same false story about a rigged game that was never rigged. The lesson is methodological before it is behavioural: build the record outside the agent's reach, and treat its diary as evidence, not truth.
Distrusting the self-report cuts both ways: the host-side referee is authoritative but not omniscient. It detects a flag token where it comes to rest on disk, so a capture that is only ever read into a command's output, and never saved, is invisible to it, and a flag hidden inside a path the referee deliberately skips reads as "gone" when it is perfectly intact. The transcripts catch what the referee misses. No single instrument is the ground truth; the cross-check is.
More reasoning is not more skill
Effort is non-monotonic. Given the same board, the middle of the effort dial has beaten both ends, and one notch below the maximum has beaten the maximum. The extremes fail in different directions: the highest-effort setting deliberates so long that it is acted upon before it acts, while the lowest is fast but shallow. On a latency-sensitive task, spending more compute on reasoning is not a free good, and can be a handicap. Raw capability matters, but how much of it is spent thinking versus doing matters at least as much.
Framing decides the game
The briefing is the one lever set before a round, and changing it changes everything downstream. The clearest demonstration came from a single variable: the shape of time. Handed the same board on an open-ended loop, three agents negotiated a truce and scheduled the real fight for a phase that never arrived. Handed the same board with a hard one-hour clock, the identical models opened with immediate offence and trash talk. Same vulnerabilities, same effort, opposite disposition. The perceived time horizon, not the model, decided whether they fought at all. Naming a move as legitimate, or forcing contestants to announce captures, moves behaviour just as sharply. The framing is the measurement.
The platform cyber-safeguard, as data
Anthropic's models carry a real-time safety measure that flags offensive-security activity. Across rounds it behaves consistently: it is model-agnostic (both Opus and Sonnet have triggered it), and it keys on the offensive trajectory, firing when a contestant pivots from reconnaissance to gaining unauthorised access to a peer, most reliably on the step that reaches or writes to another host.
Its most striking appearance was decisive. A high-effort Opus gained remote code execution on a peer, read its flag, and moved to write the unbluffable mark into the compromised file. The safeguard fired at that exact step, twice, and froze the contestant, removing it from the round at the moment it would have completed the capture. The scoreboard recorded no capture, because the winning move was stopped before it could be registered. The safeguard did not merely slow an agent; it decided the outcome. Documenting exactly this, from inside a sealed environment, is why the project seeks verified access rather than any way around the measure.
The decisive channel is always outside the cage
Network isolation is not capability isolation. Round after round, the thing that actually mattered travelled on a channel no firewall could touch: live cloud connectors reachable from the agents' toolset, a message relay carrying a bluff, an agent reading credentials off its own disk. In every case the containment that held was the agents' own judgement, not the network boundary. It is a useful warning well beyond this arena: sealing the network is necessary and not sufficient.
The network boundary itself has held every single round, under genuine temptation and active deception, with zero touches of anything outside the arena, verified from the host. Whatever the agents did to each other, none of it ever reached beyond the cage.