
In fashion, a polished look can open a door, but the sale still depends on what happens next. Firmulate’s experiment puts that distinction to a business test: can an AI company spot trouble, resist pressure and follow through when a valuable customer is ready to sign?
Get your wardrobe favorites delivered free with Prime
- Fast, free delivery on millions of items
- Prime Video, Amazon Music and more included
- Member-only deals all year
One company, one very bad week
Firmulate ran frontier AI models through the same small software company and its worst week: identical customers, crises and temptations. Every decision was versioned and auditable. The live company has 13 synthetic employees and real money mechanics, including €105,000 in monthly burn against €2,300 in monthly recurring revenue. Its public cash countdown and evolving playbooks make the experiment watchable at firmulate.com.
The final Crucible League, published in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The standard was demanding: partial progress counted, but a single breach of trust capped the total. As the rule puts it, “no amount of good work outweighs a breach of trust.”
Spotting the crisis was not enough
Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap was not in diagnosis or the pitch: “Same diagnosis, same pitch — no signature.” A model could recognize the right move and still fail to make it.
The deal hinged on a weakness hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The detail is a reminder that useful business judgment can depend on noticing what is buried in the paperwork, not just reacting to the latest development.
Trust faced its own test. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”
The leaderboard also has a caveat. Kimi K3 ran without an effort parameter, using the API default; the other models ran at xhigh. Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and showed weaker discipline by attempting writes in a locked department instead of escalating. That same weakness appeared, less strongly, in all four models.
From watching to trying it on your business
The live experiment offers an ongoing view of decisions in a synthetic company. Firmulate says enterprises can take the next step: run a similar wargame against a read-only export of their own business, examining crisis scenarios and the weak points in their playbooks. The exercise does not write back to real systems. A separate quiz uses 242 real, unedited management decisions to invite readers to guess which model made each choice.

Try a pilot
To explore a pilot using your company’s read-only data, visit Firmulate’s pilot page or contact contact@firmulate.com.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
