AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

In fashion, a polished look can open a door, but the sale still depends on what happens next. Firmulate’s experiment puts that distinction to a business test: can an AI company spot trouble, resist pressure and follow through when a valuable customer is ready to sign?

Before you orderOffer from Amazon

Get your wardrobe favorites delivered free with Prime

  • Fast, free delivery on millions of items
  • Prime Video, Amazon Music and more included
  • Member-only deals all year
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

One company, one very bad week

Firmulate ran frontier AI models through the same small software company and its worst week: identical customers, crises and temptations. Every decision was versioned and auditable. The live company has 13 synthetic employees and real money mechanics, including €105,000 in monthly burn against €2,300 in monthly recurring revenue. Its public cash countdown and evolving playbooks make the experiment watchable at firmulate.com.

The final Crucible League, published in July 2026, ranked gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. The standard was demanding: partial progress counted, but a single breach of trust capped the total. As the rule puts it, “no amount of good work outweighs a breach of trust.”

Spotting the crisis was not enough

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap was not in diagnosis or the pitch: “Same diagnosis, same pitch — no signature.” A model could recognize the right move and still fail to make it.

The deal hinged on a weakness hidden two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The detail is a reminder that useful business judgment can depend on noticing what is buried in the paperwork, not just reacting to the latest development.

Trust faced its own test. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

The leaderboard also has a caveat. Kimi K3 ran without an effort parameter, using the API default; the other models ran at xhigh. Opus 4.8 was the most thorough participant, with more than 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and showed weaker discipline by attempting writes in a locked department instead of escalating. That same weakness appeared, less strongly, in all four models.

From watching to trying it on your business

The live experiment offers an ongoing view of decisions in a synthetic company. Firmulate says enterprises can take the next step: run a similar wargame against a read-only export of their own business, examining crisis scenarios and the weak points in their playbooks. The exercise does not write back to real systems. A separate quiz uses 242 real, unedited management decisions to invite readers to guess which model made each choice.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Try a pilot

To explore a pilot using your company’s read-only data, visit Firmulate’s pilot page or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Why a Do-Nothing AI Gets 26 Points — And What It Means for Business Decisions

Discover why a simple do-nothing AI scores 26 in a real-world benchmark, revealing how trust, depth, and discipline determine AI’s true value in business decisions.

How 3D Printing Is Revolutionizing Custom Bag Hardware

Provide unique, customizable bag hardware quickly and affordably with 3D printing—discover how this innovation is transforming design possibilities.

What Fashion Brands Can Learn from AI’s Hidden Strengths in Crisis Management

AI models tested in a real company scenario reveal that execution and discipline matter more than chat quality. Fashion brands should focus on trust-building processes that deliver under pressure.

RFID‑Blocking Pockets: Do You Really Need Them?

Protect your sensitive data with RFID-blocking pockets—discover whether they’re essential or just a false sense of security.