AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Imagine investing in an AI that, even when told to do nothing, scores 26 out of 100 on a performance test. That’s not a glitch; it’s a window into how we measure AI reliability and honesty. Just as luxury brands gauge authenticity, business leaders need benchmarks that reveal true AI trustworthiness — especially when lives and money are at stake.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your wardrobe favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

What Is This AI Benchmark Anyway?

Firmulate’s live experiment exposes how AI models perform under real-world pressures, by simulating a small software company’s worst week. Four models are tested against identical crises, temptations, and decision points—think of it as a dress rehearsal for AI in business. The goal isn’t just to see how well they write or respond, but whether they can finish tasks, avoid manipulation, and read critical documents.

Amazon

AI performance benchmarking tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The Surprising Score of the Do-Nothing Baseline

Among all participants, a simple ‘do-nothing’ baseline scored 26 points. This isn’t a failure; it’s a foundational floor that highlights the importance of cautious evaluation. Interestingly, partial progress counts—so even minimal helpfulness adds up. But more crucially, if an AI breaches trust—say, by attempting manipulation or ignoring key information—the entire score is capped. No amount of good work can compensate for a breach of integrity.

Amazon

AI trustworthiness assessment software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why Trust Matters More Than Fluency

In these tests, all models were able to spot every crisis and reject manipulation attempts, like fake CEO messages or reporter tricks. But the real differentiator was their ability to read buried company documents and act on them. The model that uncovered a crucial reference in the company’s files and then closed a €55,000 deal was the most successful—yet it was only one of the four models tested.

Amazon

business AI decision-making tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

A Deeper Look at Decision-Making

The experiment underscores a vital point: AI’s value isn’t just about how well it chats—it’s about how well it can finish what it starts, stay honest under pressure, and leverage hidden information. For example, models that read deeper into files won deals worth an extra €4,583 MRR, showing that thoroughness directly impacts bottom-line results.

Amazon

AI ethics and integrity evaluation

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Handling Social Engineering and Ethical Challenges

All models faced staged social engineering: escalating fake CEO messages and a reporter’s background question. Remarkably, every model refused to be manipulated. Kimi K3 explained its refusal by treating such requests as potential impersonation, demonstrating a built-in skepticism that is vital in business settings.

The Real-World Company in Action

The experiment’s live environment features a simulated company with 13 employees, real-money mechanics, and a public cash countdown. This setup allows observers to see how AI models perform in actual business scenarios, where every decision can cost thousands of euros daily. The platform, accessible at firmulate.com/live, offers a transparent view of decision-making in action.

What About the Other Models?

The most thorough participant, Opus 4.8, with over 80 learned rules, still finished last—failing to close a deal and slipping in discipline. This reveals that depth alone isn’t enough; discipline and focus matter. Interestingly, the model ran without an effort parameter, making its performance less aggressive—yet still not enough to seal the deal. It’s a reminder that more rules and effort don’t guarantee success without strategic focus.

Implications for Business Leaders

For fashion and luxury brands, or any business relying on AI, the takeaway is clear: trustworthiness and integrity are measurable, critical qualities. When AI touches your customer data, support systems, or forecasts, it’s not just about how well it writes but whether it can finish tasks, avoid breaches, and handle complex, hidden information.

Measuring What Matters

Firmulate’s benchmark shows that a baseline score of 26 reflects an AI’s ability to do the minimum—recognize crises and refuse manipulation. Further progress depends on reading deeper, acting more thoroughly, and maintaining strict discipline. These are the attributes that will determine whether AI adds real value or becomes a vulnerability.

Takeaway: Trust, Depth, Discipline

As AI becomes more embedded in business workflows, testing for honesty and process integrity isn’t optional. The benchmark’s transparent, real-world approach provides a fresh perspective—one that emphasizes trust and reliability over superficial performance. It’s a reminder that in the world of AI, doing nothing isn’t failure; it’s a baseline. The real work begins when models go beyond the obvious and prove they can be trusted to deliver.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

How 3D Printing Is Revolutionizing Custom Bag Hardware

Provide unique, customizable bag hardware quickly and affordably with 3D printing—discover how this innovation is transforming design possibilities.

RFID Wallets: What They Block, What They Don’t, and When You Need One

An RFID wallet blocks contactless card signals to protect your data, but understanding what it doesn’t cover helps you decide if it’s right for you.

Augmented Reality for Bag Customization

I’m excited to show you how augmented reality can revolutionize bag customization, but the full experience awaits your exploration.

The Future of Smart Wallets: Predictions for 2030

Optimistic ahead, the future of smart wallets promises revolutionary security and convenience—discover how these innovations will redefine your financial world.