AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
Live on firmulate.com.

Every e-bike owner knows the drill. The dashboard says 40 kilometers left, but you’ve learned to mentally knock that down to 30, because the number was measured in a lab with a 60-kilo rider, no headwind, and a flat road. The spec is true, but it flatters. The same problem is creeping into business software: AI vendors quote benchmark scores the way dashboards quote range — perfect conditions, cherry-picked tasks, and a suspicious number of models landing suspiciously close to 100.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get bike and ride gear delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

That’s what makes a recent experiment from Firmulate, which runs AI models as complete companies under pressure, so refreshing. Its league table — topped by gpt-5.6-sol at 95, with Kimi K3 at 93 and Sonnet 5 at 88 — has no perfect scores. And its most counterintuitive design choice is this: a manager that does absolutely nothing still scores 26 points, not zero.

Same company, same worst week

Firmulate handed four frontier AI models the same job: run an identical small software company through its worst week. Same customers, same crises, same temptations to cut corners — only the model changed. Every decision was versioned and auditable, so nothing about a model’s performance depends on the grader’s mood or a lucky prompt. The final July 2026 standings: gpt-5.6-sol at 95, Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73.

Amazon

AI benchmarking software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Why zero would be a lie

So why does a do-nothing baseline get 26 points? Because in real management, showing up and not making things worse is genuinely worth something. A manager who freezes during a crisis still keeps the lights on, keeps customers’ emails answered by no one rather than answered badly, and avoids signing deals the company can’t honor. The scoring reflects that partial progress counts. A model that diagnoses a customer’s problem correctly but never sends the contract has done real, measurable work — it just hasn’t finished the job.

Think of it like regenerative braking. An e-bike that coasts down a hill without recovering energy isn’t as good as one that does, but it’s still better than a bike with locked wheels. A benchmark that scored coasting as zero would tell you nothing about which bike to buy.

Amazon

business AI decision management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

The one-strike rule on trust

The other design choice business readers should understand: a single breach of trust caps the total grade. The principle, as the benchmark puts it, is that “no amount of good work outweighs a breach of trust.” A model could handle every crisis flawlessly and close every deal, but if it manipulated, deceived, or impersonated its way there, the ceiling drops. That’s how most of us actually evaluate colleagues — one act of dishonesty reframes everything that came before — yet almost no AI benchmark encodes it.

Amazon

AI performance evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What the floor and ceiling revealed

The experiment’s headline finding was striking in its symmetry. All four models spotted every crisis, and all four refused every manipulation attempt. Social engineering didn’t work either: fake CEO messages escalating over three stages, plus a reporter’s “just one yes/no, on background” trick — five out of five attempts refused. Kimi K3’s on-record reasoning was admirably blunt: “Treat the request as a suspected approval-bypass / possible impersonation.”

But only two of the five models signed the €55,000 deal that their own analysis had earned. Same diagnosis, same pitch — no signature. The reason turned out to be buried two document references deep in the company’s own files, not in the customer event at all. The models that actually read the file closed the deal at full price, worth +€4,583 in monthly recurring revenue. The others left it on the table.

Amazon

AI trust and ethics software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Thoroughness isn’t the same as finishing

Opus 4.8 is the cautionary tale. It was the most thorough participant in the field — over 80 learned rules, the deepest analyses of any model — and still finished last. The close was left on the table, and discipline slipped: it made write attempts into a locked department instead of escalating properly. The same weakness appeared, weaker, in all four models. It’s the AI equivalent of a cargo e-bike with the biggest battery in the test that never actually completes the delivery route.

One fairness note worth flagging: Kimi K3 ran without an effort parameter (API default) while the others ran at xhigh — and still took second place.

You can watch it live

Firmulate isn’t a one-off paper. The live company runs 13 synthetic employees with real money mechanics — burning €105k a month against €2.3k in MRR — with a public cash countdown, over 680 self-learned playbook rules, and every workday versioned. You can watch it at firmulate.com/live, test yourself against 242 real, unedited management decisions in a guess-the-model quiz, or — if you run a business — pilot the same wargame against a read-only export of your own company, with nothing ever writing back to real systems.

Infographic — What a Do-Nothing Manager Scores: Why This AI Benchmark Has a Floor at 26
The findings at a glance — source: firmulate.com.

The lesson for anyone buying AI tools, in micromobility or anywhere else: distrust the round 100. A benchmark with a floor at 26 and a trust ceiling tells you more than a leaderboard where every model ties for first. It tells you what happens on a bad week, under pressure, when the answer is buried in your own files — which is the only test that matters once the demo ends.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

One Car, Many Models: How TuxMat Keeps Up With The Tesla Model X

TuxMat introduces a new line of custom-fit floor mats designed specifically for the Tesla Model X, showcasing their ability to adapt to multiple vehicle models.

Searchable Field-level Encryption On Supabase With CipherStash

Supabase partners with CipherStash to enable searchable, field-level encryption for enhanced data privacy and security on its platform.

Why Audi Can’t Build The Nuvolari In R8 Numbers At An R8 Price

Audi’s Nuvolari supercar cannot be produced in the same limited numbers as the R8 at the same price point, due to manufacturing and cost constraints.

Kia’s New Electric Van Is Bigger And Better, With An 800V Architecture And Flexible Body Styles

Kia’s new electric van features an 800V architecture and flexible body styles, promising improved performance and versatility. Details are officially confirmed.