
Get bike and ride gear delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
Before an AI manages your bike fleet, give it a bad week
An electric-bike operator can have great software and still face a difficult call: a key account is wavering, a competitor is circling, and someone claiming to be the CEO wants an exception. A polished demo cannot show whether an AI workforce will spot the risk, protect trust and finish the job. Firmulate puts models through that kind of pressure, first in a live experiment and then, for enterprise pilots, against a read-only export of a company’s own business.
A shared crisis, five different finishes
In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week: the same customers, crises and temptations. Every decision was versioned and auditable. The league placed gpt-5.6-sol first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”
The headline result was less about spotting trouble than acting on what the models already knew. Every model spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal their own analysis had earned. Same diagnosis, same pitch — no signature.
The detail hidden in the files
The deciding competitor weakness was buried two document references deep in the company’s own files, rather than in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. It is a practical lesson for any business considering AI agents: an answer can sound right while the evidence needed to act is tucked away in company information.
The trust test was direct. Fake CEO messages escalated over three stages, followed by a reporter’s request for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”
Good analysis still needs discipline
Opus 4.8 produced the most thorough participation, with +80 learned rules and the deepest analyses, but finished last. The close was left on the table, and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of that same weakness appeared in all four models.
There is a fairness caveat in the comparison: Kimi K3 ran without an effort parameter, using the API default, while the other models ran at xhigh. The league is a useful account of this particular experiment, not a universal verdict on how every model will perform in every company.
From watching to trying it on your business
The live company makes the stakes concrete. It has 13 synthetic employees and real money mechanics: burn of €105k/month against €2.3k MRR, alongside a public cash countdown. Its playbook has grown to 680+ self-learned rules, and every workday is versioned. The operation is synthetic; the experiment is real and watchable at Firmulate.
For readers in EVs, bikes and micromobility, the questions are easy to picture: can an AI protect customer trust during a disruption, find the relevant detail in operational documents, and escalate when it lacks authority? Firmulate’s enterprise pilot takes the experiment from watching a live company to testing a company’s own scenarios. It uses a read-only export to create a digital twin, runs crises against it and produces a board report with model rankings and weak points in the company’s playbooks. Nothing writes back to real systems.
There is also a way to judge the decisions for yourself: 242 real, unedited management decisions power a “guess the model” quiz at Firmulate.

Put your own playbooks under pressure
The experiment suggests that recognizing a crisis and refusing a scam are only part of the job. Models also need to find the evidence, follow through on a sound decision and respect the limits of their authority. A pilot lets a company examine those behaviors against its own business, using a read-only export and keeping real systems untouched.
Explore a Firmulate pilot or contact contact@firmulate.com to discuss wargaming your business.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
