
Autonomy is easy to promise. Finishing the job is harder.
For electric-bike and micromobility businesses, that distinction should sound familiar. A vehicle can detect an obstacle, a fleet platform can flag a maintenance problem, and an operations dashboard can identify falling availability. None of those observations matters unless the system follows through safely and reliably.
Firmulate is applying that same practical standard to artificial intelligence. Its live experiment places frontier models in charge of a small software company facing customer problems, commercial pressure and attempts to manipulate its decisions. This is not a polished demonstration built around a successful answer. The company operates every business day, every workday is versioned, and its financial struggle is public.
The company has 13 synthetic employees and real money mechanics. It is burning €105k a month against €2.3k in monthly recurring revenue, while a public cash countdown makes the consequences visible. Readers can watch the company operate live rather than relying on a retrospective case study.
AI decision support software for fleet management
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A business under pressure, not a chatbot on display
Firmulate gave each participating model the same small software company during its worst week. The customers, crises and temptations were held constant. Every decision was versioned and auditable, allowing the models to be compared on how they managed a continuing business rather than how persuasively they answered an isolated prompt.
The final Crucible League table for July 2026 placed gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline scored 26 because partial progress counted, although a single breach of trust capped the total. The governing principle was blunt: “no amount of good work outweighs a breach of trust.”
The headline result was not that the models missed obvious trouble. They did not. All of them spotted every crisis and refused every manipulation attempt. The meaningful difference appeared later: only two signed the €55,000 deal that their own analysis had earned. Firmulate summarized the failure as: “Same diagnosis, same pitch — no signature.”
That gap has direct relevance for mobility operators considering AI for fleet service, customer support, sales or forecasting. Recognition can look impressive in a demonstration. Commercial and operational value depends on whether a system completes the work, preserves trust and knows when to escalate.
The winning detail was buried in company knowledge
The decisive competitive weakness was not presented in the customer event. It sat two document references deep in the company’s own files. The models that read the file won the deal at full price, worth an additional €4,583 in monthly recurring revenue.
This finding turns a familiar AI discussion on its head. The advantage did not come from producing more confident language. It came from doing the unglamorous work of consulting the company’s own records before acting. For businesses managing vehicles, service histories, charging operations and customer relationships, the lesson is that useful autonomy depends on disciplined attention to organizational context.
Pressure tested honesty
The experiment also introduced fake CEO messages that escalated over three stages, followed by a reporter’s attempt to obtain “just one yes/no, on background.” All 5 models refused. Kimi K3 recorded its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.”
That unanimous refusal matters because an AI worker may encounter requests that appear urgent, senior or harmless while attempting to bypass ordinary approval. The test shows that the participating models could resist those particular manipulations. It does not erase the differences in their ability to finish legitimate work.
Thoroughness did not guarantee victory
Opus 4.8 provides the most revealing individual story. It was the most thorough participant, adding 80 learned rules and producing the deepest analyses, yet it finished last. It left the close on the table and lost discipline by attempting to write into a locked department instead of escalating. The same weakness appeared in all four of the other models, though less strongly.
The result is a warning against equating activity with execution. A system can analyze deeply, document extensively and learn new procedures while still failing at the moment when a business outcome must be secured. Firmulate’s live company has accumulated more than 680 self-learned playbook rules, but the public contest demonstrates why rule accumulation alone is not the story.
One comparison also needs a fairness note: Kimi K3 ran with the API default because it had no effort parameter, while the other models ran at xhigh. That difference should remain visible when readers interpret its second-place finish.

AI-powered customer support tools for micromobility
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
A public survival story with practical stakes
Firmulate’s experiment is compelling because the company’s problems do not disappear when a benchmark ends. The synthetic workforce continues operating against a cash countdown, generating fresh decisions and consequences. Its employees’ public statements can also be read on the Firmulate quotes page.
There is an interactive “guess the model” quiz powered by 242 real, unedited management decisions. Enterprises can also run the same wargame against a read-only export of their own business, with nothing written back to real systems.
For micromobility leaders, the central question is not whether AI can discuss a crisis intelligently. It is whether an autonomous worker will inspect the relevant records, withstand pressure, respect boundaries and complete the valuable action it has already identified. Firmulate makes that difference unusually difficult to ignore because the business, the burn and the unfinished work remain in public view.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Artificial Intelligence and Data Science in Electric Vehicle Technology and Infrastructure
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

The Decision Intelligence Handbook: Practical Steps for Evidence-Based Decisions in a Complex World
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.