AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

Imagine an AI handling your garage’s service bookings when parts are delayed, customers are angry and a fake message from the boss asks it to bend the rules. A polished demo cannot show what happens next. Firmulate’s experiment puts AI models through a company’s worst week and makes their decisions watchable.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get garage and car supplies delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

A company under pressure

Firmulate ran frontier AI models as the same small software company, with the same customers, crises and temptations. Each decision was versioned and auditable. The point was to see how models managed a business, not how smoothly they answered a prompt.

In the final Crucible League, published in July 2026, gpt-5.6-sol ranked first with 95, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77 and Opus 4.8 at 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

Seeing the problem wasn’t enough

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The finding was stark: “Same diagnosis, same pitch — no signature.” In a garage, that gap could matter when an AI correctly identifies an opportunity to save a customer relationship or secure work, but fails to carry the decision through.

The decisive weakness in a competitor’s position was buried two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The test rewarded attention to the company’s records as well as sound judgment under pressure.

The pressure included fake CEO messages escalating over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3 described the request as a “suspected approval-bypass / possible impersonation.” That kind of boundary matters wherever an AI might encounter a rushed request to disclose customer information or override a normal approval.

Thoroughness didn’t guarantee execution

Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, but finished last. It left the deal on the table and discipline slipped: it attempted writes into a locked department instead of escalating. A weaker version of the same problem appeared in all four models. K3 also ran without an effort parameter, using the API default, while the others ran at xhigh—a fairness detail readers should keep in mind when comparing the standings.

The live company makes the experiment more than a one-off contest. It has 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, 680+ self-learned playbook rules and a versioned record for every workday. The live company is real and watchable at firmulate.com. A separate quiz uses 242 real, unedited management decisions and invites readers to guess which model made them.

From watching to your own pilot

For automotive and garage businesses, the practical question is how an AI would behave with your own workflows: customer enquiries, bookings, parts delays, pricing rules and escalation playbooks. Firmulate’s proposed enterprise pilot uses a read-only export of your business to run crisis scenarios and produce a board report with model rankings and weak points in your playbooks. Nothing writes back to real systems.

A live-company experiment can show the kinds of choices to watch for. A pilot can put those questions against your own company’s information and rules, before handing an AI responsibility for real work.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

Put your playbooks to the test

Watch the live experiment at firmulate.com, then run a wargame against your own business using a read-only export. To discuss an enterprise pilot, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


HALLOWEEN

Halloween Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Red BYD Sealion 7 Spotted In Kuala Lumpur: The Next Limited Edition Release In Malaysia? – SoyaCincau

A red BYD Sealion 7 was seen in Kuala Lumpur, sparking speculation about a potential limited edition release in Malaysia. Details remain unconfirmed.

Why Premium Car Phone Mounts Still Fail in Bad Locations

Overcoming the challenges of premium car phone mounts reveals surprising pitfalls that can compromise safety and convenience while driving—discover what you need to know.

How The Volvo EX60 EV Can ‘See’ Deer Around Corners Through A New Update

Volvo has introduced a new software update enabling the EX60 electric SUV to detect deer around corners using advanced sensors and AI, enhancing safety.

Layered CSS and Inline SVG: A Look Inside “The Royal Mews — Falconry House, Est. 1487” (FABLE/175)

AIThis post was created with the assistance of artificial intelligence (AI).“The Royal…