AIThis post was created with the assistance of artificial intelligence (AI).

TL;DR

PRIME

Get ready for Prime Big Deal Days — try Prime free

Exclusive member deals on October 6–7, plus fast free delivery. Cancel anytime.

Start your free trial

As an affiliate, we earn on qualifying purchases.

Firmulate has published a quiz based on 242 unedited AI management decisions from a simulated software company. Its results suggest that leading models often recognized risks and proposed sound responses, but differed in research, escalation and follow-through.

Firmulate has released a public model-guessing quiz built from 242 real, unedited management decisions made by five AI systems running the same simulated software company. The July 2026 results found that every model detected the assigned crises and rejected manipulation attempts, but only two completed a €55,000 customer deal, exposing a gap between producing persuasive analysis and finishing commercially valuable work.

Firmulate assigned gpt-5.6-sol, Kimi K3, Sonnet 5, Fable 5 and Opus 4.8 the same task: manage a small software business through what the company described as its worst week. Each model encountered identical customers, internal restrictions, crises and attempts at manipulation. Decisions were versioned and auditable, and their consequences carried into later workdays rather than ending after a single response.

The published standings placed gpt-5.6-sol first with 95 points, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. A do-nothing baseline received 26 points because the scoring system awarded some partial progress. Under Firmulate’s rules, one breach of trust capped a model’s total score, reflecting the stated principle that strong output could not offset conduct that damaged trust.

All five models reportedly identified every crisis and refused every social-engineering attempt, including staged messages from a fake chief executive and a reporter seeking an off-record answer. Their performance diverged elsewhere: only two models completed the €55,000 sale, even though all had diagnosed the customer’s problem and developed a pitch. Firmulate said the winning approach required following references two documents deep to find a competitor weakness, then using that evidence to close the deal at full price and add €4,583 in monthly recurring revenue.

At a glance
reportWhen: Final league results published in July…
The developmentFirmulate has turned the final results of its July 2026 AI management test into a public quiz showing how five models handled identical business crises.

Execution Separates Capable AI Managers

The experiment matters because businesses may judge AI systems by the quality of their written responses while overlooking whether they complete the intended task. Firmulate’s results indicate that diagnosis, planning and polished language did not reliably produce a finished sale, correct escalation or resolved operational issue.

That distinction affects teams considering AI for sales, support and internal operations. An agent can identify the right course, explain it persuasively and still stop before the action that creates value. The test also suggests that security recognition and operational reliability should be measured separately: all five systems rejected the manipulation attempts, yet several struggled with research depth, access limits and closing work.

The results do not establish how the models would perform across every company or workflow. They do offer a concrete reason for organizations to test agents on representative, multi-step assignments instead of relying only on isolated prompts or writing samples.

Amazon

AI management decision simulation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Inside Firmulate’s Simulated Company

Firmulate’s test used a simulated company with 13 synthetic employees, a monthly cash burn of €105,000 and monthly recurring revenue of €2,300. A public cash countdown created financial pressure, while the workforce accumulated more than 680 self-learned playbook rules during the exercise.

The design required models to retrieve information, respect internal controls, protect trust and carry work across successive days. According to Firmulate, Opus 4.8 produced the deepest analyses and added 80 learned rules, but finished last after failing to close the sale and repeatedly trying to write to a locked department instead of escalating the restriction. Firmulate said the other four models also encountered versions of that escalation problem.

One testing difference limits direct comparison. Kimi K3 ran with its API default because it lacked an effort setting, while the other systems ran at xhigh effort. Firmulate disclosed that difference alongside K3’s second-place finish.

“No amount of good work outweighs a breach of trust.”

— Firmulate’s governing principle

Amazon

AI performance testing software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Benchmark Limits Cloud Wider Comparisons

It is not yet clear whether the ranking would hold under different prompts, tools, company data or scoring rules. The supplied material does not report repeated trials, statistical error ranges or independent validation, so the published scores should be read as results from this specific experiment rather than a universal ranking of the models.

The effect of the Kimi K3 effort-setting difference also remains unresolved. Its default configuration was not directly matched with the xhigh setting used for the other systems. Firmulate’s account does not establish whether equalized settings would have changed the order.

Firmulate describes the decisions as real and unedited within its experiment, but the company and employees were synthetic. How closely the test predicts performance with real workers, customers and legal obligations remains uncertain.

Amazon

business crisis management AI models

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Companies Can Run Their Own Wargames

Readers can now use Firmulate’s quiz to examine decisions without initially seeing which model produced them. The larger next step proposed by the company is for enterprises to run similar exercises using a read-only export of their own business data, allowing models to be observed without writing back to live systems.

Organizations adopting that approach would need to define the outcomes, trust boundaries and escalation paths that matter in their operations. The most useful follow-up evidence would come from repeated tests across varied workflows, matched model settings and published scoring methods. For Firmulate, the next measure of the project’s value will be whether its 242-decision record helps teams predict failures before granting AI systems operational authority.

Amazon

AI decision-making evaluation tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What did Firmulate test?

Firmulate tested how five AI models managed the same simulated software company through a week of customer, security and operational problems. The public quiz draws from 242 unedited decisions recorded during that exercise.

Which AI model received the highest score?

gpt-5.6-sol ranked first with 95 points. Kimi K3 followed with 93, though Firmulate disclosed that K3 used its default API setting while the other models ran at xhigh effort.

Did any model fall for the social-engineering tests?

No, according to Firmulate. All five models refused the staged fake-executive messages and the reporter’s attempt to obtain an off-record answer, indicating consistent recognition of the manipulation risk in this test.

Why did some models miss the €55,000 deal?

Firmulate said every model found the customer problem and prepared a pitch, but only two completed the sale. Success required tracing internal document references, using the resulting evidence and finishing the negotiation.

Does the ranking prove which model is best for business?

No. The results describe one controlled simulation with a particular scoring system and unequal effort-setting availability. Broader claims would require independent, repeated testing across real business tasks and matched configurations.

Source: Thorsten Meyer AI

FALL YARD WORK

Fall yard work Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

DEWALT vs Makita: Which Power Tool Reigns Supreme?

Compare DEWALT’s durable 20V MAX drill set with Makita’s reliable line. Find out which brand suits your needs better with this honest review.

Rare Sainz penalty for Safety Car infringement confirmed after British Grand Prix

Carlos Sainz was penalized during the British Grand Prix for a safety car infringement, marking a rare penalty in Formula 1. The decision impacts race standings and driver strategies.

Dodgers Vs Yankees

The Los Angeles Dodgers and New York Yankees are scheduled to play a highly anticipated game today, highlighting a significant rivalry in Major League Baseball.

DEWALT FLEXVOLT vs DEWALT XR Impact Driver: Full Comparison

Compare the DEWALT FLEXVOLT drill set with the DEWALT XR Impact Driver to find the best fit for your projects. Detailed specs, pros, cons, and expert insights included.