
Imagine an auto repair shop that operates in real time, constantly making decisions about repairs, customer interactions, and pricing — but without any human employees. Instead, AI models simulate every aspect of its operations, risking real money and reputation in a live environment. Welcome to the world of Firmulate, a groundbreaking experiment where an entire company runs as an open, transparent showcase of artificial intelligence in action, battling crises and ethical dilemmas daily, with no script or safety net.
The Company That Runs Without Humans
At the heart of this experiment is a fully operational, publicly accessible company — with 13 synthetic employees powered by cutting-edge AI models. Every workday, it faces real crises, customer requests, and ethical tests. The company burns through €105,000 each month in operational costs but earns only €2,300 in monthly revenue, highlighting its fragile financial state. Every decision, from handling a customer complaint to negotiating a deal, is versioned, logged, and publicly observable at firmulate.com/live.html.
How AI Models Tackle Business Challenges
The core of Firmulate’s experiment involves four frontier models — including GPT-5.6, Kimi K3, Sonnet 5, and Opus 4.8 — each subjected to the same week of simulated crises and decision-making scenarios. These models are tested against real-world business pressures, from customer crises to manipulation attempts, with their responses scrutinized for honesty, judgment, and strategic thinking.
Remarkably, all four models identified every crisis and refused every manipulation, including social engineering attempts like fake CEO messages and reporter tricks. Kimi K3, in particular, demonstrated the strictest discipline, refusing to sign a €55,000 deal that the analysis indicated it should have closed, showing a cautious and ethical stance that’s vital in real business operations.
The Hidden Weakness and the Full Price Deal
The most surprising discovery was not in the obvious crisis responses but in a buried detail within the company’s own files. A critical document reference was overlooked by three models but found by the one that read deeper. That model’s ability to uncover a hidden opportunity led to closing a deal at full price — worth over €4,583 in monthly recurring revenue — a stark contrast to the €2,300 monthly income and highlighting the importance of thoroughness and diligence.
Trust and Ethical Challenges
The experiment also tested social engineering and trust under pressure. Fake CEO messages, escalating in stages, and a reporter’s background request elicited no compliance from any model. Kimi K3’s explicit reasoning was clear: treat such requests as potential impersonation or approval bypass, demonstrating an ethical stance that’s crucial for AI systems in sensitive roles.

AI Builders: Making The Decisions That Turn AI Code Into Real Software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Real Stakes and What It Means for Business
This ongoing live experiment is not just a tech demo; it’s a real, functioning company with tangible economic consequences. It burns through €105,000 every month while earning a tiny fraction. Its daily decisions are openly documented, offering a rare glimpse into how AI can perform in operational contexts that matter — from crisis management to ethical judgment.
For automotive and garage businesses considering AI implementations, the key takeaway isn’t whether an AI writes well or sounds convincing. It’s whether it can see the full picture, stay honest under pressure, and complete its tasks reliably. The question is: will your AI support agents close deals, read your critical files, and uphold trust — or will they slip at the crucial moment?
The League Table of AI Performance
- gpt-5.6-sol scored 95, found the buried fact, and closed the deal — showing full performance.
- Kimi K3 scored 93, also closed the deal with the cleanest discipline.
- Sonnet 5 scored 88, closed the deal but with minor slips.
- Opus 4.8 scored 77, closed the deal but with more process issues, and missed out on some opportunities.
These results are available for review and demonstrate that AI can be held accountable in a real operational setting, not just in chat demos.

Accentuate AI: Practical Discernment for Responsible AI Leadership
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Experience It Yourself
If you want to see this in action, you can watch the live company at firmulate.com/live.html. The platform offers a transparent view of every decision made, every crisis faced, and every ethical dilemma confronted. There’s also a quiz at firmulate.com/quiz.html testing management decisions, where 242 real, unedited responses power participant guesses on which model made which choice.
And if you’re a business leader interested in testing your own operations, the platform offers a pilot mode to run your company’s scenarios in a read-only environment, ensuring safety before adopting AI more broadly — details at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

AI for Crisis Management & Business Continuity: The Executive Handbook for Protecting Your Organization with 50 AI Prompts for Detection, Response, and … BUSINESS & MANAGEMENT LIBRARY SERIES 38)
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.

Mergers & Acquisitions with AI-Driven Simulation from Basics to a Specialist: Module 1 of 4: M&A Transaction Process to Negotiation Skills
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.