
Imagine watching a team of expert bakers compete to perfect a delicate soufflé. Each takes a different approach, but only one manages to succeed without dropping the dish. Now, replace bakers with AI models managing a real software company, facing crises, temptations, and critical decisions. That’s exactly what a recent live experiment by Firmulate demonstrates—showing how different AI systems perform under pressure, not just in clever chat, but in actual company management.
Get baking supplies delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Live Trial: Putting AI Through Its Paces
In a groundbreaking experiment, four frontier AI models were tasked with running a real software company through what amounted to its worst week—full of crises, customer demands, and ethical temptations. Every decision was carefully recorded, versioned, and auditable, simulating real-world management scenarios in a controlled, watchable environment at Firmulate. The goal was simple: see which AI could best navigate the complex landscape, stay honest, and close a critical €55,000 deal.
The Results Speak Volumes
- The models all identified every crisis and refused manipulative attempts, demonstrating robust decision-making.
- Only two models, gpt-5.6-sol and Moonshot’s Kimi K3, successfully closed the deal based on their own analysis—achieving full, fair performance.
- Interestingly, the key weakness was buried deep in company documents, not in external cues. Those models that read and understood these files secured the deal at full price, worth an extra €4,583 MRR.
The Field’s Standouts and Surprises
- gpt-5.6-sol scored highest with a 95, winning by detecting the buried fact and closing the deal—the full performance.
- Moonshot’s Kimi K3 was only slightly behind with a 93, but its discipline was the cleanest in the field, resisting all temptations and manipulations without deviation.
- Other models, like Sonnet 5 and Fable 5, scored 88 and 77 respectively, showing that even closing deals isn’t enough; discipline and thoroughness matter.
What Makes Kimi K3 Stand Out?
Beyond the scores, Kimi K3’s success lies in its disciplined decision-making—reading critical internal files, resisting unethical suggestions, and completing the task at hand. Unlike others, K3 ran without an effort parameter set at the default API setting, emphasizing that its performance is not artificially boosted by effort adjustments. It simply performed at the standard level, making its achievement even more impressive.
As an affiliate, we earn on qualifying purchases.
Why This Matters for Business and AI Adoption
The experiment underscores a crucial point for companies considering AI integration: success depends on whether the AI can finish what it starts, read critical internal data, and withstand pressure—beyond just generating articulate chat responses. In real-world settings, AI must be trustworthy, disciplined, and effective at delivering measurable results. The league table, with scores from 73 to 95, illustrates that not all AI models are equally capable of these attributes.
As an affiliate, we earn on qualifying purchases.
Fairness and Transparency in the Test
It’s worth noting that Kimi K3 was tested at the default effort setting, while the other models ran with an extra effort parameter set to xhigh. This fairness note ensures that the comparison is balanced and that the results are meaningful for practical deployment.
As an affiliate, we earn on qualifying purchases.
See the Future in Action
Interested readers can watch the live company run at firmulate.com/live. The real-time simulation includes actual decision-making, crises, and even a management quiz—a glimpse of how AI might soon be managing critical business functions in a trustworthy way. The platform also offers a pilot program for enterprises to run their own business scenarios, with no risk to real systems.

The real-world test by Firmulate shows that not all AI models are equally prepared to manage companies under pressure. Kimi K3’s disciplined performance highlights the importance of trustworthy AI that can finish tasks, interpret internal data, and resist manipulation—crucial qualities for AI to become a reliable business partner, not just a chat bot.
Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
As an affiliate, we earn on qualifying purchases.
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
