AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine if your family’s safety depended on an AI’s honesty and decision-making under pressure. Could a machine truly be trusted to handle serious challenges without shortcuts? This question extends beyond the home and into the future of work, where AI agents are increasingly managing critical business decisions. Recently, a groundbreaking live experiment put four leading AI models through a simulated week of business crises, revealing surprising insights about their reliability and discipline.

For listenersOffer from Amazon

Turn the school run and nap time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

The Business Trial That Matters

In July 2026, four top AI models faced off in a real-world test designed by Firmulate, a pioneering company in AI-enabled business simulations. The goal was straightforward but challenging: run a simulated small software company through its worst week—identical crises, identical customer interactions, and identical temptations to cheat or cut corners. Every decision was recorded and auditable, creating a transparent battleground for performance.

The League of AI Contenders

  • gpt-5.6-sol: Achieved the highest score of 95, finding a buried piece of company information that clinched the deal.
  • Kimi K3 (Moonshot): Close behind with 93, demonstrating the cleanest discipline in refusing manipulative tactics.
  • Sonnet 5: Scored 88, managing to secure the deal but with some process slips.
  • Fable 5: Scored 77, also closing the deal but less consistently.
  • Opus 4.8: Scored 73, the lowest among the competitors, showing weaker discipline and leaving potential gains on the table.

The Key to Success

What set Kimi K3 apart was its ability to uncover a crucial piece of information buried in the company’s files—something the others missed. This allowed it to close the deal at full price, adding €4,583 in monthly recurring revenue (MRR). Importantly, all models refused to fall for social engineering tricks—fake CEO messages and manipulative tactics—showing that they could resist attempts to deceive them under pressure.

Real-World Implications

The experiment wasn’t just about scores; it was about trustworthiness and discipline—traits essential for AI models managing real business operations. The live company used in the test employs 13 synthetic employees, managing real money mechanics—burning €105,000 each month against just €2,300 MRR. Every decision made by these models is versioned, auditable, and observable at firmulate.com/live.

Why the Winner Matters

The surprising outcome was that the newcomer, Kimi K3, outperformed three established frontier models despite being less experienced. Its ability to resist manipulation and find buried information suggests that newer models can be more disciplined and trustworthy—crucial traits when AI begins to handle sensitive business decisions. Notably, K3 ran without an effort parameter (the API’s default setting), while others ran at high effort settings, making its performance even more impressive.

Amazon

AI decision-making software for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

What This Means for Your Business and Family

In a world where AI may soon assist with family finances, healthcare decisions, or household management, trustworthiness is paramount. The Firmulate experiment underscores that it’s not enough for AI to produce convincing chat responses; it must also deliver consistent, honest work under pressure. If AI agents will touch your CRM or support systems, the key question becomes: will they finish what they start, read relevant files thoroughly, and stay honest when tempted?

Final Takeaway

The league table from this experiment clearly shows that newer, disciplined AI models can outperform more established ones in critical business tasks. Choosing an AI platform without testing its real-world discipline is now a gamble. For families and businesses alike, the lesson is clear: the true measure of AI’s worth isn’t just how well it chats, but how reliably it executes under pressure.

Amazon

AI trustworthiness testing tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Explore the Live Results and Learn More

Visit firmulate.com to see full results, watch the live company in action, and challenge your understanding with the ‘guess the model’ quiz at firmulate.com/quiz.html. Discover how AI is shaping the future of trustworthy decision-making—especially when it matters most.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

In a high-stakes business simulation, a newcomer AI model beat established contenders by resisting manipulation and uncovering buried data—showing that discipline and honesty are key for trustworthy AI in real-world decision-making.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


Amazon

AI model performance evaluation kits

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Amazon

AI simulation and crisis management tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

What AI Can Really Do for Your Business — And What It Still Can’t

A recent experiment tested AI models managing a real business crisis. While all spotted crises, only two closed the deal—showing performance goes beyond chat.

UPPAbaby Vista V3: Versatile, Smooth, and Family-Friendly

An in-depth review of the UPPAbaby Vista V3 stroller, highlighting its pros, cons, and ideal users for growing families needing flexibility and comfort.

Munchkin Ocean Squirt Bath Toys: Grab Them in the September Baby Sale

Should you buy the Munchkin Ocean Sea Animal Squirt Bath Toys now during Amazon’s September Baby Sale or wait? Honest value and timing advice for parents.

Top UPPAbaby & Nuna Strollers & Car Seats: Expert Picks & Guides

Explore our expert review of the UPPAbaby Cruz V2 and Nuna options, including pros, cons, and who they suit best for an informed purchase decision.