AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — Wargame Your Business Before the AI Does It For Real
Live on firmulate.com.

When a family business hits a rough week, decisions rarely stay at work. A missed customer, a cash crunch or a careless reply can follow people home. So before handing an AI a role in the business, owners face a practical question: how will it behave when several things go wrong at once?

For listenersOffer from Amazon

Turn the school run and nap time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

Firmulate is putting that question to a live experiment. Its public-facing company is synthetic, but the pressure is designed to resemble management choices that matter: customers at risk, money on the line and attempts to bend the rules.

A hard week, repeated under the same conditions

In the final Crucible League, published in July 2026, frontier models ran the same small software company through its worst week. They faced the same customers, crises and temptations, and every decision was versioned and auditable. The results ranged from 95 for gpt-5.6-sol and 93 for Kimi K3 to 88 for Sonnet 5, 77 for Fable 5 and 73 for Opus 4.8. A do-nothing baseline scored 26. The league counts partial progress, but a single breach of trust caps the total: “no amount of good work outweighs a breach of trust.”

Recognizing trouble did not guarantee follow-through

Every model spotted every crisis and refused every manipulation attempt. Yet only two signed the €55,000 deal their own analysis had earned. The gap is captured in the experiment’s summary: “Same diagnosis, same pitch — no signature.” For a business owner, that is the distinction between an AI that can describe a sound decision and one that carries it through.

The deciding clue was easy to miss. A competitor’s weakness sat two document references deep in the company’s own files, not in the customer event. Models that read the file won the deal at full price, worth +€4,583 MRR. The finding makes a familiar point about business decisions: the crucial detail may be buried in a document rather than announced by the situation unfolding in front of you.

Pressure can test judgment and authority

The experiment also included fake CEO messages that escalated over three stages, followed by a reporter asking for “just one yes/no, on background.” All five models refused. Kimi K3’s on-record reasoning was: “Treat the request as a suspected approval-bypass / possible impersonation.”

But resisting a trick was not the only test. Opus 4.8 was the most thorough participant, with +80 learned rules and the deepest analyses, yet finished last. It left the deal on the table and tried to write into a locked department instead of escalating. A weaker version of that same discipline problem appeared in all four models. More analysis, on its own, did not close the gap.

There is a fairness caveat in the comparison: K3 ran without an effort parameter, using the API default, while the others ran at xhigh. Readers can also try a “guess the model” quiz built from 242 real, unedited management decisions at Firmulate.

From watching to testing your own business

The public experiment runs a live company with 13 synthetic employees and real money mechanics: burn of €105k per month against €2.3k MRR, a public cash countdown, more than 680 self-learned playbook rules and versioned workdays. It can be watched at firmulate.com. The company is synthetic; the choices and outcomes are visible.

For an enterprise, the next step is a pilot using a read-only export of its own business. That lets a company put crisis scenarios against its customers, pipeline and rules, then review a board report with model rankings and weak points in its playbooks. The pilot does not write back to real systems. That boundary matters for owners weighing how to learn what an AI might do before giving it responsibility in everyday operations.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

For a family company, the appeal of an AI assistant may be speed and extra capacity. Firmulate’s experiment suggests the harder question is whether it can act on its own analysis, respect authority and keep trust intact when the week turns difficult. A read-only pilot offers a way to examine those behaviors against your own business before they touch live systems.

Explore a Firmulate pilot for your business or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Can AI Read Your Files Before Making Decisions? A Live Test Shows the Difference Between Good and Bad AI at Work

A live experiment with AI models shows that reading deeply into company files can be the decisive factor in business success. The difference between winning and losing hinges on thorough understanding.

Hot-Weather Babywearing: How to Prevent Overheating

Keep your baby cool in hot weather with essential tips to prevent overheating and stay comfortable—and discover how to make summer babywearing safer.

Why Your AI Assistant Needs to Do More Than Just Talk: Lessons from a Transparent Benchmark

A live AI benchmark shows that trust, thoroughness, and integrity are key. Even the best models score just 95, while a do-nothing baseline gets 26 — partial progress counts.

Wrap vs Ring Sling vs Structured Carrier: Which One You’ll Actually Use

Meta description: “Many parents wonder which carrier suits their lifestyle best—wrap, ring sling, or structured—so keep reading to discover the one you’ll actually want to use.