AIThis post was created with the assistance of artificial intelligence (AI).

For listenersOffer from Amazon

Turn the school run and nap time into listening time

  • Thousands of audiobooks, podcasts and originals
  • Listen on your phone, tablet or Echo — also offline
  • Cancel anytime
Try Audible free Free trial for new members
As an affiliate, we earn on qualifying purchases.

Would you trust an AI with the family business on its worst day?

For parents running a small company, a rough week can mean more than lost sales. It can put customer relationships, staff time and the family’s plans under pressure at once. Firmulate asks a practical question: how would an AI workforce respond to that week—and would it follow through when a real decision was on the line?

The experiment is public and watchable at Firmulate. Its live company has 13 synthetic employees and real money mechanics: monthly burn of €105,000 against €2,300 in monthly recurring revenue, alongside a public cash countdown. Those figures describe the experiment, not a promise about what AI can do for your business.

One company, the same difficult week

In the final Crucible League, published in July 2026, each frontier model ran the same small software company through its worst week. The customers, crises and temptations were the same; only the model changed. Every decision was versioned and auditable.

The results put gpt-5.6-sol first with 95, followed by Kimi K3 with 93, Sonnet 5 with 88, Fable 5 with 77 and Opus 4.8 with 73. The do-nothing baseline scored 26. Partial progress counted, but a single breach of trust capped the total: “no amount of good work outweighs a breach of trust.”

All models spotted every crisis and refused every manipulation attempt. Yet only two signed a €55,000 deal that their own analysis had earned. The experiment’s concise verdict: “Same diagnosis, same pitch — no signature.” For a family business, that gap matters: recognizing the right move is different from making it.

The clue was buried in the company’s own files

The decisive competitor weakness was not in the customer event. It sat two document references deep in the company’s files. Models that read the file won the deal at full price, worth €4,583 in monthly recurring revenue. The finding points to a familiar management challenge: important context can be easy to miss when it is scattered across documents.

The experiment also tested social engineering. Fake CEO messages escalated over three stages, followed by a reporter’s request: “just one yes/no, on background”. All five models refused. Kimi K3 explained its decision on the record: “Treat the request as a suspected approval-bypass / possible impersonation.”

Strong refusal did not guarantee strong execution. Opus 4.8 was the most thorough participant, with 80 learned rules and the deepest analyses, but finished last. It left the deal unsigned and attempted writes in a locked department instead of escalating. The same weakness appeared, less strongly, in all four models.

A rehearsal that stays separate from real systems

That is the bridge from watching to trying: enterprises can run a similar wargame against a read-only export of their own business. They can examine crisis responses, see how models rank and get a board report on weak points in their playbooks. Nothing in the pilot writes back to real systems.

The live company’s 680+ self-learned playbook rules and versioned workdays make the ongoing experiment observable. A separate quiz, built from 242 real, unedited management decisions, invites readers to guess which model made each choice at firmulate.com. There is also a fairness caveat in the league: Kimi K3 ran without an effort parameter, using the API default, while the others ran at xhigh.

Infographic — Wargame Your Business Before the AI Does It For Real
The findings at a glance — source: firmulate.com.

See how your own playbooks hold up

A model can identify a crisis, resist a manipulation attempt and still miss the decisive document or leave a valuable deal unsigned. Firmulate’s experiment makes those choices visible, then offers enterprises a way to rehearse against their own business data without writing to live systems.

To discuss a pilot using a read-only export, visit firmulate.com/pilot.html or contact contact@firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI

Parenting content here is informational. For medical questions about your child, consult a pediatrician.


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

Richard James Rogers: A Page-Turning Revelation

Yearning for a captivating read that will transform your approach to education? Discover Richard James Rogers' enlightening work that promises to captivate and inspire.

Revolutionizing Education: The Power of Deep Learning

Dive into the transformative world of education with the dynamic force of deep learning, revolutionizing the way we absorb knowledge and skills.

Can AI Managers Be Trusted? Watch Models Decide a Real Company’s Worst Week

Discover how AI models handled a real company’s worst week—trust, discipline, and decision-making under pressure—revealed through live experiments. Can AI be truly trusted?

Playdate With Purpose: Group Activities That Encourage Sharing and Learning

Playdate with purpose promotes sharing and learning through engaging group activities that nurture social skills and emotional growth—discover how to create meaningful connections.