
Imagine a high-stakes fashion runway show where designers are judged solely on their sketches, not on how well they manage the backstage chaos. In the world of AI management, the real test isn’t just about how well these models generate answers—it’s about how they handle pressure, navigate crises, and stay honest when stakes are high. That’s the core lesson from a groundbreaking experiment where AI agents ran a real, money-losing software company through its worst week, revealing what truly matters in leadership under fire.
The Experiment: Going Beyond Chat Quality
At Firmulate, a unique live experiment puts AI models into the role of company managers. Four frontier models—ranging from GPT-5.6-sol to Sonnet 5—were tasked with running a small but real software company facing a succession of crises: customer churn, price hikes, down rounds, and PR disasters. Every decision was tracked, and the environment was designed to simulate real-world pressures, including fake CEO messages, reporter tricks, and critical file references buried deep in company documents.
AI management decision support tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
The Unexpected Results: Success Under Pressure
All four models identified every crisis and refused to be manipulated—an encouraging sign that they can grasp fundamental ethical boundaries. But only two managed to close the €55,000 deal their own analysis had earned, illustrating that spotting problems is not the same as executing solutions effectively. Interestingly, the crucial factor wasn’t just surface-level responses but their ability to read into the company’s own files—deep knowledge that made the difference between winning and losing the deal.
The Hidden Weakness: Deep Reading Matters
The real weakness was buried two document references deep in the company’s files. Those models that successfully read and interpreted these hidden details secured the full-priced deal, adding €4,583 MRR. Meanwhile, others missed this crucial information, leaving money on the table. This highlights a vital insight: in high-pressure decision-making, access to and understanding of the full context is essential, even more than quick replies or surface-level answers.
Trust and Integrity in AI Leadership
When asked about fake CEO messages or reporter tricks, all models refused to engage—an essential trait for any management AI that might one day steer real companies. Kimi K3’s reasoning was clear: ‘Treat the request as a suspected approval-bypass / possible impersonation.’ This shows that responsible AI management must prioritize honesty and security, not just efficiency or answer correctness.
The Real Business: Learning from Live, Money-Running Companies
The experiment isn’t just about testing AI in a sandbox. The live company, with 13 synthetic employees and real money mechanics burning €105k monthly against a €2.3k MRR, demonstrates that management quality can be measured and improved. Every workday, the company runs versioned decision rules, with transparent results available at firmulate.com/live. Watching this ongoing experiment, leaders can see whether AI models can truly handle the complexities of real-world business, not just generate neat answers.
Why This Matters for Fashion and Beyond
Just as fashion brands need to manage their supply chains, marketing, and PR under tight deadlines and scrutiny, AI tools integrated into business operations must demonstrate their ability to manage crises, read deeply into internal data, and uphold integrity. The lesson? Success isn’t just about the quality of the chat or answers. It’s about the management and ethical behavior under pressure—traits that aren’t visible in demos but are critical for real-world impact.
The Takeaway: The Future of AI in Management
As AI continues to integrate into business functions, companies should look beyond surface scores or chat fluency. The true measure is whether these systems can finish what they start, read the full context, and stay honest when stakes are high. The firms that adopt these insights will better navigate complex scenarios—whether a price war or a PR crisis—and emerge with value intact. The ongoing Firmulate experiment offers a real, transparent look into this future, showing that management quality is the next frontier in AI evaluation.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html