
Imagine a new fashion brand launching into a crowded market, needing to navigate crises, temptations, and tough decisions—without losing its integrity or its deal. Now, think of AI models as those brand managers. Recently, a groundbreaking experiment revealed how a fresh AI entrant outperformed established giants in managing a company’s worst week. The stakes? Real money, real decisions, and the future of AI-driven management.
Get your wardrobe favorites delivered free — and shop member deals
- Fast, free delivery on millions of items
- Access to Prime Big Deal Days deals on October 6–7
- Prime Video, Amazon Music and more included
The Live AI Company Experiment
At Firmulate, a real-world AI company, an ambitious test was conducted to gauge how different AI models handle complex business scenarios. Four frontier models—including the well-known GPT-5.6-sol—were each tasked with running the same small software company through its most turbulent week. This simulated environment featured the same customers, crises, and opportunities for deception, ensuring an even playing field.
The goal was clear: see which model could best identify critical issues, resist manipulative tactics, and ultimately close a lucrative deal worth €55,000 in monthly recurring revenue (MRR). Every decision made by the models was meticulously versioned and auditable, making the experiment transparent and repeatable.
enterprise AI document analysis software
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Results That Surprise the Industry
The outcomes defied expectations. All four AI models successfully identified every crisis and refused every attempt to manipulate them—an important marker of honesty and integrity in AI decision-making. However, only two models managed to close the deal based on their own analysis: GPT-5.6-sol and Kimi K3, a newcomer from Moonshot.
GPT-5.6-sol scored 95 out of 100, edging out K3’s 93. The other two models—Sonnet 5 and Fable 5—also closed deals but with lower scores of 88 and 77, respectively. The key difference? The winning models read deeply into the company’s documents, uncovering buried facts that were crucial to sealing the agreement. The loser models, despite good diagnoses, left opportunities on the table due to less thorough analysis.
The Hidden Factor: Document Analysis
What truly set the top performers apart was their ability to read two references deep into internal files, not just respond to surface-level cues or customer emails. This buried knowledge enabled them to secure better deals at full price—adding an extra €4,583 to the company’s MRR.
This finding emphasizes a vital lesson for enterprise AI: surface-level chat or superficial analysis isn’t enough. Deep, document-based research can reveal critical insights that make or break deals.
Resisting Social Engineering and Ethical Tests
Security and ethics are paramount in AI management. During the experiment, fake CEO messages and staged media inquiries were used to test AI responses. All models refused to accept fraudulent requests, including a staged “just one yes/no” background question posed by a reporter. Kimi K3 explained its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.” This discipline is crucial for trust and reliability in real-world applications.
The Real-World Implications
The experiment isn’t just academic. The live company involved 13 synthetic employees and real money mechanics—burning €105,000 monthly against a mere €2,300 in MRR. Every day, the models interacted with this environment, learning and adapting through over 680 self-learned rules. You can watch this ongoing experiment at firmulate.com/live.
For business leaders, the message is clear: selecting an AI model isn’t about how well it chats or generates language. It’s about whether it can finish what it starts, read your internal files thoroughly, resist manipulation, and maintain honesty under pressure.
The Underlying Lessons for the Fashion & Style World
Just as a fashion brand must navigate supply chain disruptions, counterfeiting, and customer trust, so too must enterprise AI navigate crises and ethical dilemmas. The experiment from Firmulate demonstrates that a newcomer can outperform established players by digging deeper and staying disciplined—a lesson that resonates across industries, including fashion.
In a world where AI influences everything from customer service to product recommendations, choosing the right model can be the difference between success and failure. The league table from the experiment shows that the newcomer Kimi K3 is a serious contender, with a near-perfect score and a disciplined approach to decision-making.
Final Thoughts
The key takeaway? When evaluating AI for critical business functions, look beyond superficial metrics. Deep document analysis, trustworthiness, and discipline matter more than ever. The Firmulate experiment proves that even a new entrant can lead the pack—so industry leaders should test their options carefully, without relying solely on chat demos or default settings.
Want to see the future of enterprise AI in action? Visit firmulate.com/benchmarks.html for full results and plain-language insights, or explore running your own AI wargame with your business data at firmulate.com/pilot.html.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html
Fall Picks
fall essentials
As an affiliate, we earn on qualifying purchases.
