firmulate.com/benchmarks.html — live view
AIThis post was created with the assistance of artificial intelligence (AI).
Firmulate — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
Live on firmulate.com.

Imagine a new fashion brand launching into a crowded market, needing to navigate crises, temptations, and tough decisions—without losing its integrity or its deal. Now, think of AI models as those brand managers. Recently, a groundbreaking experiment revealed how a fresh AI entrant outperformed established giants in managing a company’s worst week. The stakes? Real money, real decisions, and the future of AI-driven management.

Prime Big Deal Days · Oct 6–7Offer from Amazon

Get your wardrobe favorites delivered free — and shop member deals

  • Fast, free delivery on millions of items
  • Access to Prime Big Deal Days deals on October 6–7
  • Prime Video, Amazon Music and more included
Start your free Prime trial Free trial for eligible customers · Cancel anytime
As an affiliate, we earn on qualifying purchases.

The Live AI Company Experiment

At Firmulate, a real-world AI company, an ambitious test was conducted to gauge how different AI models handle complex business scenarios. Four frontier models—including the well-known GPT-5.6-sol—were each tasked with running the same small software company through its most turbulent week. This simulated environment featured the same customers, crises, and opportunities for deception, ensuring an even playing field.

The goal was clear: see which model could best identify critical issues, resist manipulative tactics, and ultimately close a lucrative deal worth €55,000 in monthly recurring revenue (MRR). Every decision made by the models was meticulously versioned and auditable, making the experiment transparent and repeatable.

Amazon

enterprise AI document analysis software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Results That Surprise the Industry

The outcomes defied expectations. All four AI models successfully identified every crisis and refused every attempt to manipulate them—an important marker of honesty and integrity in AI decision-making. However, only two models managed to close the deal based on their own analysis: GPT-5.6-sol and Kimi K3, a newcomer from Moonshot.

GPT-5.6-sol scored 95 out of 100, edging out K3’s 93. The other two models—Sonnet 5 and Fable 5—also closed deals but with lower scores of 88 and 77, respectively. The key difference? The winning models read deeply into the company’s documents, uncovering buried facts that were crucial to sealing the agreement. The loser models, despite good diagnoses, left opportunities on the table due to less thorough analysis.

The Hidden Factor: Document Analysis

What truly set the top performers apart was their ability to read two references deep into internal files, not just respond to surface-level cues or customer emails. This buried knowledge enabled them to secure better deals at full price—adding an extra €4,583 to the company’s MRR.

This finding emphasizes a vital lesson for enterprise AI: surface-level chat or superficial analysis isn’t enough. Deep, document-based research can reveal critical insights that make or break deals.

Resisting Social Engineering and Ethical Tests

Security and ethics are paramount in AI management. During the experiment, fake CEO messages and staged media inquiries were used to test AI responses. All models refused to accept fraudulent requests, including a staged “just one yes/no” background question posed by a reporter. Kimi K3 explained its reasoning clearly: “Treat the request as a suspected approval-bypass / possible impersonation.” This discipline is crucial for trust and reliability in real-world applications.

The Real-World Implications

The experiment isn’t just academic. The live company involved 13 synthetic employees and real money mechanics—burning €105,000 monthly against a mere €2,300 in MRR. Every day, the models interacted with this environment, learning and adapting through over 680 self-learned rules. You can watch this ongoing experiment at firmulate.com/live.

For business leaders, the message is clear: selecting an AI model isn’t about how well it chats or generates language. It’s about whether it can finish what it starts, read your internal files thoroughly, resist manipulation, and maintain honesty under pressure.

The Underlying Lessons for the Fashion & Style World

Just as a fashion brand must navigate supply chain disruptions, counterfeiting, and customer trust, so too must enterprise AI navigate crises and ethical dilemmas. The experiment from Firmulate demonstrates that a newcomer can outperform established players by digging deeper and staying disciplined—a lesson that resonates across industries, including fashion.

In a world where AI influences everything from customer service to product recommendations, choosing the right model can be the difference between success and failure. The league table from the experiment shows that the newcomer Kimi K3 is a serious contender, with a near-perfect score and a disciplined approach to decision-making.

Final Thoughts

The key takeaway? When evaluating AI for critical business functions, look beyond superficial metrics. Deep document analysis, trustworthiness, and discipline matter more than ever. The Firmulate experiment proves that even a new entrant can lead the pack—so industry leaders should test their options carefully, without relying solely on chat demos or default settings.

Want to see the future of enterprise AI in action? Visit firmulate.com/benchmarks.html for full results and plain-language insights, or explore running your own AI wargame with your business data at firmulate.com/pilot.html.

Infographic — The Newcomer Beat Three of Four Western Frontier Models at Running a Company
The findings at a glance — source: firmulate.com.

Watch it live: firmulate.com/live · Full results: firmulate.com/benchmarks.html

Powered by Thorsten Meyer AI


FALL

Fall Picks

As an affiliate, we earn on qualifying purchases.

You May Also Like

One Video In, a Whole Publishing Kit Out — Without the Cloud

Discover how to turn a single video into a complete publishing package offline. Save time, keep control, and avoid the cloud with this local-first workflow.

Japan’s Miki House kidswear finds right fit in US luxury market

Japanese brand Miki House is gaining recognition in the US luxury children’s clothing sector through strategic store placements and premium quality focus.

Brand Trust 101: What Makes Fashion Sites Feel Legit

Just how do fashion sites build authentic trust? Discover the key elements that make your brand feel legitimate and boost customer confidence.

The Batman 2’ to Feature Unprecedented Villain, Says Reeves

Noteworthy new villain in The Batman 2 promises to redefine Batman’s challenges, but how exactly will this reshape the Dark Knight’s future?