📊 Full opportunity report: Beyond The Demo: The AI Leaderboard That Defines Success on ThorstenMeyerAI.com — validation score, market gap, and execution plan.
TL;DR
Firmulate’s live AI benchmark tests models managing a simulated company during a crisis week. Results show management skills, trust, and decision-making are critical, not just chat quality. This shifts AI evaluation toward real-world management capabilities.
Firmulate has conducted a groundbreaking live experiment where AI models are tasked with managing a simulated company during its worst week. The results, announced in July 2026, demonstrate that management quality, trustworthiness, and decision-making are crucial metrics for AI success, beyond traditional chat or coding benchmarks.
The experiment involved five AI models competing in a simulated crisis environment, with a final ranking: gpt-5.6-sol leading at 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A baseline scored 26, emphasizing the challenge. For more on how AI models are evaluated in real-world scenarios, see the original analysis. Crucially, the models were evaluated not only on crisis detection and response but on their ability to investigate, communicate, and finalize decisions within a trust framework that penalized breaches. For insights into trust and decision-making in AI systems, see the original analysis.
Despite all models identifying crises and resisting manipulation attempts—such as fake CEO messages—they failed to close deals or execute decisions effectively. This highlights the importance of comprehensive AI benchmarking, as discussed in the original analysis. For example, only two models signed a €55,000 deal after analysis, while others identified the opportunity but did not follow through. The experiment revealed that superficial responses or extensive analysis do not guarantee effective management outcomes, especially when critical facts are overlooked or communication channels are misused.
Management Skills as a New AI Benchmark
This experiment underscores that management quality, trust, and decision execution are essential metrics for AI evaluation, especially in operational contexts. It suggests that AI systems must be tested on their ability to prioritize, read organizational context, escalate appropriately, and maintain honesty under pressure—capabilities vital for real-world management but overlooked in traditional benchmarks.
For organizations, this means shifting focus from chat or coding prowess to how models handle complex, consequence-driven tasks. The findings challenge the assumption that more detailed analysis or activity equates to better management, highlighting that effective decision-making within organizational constraints is the true measure of AI readiness for operational roles.
As an affiliate, we earn on qualifying purchases.
From Traditional Benchmarks to Real-World Management
Historically, AI evaluation has centered on technical performance—coding accuracy, chat fluency, or game scores. These benchmarks, while useful, do not capture how models perform in managing real organizations under duress. The Firmulate experiment introduces a new approach: live, consequence-based testing within a simulated business environment, replicating the pressures and trust requirements of actual management.
Since its launch, the experiment has involved a simulated company with real money mechanics, self-learned rules, and versioned decisions. The goal is to observe whether AI models can prioritize, read organizational files, escalate issues, and maintain honesty, thus providing a more meaningful measure of operational competence.
“Traditional benchmarks only measure superficial capabilities. Our live experiment reveals that true management involves trust, prioritization, and accountability—areas where AI models still have much to prove.”
— Thorsten Meyer, founder of Firmulate
AI decision-making tools for business
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Unclear Aspects of AI Management Performance
While the experiment shows promising results, it remains unclear how these models will perform in diverse real-world organizations with different cultures, structures, and crisis types. The long-term reliability of AI in sustained management roles, especially under unpredictable or high-stakes conditions, has yet to be established. Additionally, the impact of model training, configuration, and contextual understanding on performance variability is still being studied.
AI trustworthiness assessment tools
As an affiliate, we earn on qualifying purchases.
As an affiliate, we earn on qualifying purchases.
Next Steps for AI Management Evaluation and Adoption
Following the July 2026 results, firms are expected to adopt similar live testing frameworks tailored to their specific operational contexts. Further research will focus on expanding the scope of scenarios, testing models over longer periods, and integrating human oversight. The industry will also explore developing standardized management benchmarks that reflect real organizational challenges, emphasizing trust, decision quality, and accountability.
As AI models improve, organizations will likely pilot these systems in limited roles, gradually increasing responsibility while monitoring performance and trust metrics. The ultimate goal is to establish a comprehensive evaluation system that predicts AI effectiveness in managing complex, real-world operations.
As an affiliate, we earn on qualifying purchases.
Key Questions
What makes this AI benchmark different from traditional tests?
This benchmark evaluates AI models on their ability to manage a simulated company during crises, focusing on decision-making, trustworthiness, escalation, and execution—beyond just generating responses or code.
Can these AI models replace human managers?
Currently, the models are tested as decision-support tools in controlled simulations. Their performance in real management roles requires further validation, especially in unpredictable environments.
What are the main weaknesses identified in the models?
While models can identify crises and resist manipulation, they often fail to follow through on decisions, escalate properly, or prioritize effectively, revealing gaps in execution and organizational understanding.
Will this lead to new AI management standards?
Yes, the experiment suggests that management skills—trust, decision quality, escalation—should become core metrics in AI evaluation, influencing future benchmarks and deployment strategies.
How can organizations use this testing approach?
Businesses can run similar live simulations internally, testing AI agents against their specific operational scenarios to assess readiness before full deployment.
Source: ThorstenMeyerAI.com