AIThis post was created with the assistance of artificial intelligence (AI).

📊 Full opportunity report: Beyond The Demo: The AI Leaderboard That Defines Success on ThorstenMeyerAI.com — validation score, market gap, and execution plan.

TL;DR

Firmulate’s live AI benchmark tests models managing a simulated company during a crisis week. Results show management skills, trust, and decision-making are critical, not just chat quality. This shifts AI evaluation toward real-world management capabilities.

Firmulate has conducted a groundbreaking live experiment where AI models are tasked with managing a simulated company during its worst week. The results, announced in July 2026, demonstrate that management quality, trustworthiness, and decision-making are crucial metrics for AI success, beyond traditional chat or coding benchmarks.

The experiment involved five AI models competing in a simulated crisis environment, with a final ranking: gpt-5.6-sol leading at 95 points, followed by Kimi K3 at 93, Sonnet 5 at 88, Fable 5 at 77, and Opus 4.8 at 73. A baseline scored 26, emphasizing the challenge. For more on how AI models are evaluated in real-world scenarios, see the original analysis. Crucially, the models were evaluated not only on crisis detection and response but on their ability to investigate, communicate, and finalize decisions within a trust framework that penalized breaches. For insights into trust and decision-making in AI systems, see the original analysis.

Despite all models identifying crises and resisting manipulation attempts—such as fake CEO messages—they failed to close deals or execute decisions effectively. This highlights the importance of comprehensive AI benchmarking, as discussed in the original analysis. For example, only two models signed a €55,000 deal after analysis, while others identified the opportunity but did not follow through. The experiment revealed that superficial responses or extensive analysis do not guarantee effective management outcomes, especially when critical facts are overlooked or communication channels are misused.

At a glance
reportWhen: ongoing, with final results published i…
The developmentFirmulate launched a live experiment testing AI models managing a simulated company during a week of crises, highlighting management performance as a new benchmark.
Crypto market snapshot
Fear & Greed Index
71/100 — Greed
Bitcoin BTC$78,749▼ 0.3%
Ethereum ETH$2,489▲ 1.0%
Tether USDT$0.9998▼ 0.0%
BNB BNB$705.02▲ 0.8%
XRP XRP$1.4▼ 2.4%
USDC USDC$0.9999▼ 0.0%
Solana SOL$101.07▲ 4.2%
TRON TRX$0.3349▼ 0.9%
Live data · CoinGecko · alternative.me (24h change)

Management Skills as a New AI Benchmark

This experiment underscores that management quality, trust, and decision execution are essential metrics for AI evaluation, especially in operational contexts. It suggests that AI systems must be tested on their ability to prioritize, read organizational context, escalate appropriately, and maintain honesty under pressure—capabilities vital for real-world management but overlooked in traditional benchmarks.

For organizations, this means shifting focus from chat or coding prowess to how models handle complex, consequence-driven tasks. The findings challenge the assumption that more detailed analysis or activity equates to better management, highlighting that effective decision-making within organizational constraints is the true measure of AI readiness for operational roles.

Amazon

AI management simulation software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

From Traditional Benchmarks to Real-World Management

Historically, AI evaluation has centered on technical performance—coding accuracy, chat fluency, or game scores. These benchmarks, while useful, do not capture how models perform in managing real organizations under duress. The Firmulate experiment introduces a new approach: live, consequence-based testing within a simulated business environment, replicating the pressures and trust requirements of actual management.

Since its launch, the experiment has involved a simulated company with real money mechanics, self-learned rules, and versioned decisions. The goal is to observe whether AI models can prioritize, read organizational files, escalate issues, and maintain honesty, thus providing a more meaningful measure of operational competence.

“Traditional benchmarks only measure superficial capabilities. Our live experiment reveals that true management involves trust, prioritization, and accountability—areas where AI models still have much to prove.”

— Thorsten Meyer, founder of Firmulate

Amazon

AI decision-making tools for business

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Unclear Aspects of AI Management Performance

While the experiment shows promising results, it remains unclear how these models will perform in diverse real-world organizations with different cultures, structures, and crisis types. The long-term reliability of AI in sustained management roles, especially under unpredictable or high-stakes conditions, has yet to be established. Additionally, the impact of model training, configuration, and contextual understanding on performance variability is still being studied.

Amazon

AI trustworthiness assessment tools

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Next Steps for AI Management Evaluation and Adoption

Following the July 2026 results, firms are expected to adopt similar live testing frameworks tailored to their specific operational contexts. Further research will focus on expanding the scope of scenarios, testing models over longer periods, and integrating human oversight. The industry will also explore developing standardized management benchmarks that reflect real organizational challenges, emphasizing trust, decision quality, and accountability.

As AI models improve, organizations will likely pilot these systems in limited roles, gradually increasing responsibility while monitoring performance and trust metrics. The ultimate goal is to establish a comprehensive evaluation system that predicts AI effectiveness in managing complex, real-world operations.

Amazon

AI crisis management software

As an affiliate, we earn on qualifying purchases.

As an affiliate, we earn on qualifying purchases.

Key Questions

What makes this AI benchmark different from traditional tests?

This benchmark evaluates AI models on their ability to manage a simulated company during crises, focusing on decision-making, trustworthiness, escalation, and execution—beyond just generating responses or code.

Can these AI models replace human managers?

Currently, the models are tested as decision-support tools in controlled simulations. Their performance in real management roles requires further validation, especially in unpredictable environments.

What are the main weaknesses identified in the models?

While models can identify crises and resist manipulation, they often fail to follow through on decisions, escalate properly, or prioritize effectively, revealing gaps in execution and organizational understanding.

Will this lead to new AI management standards?

Yes, the experiment suggests that management skills—trust, decision quality, escalation—should become core metrics in AI evaluation, influencing future benchmarks and deployment strategies.

How can organizations use this testing approach?

Businesses can run similar live simulations internally, testing AI agents against their specific operational scenarios to assess readiness before full deployment.

Source: ThorstenMeyerAI.com

Nothing in this article is financial or investment advice. Cryptocurrency and precious-metal investments carry significant risk — do your own research and consider a licensed advisor.
You May Also Like

Angel Reese’s Astonishing Rise—The Unexpected Millions of a College Athlete

Prepare to be inspired by Angel Reese’s astonishing rise, as she transforms college athletics into a lucrative platform—discover how she did it!

Forward-Deployed Engineer Economics 2.0: The Unit Economics Math, Six Months Later

Six months after initial analysis, FDE economics reveal profitability at enterprise scale but risks at lower levels, impacting AI lab scaling strategies.