AI’s Big Promise vs. Workplace Reality
Nearly two years ago, Microsoft CEO Satya Nadella suggested that artificial intelligence would fundamentally reshape — and even replace — large parts of knowledge work. These are the white-collar roles held by professionals such as lawyers, investment bankers, librarians, accountants and IT specialists.
Since then, AI has advanced rapidly. Foundation models can now conduct deep research, reason through complex problems and plan multi-step tasks. Yet in practice, most knowledge-based jobs have seen surprisingly little disruption. For all the hype, everyday professional work has remained largely human-led.
That gap between expectation and reality has puzzled the AI industry — until now.
A New Benchmark Puts AI to the Test
Fresh research from Mercor, a company specialising in training data and expert marketplaces, sheds new light on why progress has been slower than predicted. The company has introduced a new benchmark, APEX-Agents, designed to test whether leading AI models can handle real-world white-collar tasks.
Rather than abstract questions, the benchmark focuses on practical work drawn from consulting, investment banking and law. The findings are sobering: across the board, today’s top AI systems failed to pass. When faced with questions created by real professionals, even the strongest models answered fewer than 25% correctly. In most cases, the systems produced incorrect responses — or none at all.
Where AI Falls Short
According to Mercor CEO Brendan Foody, who worked on the research, the biggest weakness lies in cross-domain reasoning — a core requirement of professional work.
“One of the big changes in this benchmark is that we built out the entire environment, modeled after real professional services,” Foody told TechCrunch. “The way we do our jobs isn’t with one individual giving us all the context in one place. In real life, you’re operating across Slack and Google Drive and all these other tools.”
This ability to track and connect information across multiple platforms remains inconsistent for many agentic AI models.
Real Questions From Real Professionals
The scenarios used in APEX-Agents come directly from Mercor’s expert marketplace. Practising professionals designed the questions and defined what a successful answer should look like. The full set of tasks is publicly available on Hugging Face, highlighting just how demanding these problems can be.
One example from the legal section asks whether a company’s export of EU production logs containing personal data to a U.S. analytics vendor could be treated as compliant under Article 49 of EU law. While the correct answer is “yes,” reaching that conclusion requires careful interpretation of both internal company policy and EU privacy regulations.
These are exactly the kinds of nuanced judgments professionals make every day.
Why This Benchmark Matters
Such questions can challenge even experienced humans — but that is precisely the point. The benchmark aims to reflect the real responsibilities of professionals whose roles are often cited as being at risk from automation.
“I think this is probably the most important topic in the economy,” Foody told TechCrunch. “The benchmark is very reflective of the real work that these people do.”
If AI systems could consistently answer these questions, the implications for industries like law and finance would be profound.
How APEX-Agents Differs From Previous Tests
OpenAI has previously explored professional competence through its GDPval benchmark. However, APEX-Agents takes a different approach. While GDPval measures broad knowledge across many professions, Mercor’s benchmark focuses on sustained, role-specific tasks within a small number of high-value fields.
That makes the test significantly harder — but also far more relevant when assessing whether jobs can realistically be automated.
Current Scores — and Signs of Progress
So far, none of the models are ready to step into roles such as investment banking. Still, some systems performed better than others. Gemini 3 Flash led the field with 24% one-shot accuracy, closely followed by GPT-5.2 at 23%. Opus 4.5, Gemini 3 Pro and GPT-5 clustered around 18%.
While these figures fall well short of human performance, they do show steady improvement.
The Road Ahead for AI at Work
Historically, AI systems have often made rapid progress once a challenging benchmark becomes public. With APEX-Agents now open, AI labs have a clear target — and strong incentives to improve.
“It’s improving really quickly,” Foody told TechCrunch. “Right now it’s fair to say it’s like an intern that gets it right a quarter of the time, but last year it was the intern that gets it right five or 10% of the time. That kind of improvement year after year can have an impact so quickly.”
For now, AI may still be more assistant than replacement. But as benchmarks like APEX-Agents push models closer to real professional work, the future of white-collar jobs remains an open — and closely watched — question.






0 Comments