Here come the first real tests to see if artificial intelligence is actually ready to work in corporate IT departments. Spoiler: not quite. The best AIs out there can't get past 50% success.
In 30 seconds
01ITBench-AA benchmark shows top AI models (GPT-4, Claude) fail above 50% on real IT workplace tasks.
02
→
💡
What this means for you
If you work at a company and someone promised you that AI would solve your IT problems: wait. It's not time yet. If you work in IT and worry about being replaced by a machine: sleep easy, at least for now.
Thought slapping 'AI' next to a company name guaranteed its stock would soar? Well, the market had a bitter surprise this year.
·1 min·2·Beginner
AI struggles with sequential processes, long-term memory, and self-correction mechanisms.
03Test deflates hype: enterprise AI automation is not reliable yet, major improvements needed.
Picture this: you tell your company's AI "Hey, go manage my server for me." Yeah, don't try that experiment. Enter ITBench-AA, the first real benchmark testing what "agentic" AI models (the kind that should act autonomously) actually do when facing real-world enterprise IT tasks. This test came from a collaboration between IBM Research and Artificial Analysis, so it's not exactly a joke benchmark.
The verdict? Even the best AI models available right now — we're talking GPT-4, Claude, and the usual suspects — don't crack 50%. It's like telling these digital titans: "Look, you're amazing at benchmarks, but when it comes to actually solving a problem, you ghost me halfway through." On tougher tasks, it gets downright ugly.
What they tested are real-world tasks: configuring environments, fixing network errors, managing permissions, integrating systems. Not sci-fi stuff — it's what IT managers ask their teams to do every single day. To get a sense of the difficulty, just imagine an actual IT expert tackling these same tasks. Their success rate? Way higher. So AI is still far from being a trustworthy colleague.
📬 Enjoying this article?
Get the best AI news every week, straight to your inbox.
Why does it matter? Because over the past few years there's been endless hype about AI and enterprise automation. Every vendor in existence was screaming "AI will revolutionize your processes." Well, ITBench-AA is the first benchmark that says: "Hold up, let's actually see what we're dealing with." It's not a smack against AI, it's just a cold shower — a useful one, for once.
On one hand, the test is bad news for anyone who's already thrown billions at "intelligent" solutions. On the other hand, it's a chance to understand what needs fixing. The researchers pinpointed the weak spots: AI struggles when tasks are sequential, when memory is required (remembering what it did five steps ago), and when it needs to correct its own mistakes — the exact stuff a human would do naturally.
Bottom line: don't go buying that digital robot thinking it'll handle all your IT magic. But now we at least know exactly where we stand and where we're going. ITBench-AA isn't bad news for AI, it's honest news.
While the tech world was buzzing about OpenAI, Anthropic made its move. They just dropped Opus 5, a model they claim is almost as good as their legendary Fable 5.