Forget lab tests, someone decided to put AI to the real-world test. A new benchmark had 123 people judge top models directly.
In 30 seconds
01A test in India evaluated GPT-4, Gemini Pro, and Mixtral 8x7B on coding, reasoning, and creativity.
02GPT-4 achieved the best overall results, particularly in code and logic, despite being the most expensive.
→
💡
What this means for you
This means that, for now, paid models like GPT-4 still offer the best performance, but open-source models are catching up. For users, the choice will depend on how much they're willing to spend and for which specific tasks they use AI.
Tired of just listening? Soon we could all be mini-Mozarts, or at least DJs. Suno's CEO has clear ideas on how AI will change music.
·2 min·Beginner
03Gemini Pro excelled in creative writing and moderation, while the open-source Mixtral showed good potential.
0101
Who Put AI to the Test and How?
Someone decided to step outside the usual synthetic benchmarks, the ones only tech geeks understand. They enlisted 123 people in India to evaluate the performance of three leading AI models: OpenAI's GPT-4, Google's Gemini Pro, and the open-source Mixtral 8x7B. This experiment was a submission for the Kaggle Benchmarking Challenge.
Participants judged the models across four crucial areas. These ranged from generating Python code to solving logic problems. Then there was creative writing, producing short stories, and content moderation. Quite a mix, really, to see who performs best in the field, not just on spreadsheets. All this to understand how these AIs behave when an average person uses them.
0202
Who Won the "Human" Race?
In the conducted benchmark, GPT-4 showed the best overall performance, particularly in code generation and logical reasoning. Not a huge surprise, given it's the flagship model and not exactly cheap to use. But it's nice to have "human" confirmation of that.
📬 Enjoying this article?
Get the best AI news every week, straight to your inbox.
Google's Gemini Pro, on the other hand, held its own, shining especially in creative writing and content moderation. Google has always had a soft spot for text, after all. And Mixtral 8x7B? The open-source model showed promising performance, though still a step behind the proprietary giants. It's a bit like the David in this situation, but with an arsenal still needing refinement. It demonstrated that open source is gaining ground, albeit slowly.
0303
Why Does This "Indian Test" Matter to Us?
This type of evaluation, based on real people and not just automated metrics, offers a valuable perspective. It tells us how AI is perceived and used in everyday life. It's a great complement to "lab tests" that often miss human nuances. It's not just about numbers, but about usability and perceived quality.
Sure, 123 people in India aren't the whole world, but it's an interesting start. The methodology involved 123 people in India evaluating the models' responses, providing human-experience-based analysis. It makes us realize that while proprietary models still lead, the open-source push is strong. And one day, perhaps, even the underdog David will beat Goliath. Or at least give him a good run for his money.
Making big-budget films on a shoestring? It sounds like sci-fi, but someone's actually doing it. A new studio is betting big on AI to produce blockbusters with tiny budgets.