Artificial Intelligence
May. 18th, 2025 04:18 pm![[personal profile]](https://www.dreamwidth.org/img/silk/identity/user.png)
Professors Staffed a Fake Company Entirely With AI Agents
As Business Insider first reported, the results were dismal. The best-performing model was Anthropic's Claude 3.5 Sonnet, which struggled to finish just 24 percent of the jobs assigned to it. The study's authors note that even this meager performance is prohibitively expensive, averaging nearly 30 steps and a cost of over $6 per task.
Google's Gemini 2.0 Flash, meanwhile, averaged a time-consuming 40 steps per finished task, but only had an 11.4 percent rate of success — the second highest of all the models. The worst AI employee was Amazon's Nova Pro v1, which finished just 1.7 percent of its assignments at an average of almost 20 steps.
While corporations may wish to replace human employees with software, it is not yet feasible for complex tasks. Only the simplest jobs are really at risk.
As Business Insider first reported, the results were dismal. The best-performing model was Anthropic's Claude 3.5 Sonnet, which struggled to finish just 24 percent of the jobs assigned to it. The study's authors note that even this meager performance is prohibitively expensive, averaging nearly 30 steps and a cost of over $6 per task.
Google's Gemini 2.0 Flash, meanwhile, averaged a time-consuming 40 steps per finished task, but only had an 11.4 percent rate of success — the second highest of all the models. The worst AI employee was Amazon's Nova Pro v1, which finished just 1.7 percent of its assignments at an average of almost 20 steps.
While corporations may wish to replace human employees with software, it is not yet feasible for complex tasks. Only the simplest jobs are really at risk.
Well ...
Date: 2025-05-19 04:48 pm (UTC)Computers and humans are just good at doing totally different things. If it's a task based on logic, precision, or math then a computer will often excel. But if it requires creativity, intuition, or making something from scratch then you're better off with a human. So while we may see AI stick around for certain things like telling you which of 20,000 lidar images have ruins in them, it is not good enough to do most human jobs.