
AI models now power all manner of agents, from coding assistants that write and debug software to customer service systems
The post Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests. appeared first on The New Stack.



























