OpenAI says two Responses API settings sharply improved GPT-5.6 Sol on ARC-AGI-3
OpenAI says retained reasoning and compaction lifted GPT-5.6 Sol’s ARC-AGI-3 score nearly 3x, highlighting how harness design shapes AI benchmarks.
OpenAI says retained reasoning and compaction lifted GPT-5.6 Sol’s ARC-AGI-3 score nearly 3x, highlighting how harness design shapes AI benchmarks.
NVIDIA says LangChain-tuned Nemotron 3 Ultra reached top open-model agent benchmark results, highlighting cheaper open stacks for enterprise AI.
Arena, the Berkeley-born startup behind one of the most widely watched crowdsourced AI model leaderboards, told TechCrunch it has reached $100 million in annualized run-rate revenue roughly eight months after launching its commercial evaluation product. The milestone suggests that AI benchmarking and post-training feedback are becoming a business in their own right, not just a research exercise, though the company says its revenue is consumption-based rather than contracted recurring ARR.
Latest News and Analysis in AI Performance