Install
The New Stack is a media platform for the people who build and manage software the world relies on. We provide context and explanation of at-scale technologies to advance knowledge and create conversations through our coverage of modern architectures, components of the software development life cycle, and operations to
- 153articles · 30d
- 18+ hour agolatest article
- Aug 14, 2026earliest in window
- 96%with images
- 87avg words
- Science & Technology 140
- Software Dev. 101
- Computers & Electronics 96
- News 38
- Software 22
- Science & Nature 12
- Economy, Business & Finance 9
- Finance & Business 9
Please confirm you are human
This browser or connection looks automated. Press and continuously hold the control for 3 seconds to enable Google-hosted web results and, when separately allowed, AI-assisted answers.
A successful check enables 100 search requests. Interactive access does not authorize scraping, systematic collection, or reuse of search output.
News
Claude did best on a new benchmark for agents that build agents. It still passed fewer than a quarter of the tests.
3+ day, 13+ hour ago (492+ words) Sierra has open-sourced Hyper-𝜏-bench, a follow-up to its 2024 τ-bench that tests how well AI agents can build other agents....
K2 Horizon just shipped as six new fully open models — developers aren't fully convinced
3+ day, 21+ hour ago (632+ words) Based in the Emirati capital, Abu Dhabi, the Institute of Foundation Models (IFM) introduced K2 Horizon last week. This group of six AI foundation models, ranging from 0.9 billion to 375 billion parameters, is claimed to be the “largest fully open-source fleet of…...
DeepSeek is hiring 150 engineers, and none of them will touch a model
4+ day, 14+ hour ago (172+ words) DeepSeek announced 150 backend engineering roles to scale DSec, the sandbox system running hundreds of thousands of concurrent agent environments on a single cluster....
AI agents are creating more work, not less — and OpenAI's own numbers back it up
5+ day, 14+ hour ago (317+ words) OpenAI's agents log 3.1 workdays for every human one, but more hours don't mean faster breakthroughs — and supervision may be the real bottleneck....
OpenAI's new model costs 2.5x more per token — and developers are saving money anyway
5+ day, 17+ hour ago (300+ words) OpenAI says GPT-6 Astra at low reasoning beats Sol at high — and can cost less per task despite 2.5x token prices. Here's what the benchmarks show...
“Twenty years of brand building simply froze in time”: How coding agents select their tools of choice
5+ day, 21+ hour ago (311+ words) A new study shows Claude Code, Cursor, and Codex pick wildly different tools for the same job - and vendors are racing to influence the choice....
GPT-6 Astra aced the hardest AI benchmark. The asterisk matters more than the score.
1+ week, 2+ day ago (217+ words) OpenAI says GPT-6 Astra scored 98.6% on ARC-AGI-3, up from 7.8% six months ago. But undisclosed test settings and opaque reasoning complicate the AGI question....
The systems guide to production token optimization
1+ week, 2+ day ago (542+ words) Scale AI systems efficiently by treating token optimization as a distributed systems and hardware utilization challenge....
It cost $33 to build a virtual Union Square. Here's what the agents got wrong.
1+ week, 2+ day ago (152+ words) AI coding agents built a 3D browser replica of San Francisco's Union Square in two hours, then used Playwright screenshots to catch visual errors no test could....
“Google was ahead only a few hours”: Muse Spark 1.3 edges out Gemini as Meta claims its biggest coding leap yet
1+ week, 2+ day ago (682+ words) ...