OpenAI released GPT-6 "Astra" on Thursday 3 September, describing it as "a new frontier on computer and browser use" and - the line that matters for this publication - "the best model for software engineering to date". President Greg Brockman called it the company's "most intelligent and, also very importantly, our most aligned model yet".
What shipped
Astra went first to customers of Daybreak, OpenAI's cybersecurity programme, and is rolling out over the following week to Pro, Plus, Enterprise and Business accounts, with API access alongside. The same week OpenAI opened a managed Agents API in public beta - orchestration, long-running sessions and context management for autonomous agents, with sandbox compute from OpenAI itself or from partners such as Vercel and DigitalOcean - and announced DevDay for 29 September in San Francisco.
The claims
OpenAI says Astra beats its own Sol model and Anthropic's Fable on benchmarks for bug-finding, terminal-task execution and codebase questions - the three things a coding agent actually does all day. It is also pitching the model's ability to identify and develop zero-day exploits as a defensive tool, which is why security customers got it first. These are the company's numbers, not independent ones; the comparison to watch is in the editors and agents that will swap it in over the next month.
Every frontier release now leads with the same benchmark: can it find the bug, run the terminal, answer the question about the repo. The model race has become a race to be the best employee in a vibe coder's workflow.
The catch
Astra uses a reasoning technique OpenAI calls "opaque recurrence", which obscures the chain of thought that safety teams use to monitor what a model is doing. Chief scientist Jakub Pachocki acknowledged that "as model capabilities are increasing, monitorability is getting more challenging". Brockman also said there is "no contractual AGI triggering anymore" - AGI is now "a mission concept or spiritual concept" rather than a clause. For a builder, the practical version of the concern is simpler: an agent that is harder to inspect is harder to trust with production access.
What it means for builders
Two moves this week. First, if you build on OpenAI, budget an afternoon to re-run your evals on Astra; the coding gains, if they hold, land directly in agent reliability. Second, the Agents API changes the make-or-buy question for anyone who has been hand-rolling orchestration: sandboxed compute and long-running sessions as a managed service is exactly the layer most one-person companies were building themselves. Cursor's new owner and Anthropic will answer within weeks; the pace is now the product.
Read the newsletter that top founders & VCs keep in their inbox.