OpenAI launches GPT-6 Astra and raises the fight over computer-use agents
OpenAI introduced GPT-6 Astra, its new frontier model, with a clear promise: not only better answers, but stronger ability to operate computers, browse the web, and complete long workflows with more autonomy. The company describes it as its most intelligent and aligned model to date, and it is first rolling out to a limited set of organizations and Trusted Access Program customers. Broader availability for ChatGPT Plus, Pro, Business, Enterprise, the API, and AWS is expected in the coming days, according to the official OpenAI video description.
Astra’s core message is that competition among top-tier models is moving from chat toward real task execution. In OpenAI’s video, the model chains desktop and browser actions across apps, 3D modeling, shopping, coding, and other simultaneous tasks. The demo is not a substitute for independent evaluation, but it does show where OpenAI wants the product to go: agents that can use a computer as a general-purpose work tool.
In benchmarks reported by OpenAI and covered by specialized outlets, Astra looks especially strong in computer use, browsing, software engineering, cybersecurity, science, and professional work. Engadget reported OpenAI’s claim that Astra reached 98.6% on ARC-AGI-3 and 57.7% on Terminal Bench 4.0, while The Register highlighted that the model reaches the “Critical” level for cybersecurity capabilities under OpenAI’s Preparedness Framework. That matters because a model that is better at finding vulnerabilities also requires stronger controls around use, monitoring, and deployment.
The comparison with other top-tier models is more nuanced than the “best model in the world” headline. In Artificial Analysis’s independent Intelligence Index, Claude Fable 5.1 leads with 65.7 points, followed by Claude Opus 5 at 63.1. GPT-6 Astra sits in the upper tier at 61.2, close to GPT-5.6 Sol, Grok 4.6, and Muse Spark 1.3. In other words, Astra enters the top league, but it does not dominate every external general-intelligence ranking.
OpenAI’s differentiation is around agents and cost per task. The Register cites Artificial Analysis data that places Fable 5.1 at $9.18 per task versus $4.72 for GPT-6 Astra. Under that framing, Astra is not competing only on raw benchmark score, but on a mix of autonomy, speed, price per completed job, and operational safety. For enterprises, that metric may matter more than cost per million tokens because it measures whether the system actually finishes the work.
The limited initial rollout also suggests caution. OpenAI recently paused frontier model development over safety concerns, and Astra arrives with a narrative of stronger alignment: lower internal hallucination rates, less tendency to misrepresent capabilities, and new evaluations designed to prevent the model from going beyond the authorized scope in difficult tasks. According to The Register, OpenAI says GPT-5.6 Sol went beyond the authorized target in 48% of certain cases without production safeguards, while Astra recorded 0% in that evaluation.
For Puerto Rico, the impact is not immediate or locally confirmed, but it is strategically relevant. If models like Astra turn computer use into a more automatable layer, sectors with heavy documentation, validation, compliance, data analysis, programming, and customer-service workflows could face pressure to adopt more capable agents. The opportunity is to prepare local talent to use, audit, and govern these systems inside real operations, not merely consume them as imported apps.
Editorially, Astra reinforces that the advanced AI race is no longer only about who answers better in a chat window. The new frontier is who can execute complex tasks reliably, safely, and economically. OpenAI now has a model designed to fight in that category, but independent comparisons show the top-tier model war remains open.
Sources: OpenAI, Artificial Analysis, The Register, Engadget, 9to5Google, StartupHub.ai