OpenAI’s GPT-6 Astra has sparked a major debate in the field of artificial intelligence: has AI become more intelligent than humans? The company’s new flagship product is a significant stride towards the future of more independent AI, covering computer usage, software development, scientific studies, cybersecurity, and professional workflows.
According to OpenAI, Astra is “world’s most intelligent and aligned model” that can perform complex tasks on computers and browsers with minimal human input. GPT-6 Astra can take multiple steps to achieve an objective, utilize tools, respond to errors, and then keep going until it reaches a goal, unlike traditional AI systems that would only follow a set of instructions.
However, impressive benchmark scores are not a sign of the arrival of artificial general intelligence (AGI). Astra might show super-human abilities in a number of things but not reach the level of general intelligence.
What Is GPT-6 Astra?
GPT-6 Astra is meant to take AI beyond just providing answers. It is a great leap forward in that it can act on behalf of a user. OpenAI says that Astra can complete online forms, update customer records, manage calendars, do online research, draft documents and emails, create websites, analyse scientific data, install, test and troubleshoot software.
It can also navigate computers, websites and software directly, enabling users to hand over multi-step workflows without having to walk the model through each step. This agentic ability may have far reaching consequences for businesses and professionals. Users can increasingly provide an objective and let the AI system figure out how to accomplish a task.
GPT-6 Astra Benchmark Results

The results of Astra’s performance have caught the eye because of the results that have been achieved in several difficult benchmarks.
Astra reportedly achieved a score of 72.6% on OSWorld 2.0, a test that assesses computer-use skills, while GPT-5.6 Sol came in at 65.7%. OpenAI’s data also show that on average, Astra can complete these tasks in about 40 minutes, while Sol takes about 75 minutes.
The model achieved a score of 59.3% in Agents’ Last Exam, a test of complex professional tasks. It scored 64.6% on its Terminal-Bench Science 0.1 score, with the benchmark emphasizing scientific research workflows related to coding, data analysis, simulations and model fitting.
Astra also achieved around 57.9% on Terminal-Bench 4.0, beating Claude Fable 5.1 with 55.8% and GPT-5.6 Sol with 37.3%. These are all indicators of a major change. Astra isn’t just answering questions, it can make decisions on what action to take next, and respond if things go wrong.
What Does the 99.9% ARC-AGI Score Mean?
The most striking GPT-6 Astra result is that it reportedly achieved a score of 99.9% on ARC-AGI-3. The benchmark aims to assess the ability of AI systems to learn rules in new environments without just memorizing the information they’ve seen while being trained. Astra exceeded the human action-efficiency baseline in 96% of the levels of the benchmark, according to ARC Prize Foundation’s Greg Kamradt.
However, the 99.9% figure requires important context.
The supplied benchmark analysis states that Astra was evaluated under two different setups. It got a 99.9% in OpenAI’s provider-adapter setup, where the reasoning state of the model can be maintained across multiple actions. The score under provider neutral harness was 62.7%. The higher result is not irrelevant, but it illustrates the need for taking testing conditions into account when making a benchmark comparison.
Most importantly, ARC-AGI-3 does not constitute AGI itself. The benchmark is in a controlled setting that has rules, while the real world is unpredictable and open-ended.
Has GPT-6 Astra Surpassed Human Intelligence?
The answer is very much context dependent on the definition of “human intelligence.” AI systems are already better than humans at some specific tasks, such as some mathematical problems, games, coding challenges and pattern recognition problems. Astra brings these skills to more interactive computer settings and excels in math, science and cyber security.
Humans are able to transfer knowledge from one situation to another, comprehend social context, set goals, function in an unpredictable physical environment and adapt when rules are ambiguous. If an AI model outperforms humans in a certain metric, it does not necessarily imply that it is more intelligent than humans in general.
So, an AI system can do a specific task better than humans without having general intelligence.
GPT-6 Astra vs Claude and Other AI Models
Astra is very capable, but the evidence in the benchmarks available does not demonstrate its superiority over all other competing AI models. Based on the provided analysis, Artificial Analysis has an overall Intelligence Index of 61, which is the same as GPT-5.6 Sol, while Claude Fable 5.1 has an overall Intelligence Index of 66.
Fable 5.1 has a Coding Agent Index of 70, whereas Astra’s is 67. Astra is superior on some special tests. On Terminal-Bench 4.0 it gets approximately 57.9% while Fable 5.1 gets 55.8%. However, this disparity is significantly smaller on Deep SWE, with Astra at around 74.1% and Claude Opus 5 at 73.7%. The overall situation is thus more complicated than the notion that Astra has just turned into the best AI model at anything.
Where GPT-6 Astra Makes Its Biggest Leap

The most promising development for Astra seems to be the use of agents in computers. It reportedly has a ScreenSpot Pro score of 92.7%, while GPT-5.6 Sol has a score of 76.9%. The analysis provided also shows higher long-context retrieval performance, achieving 96% accuracy on about 1 million tokens compared to 74% for Sol.
OpenAI claims that Astra scored 100% on ExploitBench and 42.4% on ExploitGym. In the tests, the model identified and exploited two previously unknown zero-day vulnerabilities in an exploit chain, the company claims.
These capabilities have led OpenAI to rate Astra as a Critical cyber capability and state that access to its most advanced cyber capabilities will be limited to vetted testers and selected defensive programmes at first.
Why AI Safety Is Becoming More Important
There are two different incidents involving OpenAI-related AI agents: one in July when Hugging Face was involved, and another with a German programming wiki. OpenAI has stressed that Astra itself was not involved in these incidents. The episodes nevertheless point to worries that autonomous systems could act in ways that are harmful when they are given a goal to achieve but not a step-by-step plan.
This is intimately related to the issue of AI alignment. OpenAI claims Astra is its most aligned model, and has added extra safeguards. But the company is also aware that more sophisticated models may be more difficult to understand and control. More intelligence does not necessarily mean more in line with human intentions.
Is GPT-6 Astra the Beginning of AGI?
According to OpenAI’s Charter, AGI is defined as highly autonomous systems that can outperform humans at most economically valuable tasks. If that’s the case, then a high score on one benchmark is not sufficient.
Astra clearly shows an increase in autonomy. It can carry out more complex digital operations, research, write and troubleshoot software, navigate computers and operate through a series of operations. However, there is not enough evidence to say that it does outperform humans in most economically significant tasks today.
Singularity is not to be confused with AGI. The technological singularity is a hypothetical future point at which the rate of technological progress becomes so rapid that humans will no longer be able to reliably predict or control it, whereas AGI is a level of machine capability.
Final Verdict
GPT-6 Astra is a significant leap in AI technology, especially in the realm of computer usage, self-directed workflows, coding, scientific research, and cybersecurity. It has already shown itself to be more efficient than humans at certain specialized tasks. It’s not yet proven that Astra is better than human intelligence, nor that AGI has definitely arrived.
Perhaps the biggest change is that AI systems can now perform more than just answer questions, they can take action. The central issue will be whether humans can anticipate, monitor, and manage the way AI approaches a problem to solve it, as models become more adept at achieving goals over extended series of actions.
Astra might not be the proof of the arrival of AGI. It has shrunk the distance between an AI assistant that follows commands and an AI agent that goes out on its own to achieve goals.






