In a monumental stride forward for artificial intelligence, OpenAI has officially unveiled its GPT-5.4 'Thinking' model, demonstrating unprecedented capabilities that include achieving human-level performance on complex economic tasks. Released in early March 2026, this latest iteration is poised to redefine professional workflows and significantly augment human decision-making across various industries.
The GPT-5.4 model, available through OpenAI's API and ChatGPT, represents a unification of advanced reasoning, coding, and agentic workflows. It incorporates the sophisticated coding prowess of its predecessor, GPT-5.3-Codex, into a single, comprehensive frontier model.
Unrivaled Performance in Knowledge Work
The most striking aspect of GPT-5.4's capabilities lies in its performance on professional knowledge work. According to OpenAI's internal GDPval evaluation, which assesses the model's aptitude across 44 diverse occupations, ranging from legal analysis to intricate financial modeling, GPT-5.4 either matched or exceeded human professionals in an astounding 83% of comparisons. This marks a substantial improvement from GPT-5.2's 70.9% in similar evaluations. This benchmark indicates that the model can save over four and a half hours on a typical seven-hour professional task, showcasing its immense economic value.
Beyond abstract knowledge tasks, GPT-5.4 also excels in practical computer use. On the OSWorld-V benchmark, designed to simulate real-world desktop productivity tasks, the model achieved a 75% success rate, outperforming the human baseline of 72.4%. This capability allows GPT-5.4 to navigate desktop applications, accurately fill forms, and interact with web browsers with a proficiency surpassing that of expert human testers.
Architectural Advances and Enhanced Reliability
OpenAI has engineered GPT-5.4 with a massive 1-million-token context window, enabling it to process and analyze vast amounts of information—equivalent to hundreds of pages of text—in a single request. This feature is particularly impactful for tasks requiring extensive data comprehension, such as reviewing entire codebases or large document collections.
Furthermore, the model demonstrates enhanced reliability, exhibiting a 33% reduction in factual errors and an 18% decrease in overall mistakes compared to GPT-5.2. This focus on accuracy underscores OpenAI's commitment to developing more dependable and trustworthy AI systems.
Performance Comparison: GDPval Benchmark Scores
The chart below illustrates the significant leap in performance achieved by GPT-5.4 on the GDPval benchmark, showcasing its ability to match or exceed human professionals in a broader range of knowledge work tasks.
The introduction of GPT-5.4 'Thinking' marks a pivotal moment in the advancement of AI. With its enhanced reasoning, comprehensive computer-use capabilities, and remarkable accuracy, OpenAI is not just pushing the boundaries of what AI can achieve but is also laying the groundwork for a future where AI acts as a truly intelligent digital co-worker, streamlining complex tasks and fostering unprecedented levels of productivity across the global economy.