As AI workloads evolve toward inference and agentic AI, the criteria for system performance are also changing. Computing power alone is no longer enough. Where data is stored, how quickly it moves, and how efficiently it is processed now have a direct impact on the performance and efficiency of AI systems.
This article series explores the transformation of the AI ecosystem, where software, data center infrastructure, semiconductors, and memory technologies work together. It also examines why memory semiconductors have become a foundational technology for the AI era and outlines SK hynix’s vision for building the next-generation AI memory ecosystem.
① The paradigm shift in AI computing – Professor Jinwoo Shin, KAIST
② The real bottleneck: Data, not compute
③ Redesigning infrastructure: Why architecture determines performance
④ The semiconductor paradigm shift: Memory’s growing role
⑤ Completing AI: Physical intelligence and the role of memory
From training to inference, from AI that answers to AI that acts
Since the launch of ChatGPT in late 2022, AI has spread rapidly across industries and everyday life. Yet one of the most significant changes taking place within the AI industry remains largely unnoticed. The long-held approach of advancing AI by training ever-larger models on ever-growing datasets — known as Scaling Laws* — is beginning to reach its limits.

The focus of AI competition is shifting from how much a model has learned to how well it can reason, and ultimately to how effectively it can act on its own. AI is evolving from training-centric systems to reasoning-centric systems, and from tools that simply generate answers to systems capable of autonomous action.
This transition extends far beyond AI model development. It is also redefining the requirements for the semiconductor and memory technologies that underpin AI infrastructure. This article explores the nature of this paradigm shift.
The era of Scaling Laws — and what comes next
ChatGPT’s success was built on three key pillars: the Transformer* architecture, which enabled large-scale neural network training; the pre-training Scaling Laws, which showed that performance improves as models and datasets grow; and post-training*, which aligns models with human intent. The scale of this approach is evident in reports that Meta devoted approximately 30 million GPU hours to training Llama 3, underscoring the enormous computing resources required.
The formula was simple but remarkably effective. It fueled the rapid emergence of frontier models such as ChatGPT, Claude, and Gemini. Since 2024, however, two fundamental limitations have become increasingly apparent.
* Post-training: The stage following pre-training in which a base LLM is refined using carefully curated examples so that it can follow user instructions and produce desired behaviors. This process enables instruction-following and conversational AI models.

The first challenge is data scarcity. Ilya Sutskever, one of the pioneers behind ChatGPT, argued that “pre-training as we know it will end” during his acceptance speech for the Test-of-Time Award* at NeurIPS 2024, one of the world’s leading conferences on AI and machine learning. While computing resources can continue to grow through better hardware, improved algorithms, and larger computing clusters, the supply of high-quality training data cannot expand in the same way. Sutskever compared internet text to “the fossil fuel of AI.” Researchers have similarly projected that by around 2028, the total amount of text that can be extracted from the internet will roughly match the amount already used to train AI models.
The second challenge is cost efficiency. Simply doubling a model’s size does not double its performance. Instead, each additional increase in model size, dataset size, or computing power produces a smaller improvement than the one before. If greater gains can be achieved with the same investment through other approaches, there is less reason to devote ever-increasing resources solely to large-scale pre-training.
Together, these constraints are reshaping the direction of AI development. Rather than continuing to scale models and datasets indefinitely, the industry is increasingly investing more computing resources in the reasoning phase — when a model generates an answer.
This does not mean AI progress is slowing. It means the path to progress is changing.
In the past, the most reliable way to improve performance was to train larger models for longer periods. Today, the emphasis is shifting toward how effectively a model can apply and maximize the knowledge it has already learned during inference.
Even the same model can produce very different results depending on how a problem is decomposed, how carefully intermediate reasoning steps are carried out and how effectively external tools are used when needed. In other words, the center of AI competition is moving from the training phase to the inference phase.
The rise of reasoning: The third Scaling Law
Across several keynote presentations, NVIDIA Founder and CEO Jensen Huang has used the phrase “from one to three” to describe a fundamental shift in AI’s development process. AI performance was once explained primarily through a single Scaling Law centered on pre-training, but today, it is increasingly driven by three complementary forms of scaling: pre-training, post-training, and test-time (Inference-time) scaling.
At the center of this shift is reasoning during inference. Just as people do not immediately blurt out the answer to a difficult math problem but instead work through it step by step, AI models can achieve significantly better results by performing a longer and more sophisticated reasoning process before generating a response. One of the most influential techniques enabling this capability is Chain of Thought*.
Reasoning-first models such as OpenAI’s o1 and o3 series and DeepSeek-R1 exemplify this trend. Among them, o1 clearly demonstrates how inference-time computation can dramatically improve performance. On AIME* 2024, a benchmark based on competition-level mathematics, the general-purpose multimodal model GPT-4o scored 13.4%, while o1 achieved 83.3%. On GPQA Diamond*, a benchmark consisting of Ph.D.-level scientific questions, o1 scored 78%, surpassing the average human expert accuracy of 69.7%. These results suggest that allocating more computation to the reasoning process itself can substantially improve model performance.
Another milestone came in September 2025, when a paper on DeepSeek-R1 from Chinese startup DeepSeek was featured on the cover of “Nature.” In the paper, the researchers challenged a widely held assumption about how advanced reasoning models are trained. Previously, advanced reasoning capabilities were widely believed to require large volumes of expensive supervised fine-tuning (SFT) data. The DeepSeek research team, however, demonstrated that reinforcement learning alone could enable a model to develop its own chain-of-thought reasoning process. The work suggested that sophisticated reasoning capabilities are no longer exclusive to a handful of major technology companies.

▲ Image Source: Cover of Nature, Vol. 645, Issue 8081 (18 September 2025).
This shift is also changing how AI systems perform computation.
Conventional LLMs are primarily optimized to generate relatively short responses as quickly as possible. By contrast, reasoning models may perform internal reasoning equivalent to hundreds or even thousands of tokens before producing a final answer. Measurements show that for a simple arithmetic problem that a conventional model typically answers using 7-12 output tokens, reasoning models such as o1 and DeepSeek-R1 can consume more than 900 tokens before arriving at the same answer. In other words, the number of output tokens processed by the same GPU can increase by nearly 100 times. Inference is no longer simply the execution stage — it is becoming one of the primary computational workloads in AI services.
The economics of AI services are changing as well.
Pre-training requires a massive upfront investment in compute. Inference costs, however, are incurred every time a user submits a query. As AI systems spend more time reasoning and execute more intermediate steps before generating an answer, the computational demand and infrastructure burden of operating AI services continue to grow.
The rise of reasoning-centric AI therefore represents far more than the arrival of smarter models. It is transforming the economics of AI services, response latency, user experience, and the way AI infrastructure itself is designed.
* American Invitational Mathematics Examination (AIME): A highly challenging mathematics competition administered by the Mathematical Association of America (MAA). It is widely used as a benchmark for evaluating complex mathematical reasoning in AI models.
* Graduate-Level Google-Proof Q&A Diamond (GPQA Diamond): A benchmark consisting of Ph.D.-level questions in biology, chemistry, and physics, used to evaluate scientific reasoning capabilities in AI models.
Agentic AI: From answering questions to taking action
As AI reasoning capabilities improve, the way people use AI is changing as well. At the center of this shift is agentic AI. Unlike a conversational model that simply responds to a user’s prompt, an agentic AI system can receive a goal, develop its own plan, invoke the tools it needs, evaluate the results, and complete a task autonomously.
Consider a simple request: “Book a Korean restaurant near Times Square for two people tonight.” Rather than generating a single response, an agentic AI breaks the request into multiple subtasks — searching for restaurants, comparing options, checking business hours, finding available reservation times, incorporating the user’s preferences, making the reservation, and adding it to a calendar. At each step, it can call external tools such as search engines, maps, reservation application programming interfaces (APIs), and calendar services as needed. Tasks that once required users to search, compare, and enter information manually can instead be carried out by AI.

This highlights the difference between chatbots and agentic AI. Chatbots are designed primarily to respond to user queries. Agentic AI, by contrast, is designed to accomplish goals. Rather than requiring users to specify every step, it decomposes a task, determines the sequence of actions, gathers the necessary information, evaluates intermediate results, and decides what to do next. In the process, AI expands beyond a standalone language model to work with search engines, code execution environments, databases, enterprise software, and even robotic control systems. This is why agentic AI is gaining traction: AI is evolving from a conversational interface into an active participant in digital work.
The market has already begun to recognize this trend. In its 2024 Hype Cycle* for Artificial Intelligence, Gartner classified AI Agents at the Peak of Inflated Expectations.* The same report also identified World Models* and Embodied AI* as key technologies in the early stages of innovation. Together, these developments suggest that AI capable of taking action — not simply generating responses — could reshape industries over the next 5-10 years.
* Peak of Inflated Expectations: A stage in which interest and expectations surrounding a new technology rise rapidly. While early success stories receive significant attention, the technology’s actual maturity and commercial impact often fall short of market expectations.
* World Models: AI systems designed to understand how the physical world works, including physical laws and spatial relationships. By learning from text, images, video, audio and motion, they can predict future states and infer what is likely to happen next.
* Embodied AI: AI integrated into physical systems that can interact with the real world. Applications include general-purpose robots, humanoid robots, autonomous vehicles, factories and logistics centers. By combining machine learning with sensors and computer vision, embodied AI can perceive its surroundings, make decisions and take actions autonomously.
Agentic AI also changes the nature of computing demand. A single user request may trigger dozens of LLM calls and dozens of tool invocations behind the scenes. While reasoning models increase the amount of computation devoted to thinking through a problem, agentic AI dramatically increases the number of interactions required to complete a task. As reasoning and action converge, the total number of tokens that data centers must process grows substantially.

The past decade of AI development can be summarized in a single phrase: build bigger models. The next phase, however, will be far more complex. The center of competition is shifting from pre-training to post-training and now to Inference-time computation. At the same time, AI is evolving from a system that provides one-off responses into an autonomous agent capable of carrying out tasks. The standards for evaluating AI are changing as well. Success is no longer measured only by how often a model produces the correct answer, but also by how consistently it reasons, how safely it acts, and how reliably it delivers trustworthy results.
How will AI infrastructure keep pace with this evolution?
Reasoning models and agentic AI are fundamentally changing both the volume and the nature of AI computation. Responses are becoming longer, model and tool invocations are increasing, and context must be maintained over much longer interactions. As a result, simply adding more compute is no longer enough to improve AI system performance.
Part 2 of the series will examine this challenge from a system-level perspective by answering a fundamental question: What is the true bottleneck in AI performance?

Disclaimer: The opinions expressed in this article are solely those of the author and do not necessarily reflect the official position of SK hynix.