{"id":13417,"date":"2026-10-02T08:59:03","date_gmt":"2026-10-01T23:59:03","guid":{"rendered":"https:\/\/news.skhynix.com\/en\/?p=13417"},"modified":"2026-10-02T08:59:03","modified_gmt":"2026-10-01T23:59:03","slug":"ai-ecosystem-series-ep3","status":"publish","type":"post","link":"https:\/\/news.skhynix.com\/en\/ai-ecosystem-series-ep3\/","title":{"rendered":"[AI Ecosystem] Redesigning infrastructure: Why architecture determines performance"},"content":{"rendered":"<div class=\"post-intro\">\n<p>As AI workloads evolve toward inference and agentic AI, the criteria for system performance are also changing. Computing power alone is no longer enough. Where data is stored, how quickly it moves, and how efficiently it is processed now have a direct impact on the performance and efficiency of AI systems.<\/p>\n<div style=\"height: 16px; line-height: 4px;\"><\/div>\n<p>This article series explores the transformation of the AI ecosystem, where software, data center infrastructure, semiconductors, and memory technologies work together. It also examines why memory semiconductors have become a foundational technology for the AI era and outlines SK hynix\u2019s vision for building the next-generation AI memory ecosystem.<\/p>\n<div style=\"height: 16px; line-height: 4px;\"><\/div>\n<p>\u2460 The paradigm shift in AI computing<br \/>\n\u2461 The real bottleneck: Data, not compute<br \/>\n<strong>\u2462 Redesigning infrastructure: Why architecture determines performance<\/strong><br \/>\n<strong>\u2013 Professor Onur Mutlu, ETH Zurich<\/strong><br \/>\n\u2463 The semiconductor paradigm shifts toward memory<br \/>\n\u2464 Completing AI: Physical intelligence and the role of memory<\/p>\n<\/div>\n<div style=\"height: 16px; line-height: 16px;\"><\/div>\n<div style=\"text-align: center;\">\n<div style=\"display: inline-block; max-width: 748px; width: 100%; text-align: left;\">\n<div style=\"height: 2px; background: #666666; margin-bottom: 6px;\"><\/div>\n<h3 class=\"sub-title\" style=\"margin: 0; line-height: 1.4;\">From faster components to systems that move less data<\/h3>\n<div style=\"height: 2px; background: #666666; margin-top: 6px;\"><\/div>\n<\/div>\n<\/div>\n<p>For a large language model(LLM) service such as ChatGPT, Gemini, or Claude to answer a single question, the model must repeatedly access vast amounts of model weights, previous context, and intermediate results. As responses become longer and involve more reasoning steps, the amount of data that must be read also increases. Therefore, one of the key challenges for AI systems has been how quickly and reliably they can supply massive volumes of data to compute devices.<\/p>\n<p>This challenge has become increasingly important as the center of gravity in AI shifts from model development to large-scale service operations. Pre-training requires a massive upfront investment, while inference occurs repeatedly every time a user submits a query. As generative AI services, enterprise copilots, and agentic AI proliferate, the number of requests grows, and the context that each request must maintain becomes longer. As a result, AI infrastructure competition is moving beyond simply deploying faster GPUs\/NPUs\/TPUs and larger clusters. The focus is increasingly on how efficiently GPUs\/NPUs\/TPUs, memory, storage, and networks can be co-designed and connected as an integrated system and how efficiently data can flow through that system. Or more appropriately, how minimally the system moves data across components.<\/p>\n<div style=\"height: 16px; line-height: 16px;\"><\/div>\n<div style=\"text-align: center;\">\n<div style=\"display: inline-block; max-width: 748px; width: 100%; text-align: left;\">\n<div style=\"height: 2px; background: #666666; margin-bottom: 6px;\"><\/div>\n<h3 class=\"sub-title\" style=\"margin: 0; line-height: 1.4;\">The cost of data movement in processor-centric architectures<\/h3>\n<div style=\"height: 2px; background: #666666; margin-top: 6px;\"><\/div>\n<\/div>\n<\/div>\n<p>Today\u2019s computing systems provide a very powerful infrastructure that has enabled increasingly important workloads such as AI and machine learning, genomic analysis, and large-scale data analytics. Yet their fundamental architecture remains processor-centric. Data must continuously move across the memory hierarchy \u2014 from sensors, storage, memory, and caches all the way to the compute device(s) and back again \u2014 before and after being processed.<\/p>\n<p>This architecture has long been effective at improving the performance of general-purpose computations. As AI workloads have grown dramatically in both data volume and access frequency, however, data access and data movement have become major bottlenecks in system efficiency. Models continuously access vast amounts of weights and intermediate results, while inference repeatedly reads previous context and newly generated tokens. The farther and more data must travel, the greater the latency, power consumption, and hardware cost.<\/p>\n<p>To mitigate these data access and data movement bottlenecks, system designers have complicated the design of computers greatly. Modern computing systems employ many complex and expensive mechanisms, such as cache hierarchies<span style=\"color: #ff0000;\">*<\/span>, prefetching data before it is needed, multithreading to execute multiple tasks in parallel, and out-of-order execution<span style=\"color: #ff0000;\">*<\/span>. These technologies have had limited effectiveness at reducing the latency of compute devices, but they have come at significant complexity, energy inefficiency, and hardware cost in the system.<\/p>\n<div class=\"footnote\"><span style=\"color: red;\">* <\/span>Cache hierarchies: A hierarchical structure that temporarily stores recently or frequently used data close to compute devices to reduce memory access time. It typically consists of multiple levels, such as L1, L2, and L3 caches<br \/>\n<span style=\"color: red;\">* <\/span>Out-of-order execution: A technique that reduces compute device idle time by executing instructions that are ready to execute ahead of others, even when this differs from the order specified in the program<\/div>\n<p>In today\u2019s computing systems, computation is relatively cheap, yet memory access and data movement are expensive \u2014 in both performance and energy. The energy required to retrieve data from DRAM once can be 150 to 2,000 times greater than that required for a simple arithmetic operation. Research involving neural network models running on Tensor Processing Unit(TPU) systems has demonstrated the huge energy and performance impact of this disparity<span style=\"color: #ff0000;\">*<\/span>. For large machine learning models, more than 90% of system energy can be consumed by memory access and data movement rather than computation<span style=\"color: #ff0000;\">*<\/span>.<\/p>\n<div class=\"footnote\"><span style=\"color: red;\">* <\/span>Amirali Boroumand, et al., \u201c<em><a href=\"https:\/\/arxiv.org\/abs\/2109.14320\" target=\"_blank\" rel=\"noopener\">Google Neural Network Models for Edge Devices: Analyzing and Mitigating Machine Learning Inference Bottlenecks<\/a><\/em>,\u201d 30th International Conference on Parallel Architectures and Compilation Techniques(PACT) (2021)<br \/>\n<span style=\"color: red;\">* <\/span>Onur Mutlu et al., \u201c<a href=\"https:\/\/arxiv.org\/pdf\/2012.03112\" target=\"_blank\" rel=\"noopener\"><em>A Modern Primer on Processing in Memory<\/em><\/a>,\u201d in Emerging Computing: From Devices to Systems \u2013 Looking Beyond Moore and Von Neumann, Springer, 2022; updated 2025<\/div>\n<p>The compute-memory disparity becomes even more pronounced with large language models(LLMs). Each time an LLM generates a token, it refers to the previous context and handles vast amounts of intermediate data through the KV cache<span style=\"color: #ff0000;\">*<\/span> and attention operations. As models reason for longer periods and maintain longer contexts, their requirements for memory capacity and bandwidth<span style=\"color: #ff0000;\">*<\/span> drastically increase. The memory bottleneck becomes larger and larger as these memory demands increase. No matter how fast the compute device becomes, system efficiency declines because the system spends most of its time waiting for data to arrive at the compute units.<\/p>\n<p>As AI workloads scale, High Bandwidth Memory becomes increasingly important for supplying GPUs with larger volumes of data at higher speeds. Positioned next to the GPU\/NPU\/TPU, HBM delivers large amounts of data at ultra-high speeds, helping reduce compute device latency. To fully utilize HBM performance, however, the entire data path must be designed as an integrated system \u2014 not only the GPU\/NPU\/TPU and HBM, but also storage, networks, and interconnects<span style=\"color: #ff0000;\">*<\/span>.<\/p>\n<div class=\"footnote\"><span style=\"color: red;\">* <\/span>Key-value cache(KV cache): A memory area where an LLM stores attention computation results from previous tokens for use when generating subsequent tokens. As context length increases, the KV cache grows, increasing memory capacity and bandwidth requirements<br \/>\n<span style=\"color: red;\">* <\/span>Memory Bandwidth: The amount of data that can be transferred from memory within a given period of time. Higher memory bandwidth allows compute devices to receive and process more data at once<br \/>\n<span style=\"color: red;\">* <\/span>Interconnect: A communication pathway connecting system components such as processors, memory, accelerators, and storage. In AI infrastructure, interconnects are a key factor in determining the speed and efficiency of data movement<\/div>\n<div style=\"height: 16px; line-height: 16px;\"><\/div>\n<div style=\"text-align: center;\">\n<div style=\"display: inline-block; max-width: 748px; width: 100%; text-align: left;\">\n<div style=\"height: 2px; background: #666666; margin-bottom: 6px;\"><\/div>\n<h3 class=\"sub-title\" style=\"margin: 0; line-height: 1.4;\">Why system architecture matters more than component speed<\/h3>\n<div style=\"height: 2px; background: #666666; margin-top: 6px;\"><\/div>\n<\/div>\n<\/div>\n<p>The fundamental assumptions behind system design also need to change. Processor-centric architectures largely operate by moving data to compute devices for processing. As data volumes grow and access frequency increases, architectures that send all data to a central compute device quickly reach their limits. Now, the system must instead be redesigned around where data resides and the path it takes to be processed.<\/p>\n<p>One concept developed to address this challenge is memory-centric computing<span style=\"color: #ff0000;\">*<\/span>. Solving the memory bottleneck in a fundamentally better way requires reducing data movement and enabling computation closer to memory. This is the basic idea behind memory-centric computing. Memory-centric computing takes data location and movement costs as key considerations in system design, optimizing memory placement and compute architecture together<span style=\"color: #ff0000;\">*<\/span>. Some data may reside in HBM next to the GPU\/NPU\/TPU, while other data is managed in a shared memory pool<span style=\"color: #ff0000;\">*<\/span>, with certain computations performed closer to memory. The best approach depends on factors such as model architecture, context length, the number of user requests, and network configuration.<\/p>\n<p>The good news is that such memory-centric systems are already being designed and evaluated not only in academia but also in the industry. There is an increasing push toward eliminating data movement bottlenecks, and this is best done by not moving data as much as possible.<\/p>\n<div class=\"footnote\"><span style=\"color: red;\">* <\/span>Memory-centric computing: A computing paradigm designed to process data closer to memory, moving away from the existing structure that moves data to\/from the compute device. It focuses on reducing data movement to improve both performance and power efficiency<br \/>\n<span style=\"color: red;\">* <\/span>Onur Mutlu, et al., \u201c<em><a href=\"https:\/\/arxiv.org\/abs\/2505.00458\" target=\"_blank\" rel=\"noopener\">Memory-Centric Computing: Solving Computing\u2019s Memory Problem<\/a><\/em>,\u201d 17th IEEE International Memory Workshop(IMW) (2025)<br \/>\n<span style=\"color: red;\">* <\/span>Memory pool: A shared memory resource that can be accessed by multiple devices. By allowing each device to use memory capacity as needed, memory pooling can improve overall system resource utilization<\/div>\n<div style=\"height: 16px; line-height: 16px;\"><\/div>\n<div style=\"text-align: center;\">\n<div style=\"display: inline-block; max-width: 748px; width: 100%; text-align: left;\">\n<div style=\"height: 2px; background: #666666; margin-bottom: 6px;\"><\/div>\n<h3 class=\"sub-title\" style=\"margin: 0; line-height: 1.4;\">High-speed interconnects: The two pillars of memory expansion and GPU connectivity<\/h3>\n<div style=\"height: 2px; background: #666666; margin-top: 6px;\"><\/div>\n<\/div>\n<\/div>\n<p>Processors, memory, and accelerators have long been connected through a variety of interfaces. As the volume of data processed by AI workloads grows rapidly, however, conventional connectivity architectures are struggling to provide the required communication bandwidth and memory access efficiency. Consequently, the importance of high-speed interconnects that enable faster and more flexible data paths between devices is increasing. Today, CXL<span style=\"color: #ff0000;\">*<\/span> and NVLink\/NVSwitch<span style=\"color: #ff0000;\">*<\/span> are two representative technologies driving this shift in different areas and various academic works to tightly integrate components exist.<\/p>\n<div class=\"footnote\"><span style=\"color: red;\">* <\/span>Compute Express Link(CXL): A next-generation interconnect standard that provides high-speed connectivity among CPUs, GPUs\/NPUs\/TPUs, memory, and accelerators. CXL is a key technology for enabling memory expansion and sharing<br \/>\n<span style=\"color: red;\">* <\/span>NVLink and NVSwitch: High-speed interconnect technologies developed by NVIDIA. NVLink provides high-bandwidth connections between GPUs or between GPUs and other devices, while NVSwitch extends this connectivity to enable high-speed communication among larger numbers of GPUs. These technologies are used to reduce GPU-to-GPU data bottlenecks in large-scale AI clusters<\/div>\n<p>CXL is a key interconnect technology that enables memory expansion and sharing. In traditional server architectures, memory has largely functioned as a dedicated resource attached to a specific CPU or GPU\/NPU\/TPU. Even when one device had spare memory capacity while another faced a shortage, it was difficult to share that capacity flexibly. CXL expands connectivity among processors, accelerators, and memory devices, providing a foundation for flexible memory expansion, sharing, and pooling. With CXL, computation can also be offloaded to devices or controllers that perform processing near memory, while CXL devices can communicate with one another through the flexible interface.<\/p>\n<p>NVLink and NVSwitch provide high-bandwidth connectivity within GPU clusters. When large-scale AI models run across multiple GPUs and nodes, intermediate data such as model parameters, activations, and the KV cache must move between GPUs. If GPU-to-GPU connectivity is slow or inflexible, overall processing performance can suffer regardless of how fast each individual GPU is.<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-13424\" src=\"https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/28155459\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_03_graphic_2026.jpg\" alt=\"\" width=\"1600\" height=\"1180\" srcset=\"https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/28155459\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_03_graphic_2026.jpg 1600w, https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/28155459\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_03_graphic_2026-300x221.jpg 300w, https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/28155459\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_03_graphic_2026-1024x755.jpg 1024w, https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/28155459\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_03_graphic_2026-768x566.jpg 768w, https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/28155459\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_03_graphic_2026-1536x1133.jpg 1536w, https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/28155459\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_03_graphic_2026-1200x885.jpg 1200w\" sizes=\"auto, (max-width: 1600px) 100vw, 1600px\" \/><\/p>\n<p>CXL and NVLink\/NVSwitch serve different roles: CXL focuses on memory expansion and sharing, while NVLink\/NVSwitch provide high-bandwidth GPU connectivity. Both, however, support the same broader goal of reducing data bottlenecks in AI infrastructure by enabling the delivery of the required data quickly to where it is needed. When such a connection architecture is established, memory and accelerators cease to become fixed components and become system resources that can be configured around workload requirements.<\/p>\n<p>The flexible connectivity enabled by high-speed interconnects is also changing how compute devices and memory are arranged. As more data paths become available between devices, compute and memory resources can be positioned according to workload requirements. This shift points toward an architecture that does not simply move more data over longer distances but instead performs the required processing closer to where the data resides.<\/p>\n<div style=\"height: 16px; line-height: 16px;\"><\/div>\n<div style=\"text-align: center;\">\n<div style=\"display: inline-block; max-width: 748px; width: 100%; text-align: left;\">\n<div style=\"height: 2px; background: #666666; margin-bottom: 6px;\"><\/div>\n<h3 class=\"sub-title\" style=\"margin: 0; line-height: 1.4;\">Tightly connected near-memory accelerators: Processing data closer to memory through cooperative computing<\/h3>\n<div style=\"height: 2px; background: #666666; margin-top: 6px;\"><\/div>\n<\/div>\n<\/div>\n<p>Reducing data movement bottlenecks can go a step beyond simply increasing connection speeds. While CXL and NVLink\/NVSwitch improve data paths between memory and GPUs\/NPUs\/TPUs in different ways, near-memory acceleration reduces the amount of data that must be moved by placing computation itself closer to the data.<\/p>\n<p>Instead of having a central GPU\/NPU\/TPU retrieve and process all data, accelerators located in or near memory devices divide the workload among themselves. These accelerators are tightly coupled through high-speed interconnects, allowing data to be processed close to memory first, with only the necessary results exchanged with other devices, rather than sending the entire dataset to a central GPU\/NPU\/TPU. When computation takes place near 3D-stacked high-bandwidth memory<span style=\"color: #ff0000;\">*<\/span> such as HBM or other high-capacity memory, much less data needs to travel back and forth. This can reduce latency and power consumption while making it easier to scale across multiple nodes. This architecture is known as a distributed system of tightly connected near-memory accelerators<span style=\"color: #ff0000;\">*<\/span>.<\/p>\n<p>The potential of this approach can be seen in Tesseract<span style=\"color: #ff0000;\">*<\/span>. Tesseract is a research architecture(published at the 42nd ACM\/IEEE International Symposium on Computer Architecture in 2015) that applies distributed near-memory acceleration to parallel graph processing. Multiple accelerators located close to memory process their assigned data independently and communicate with one another only when necessary. For various parallel graph-processing workloads, the research demonstrated an order-of-magnitude performance improvement along with a similar energy reduction compared with conventional processor-centric designs.<\/p>\n<p>The significance of these results extends beyond the numbers themselves. By processing each node\u2019s assigned data locally and reducing communication volume, the architecture can achieve greater scalability and energy efficiency than conventional designs. As more and diverse accelerators are added, memory capacity, memory bandwidth<span style=\"color: #ff0000;\">*<\/span>, and computation capability all scale proportionally. In other words, the key is not simply distributed processing, but an architecture in which multiple near-memory accelerators are tightly connected and perform computation cooperatively with minimal data movement.<\/p>\n<div class=\"footnote\"><span style=\"color: red;\">* <\/span>3D-stacked high-bandwidth memory: A memory architecture that vertically stacks multiple memory dies to provide high bandwidth. HBM is a representative example<br \/>\n<span style=\"color: red;\">* <\/span>Distributed system of tightly connected near-memory accelerators: An architecture that connects multiple accelerators located close to memory via high-speed interconnects, enabling each accelerator to process its assigned data and exchange only the necessary results<br \/>\n<span style=\"color: red;\">* <\/span>Tesseract: A research architecture for near-memory acceleration in which multiple accelerators located close to memory are tightly connected through high-speed interconnects. Each accelerator processes its assigned data locally and communicates with others only when necessary<br \/>\n<span style=\"color: red;\">* <\/span>The original architecture is published at the 42nd ACM\/IEEE International Symposium on Computer Architecture in 2015, with a paper entitled \u201c<em><a href=\"https:\/\/ieeexplore.ieee.org\/document\/7284059\" target=\"_blank\" rel=\"noopener\">A Scalable Processing-in-Memory Accelerator for Parallel Graph Processing<\/a><\/em>\u201d.<\/div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-13431\" src=\"https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/28155621\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_03_motion_2026.gif\" alt=\"\" width=\"1600\" height=\"1050\" \/><\/p>\n<p>More recently, research has also targeted LLM inference. For example, a CXL-based LLM inference study presented at ASPLOS<span style=\"color: #ff0000;\">*<\/span> 2025 demonstrated how memory expansion architectures and processing units located close to memory can help to greatly alleviate capacity and bandwidth constraints in large-scale inference infrastructure(called the CENT design)<span style=\"color: #ff0000;\">*<\/span>. As memory capacity and bandwidth requirements increase, architectures that concentrate all computations on a central GPU\/NPU\/TPU alone face growing challenges in improving efficiency. Such distributed near-memory architectures like the CENT design that minimize data movement using near-memory processing and efficient &amp; scalable interconnects in intelligent and novel ways can improve all metrics(performance, energy, cost, hardware area) at the same time.<\/p>\n<div class=\"footnote\"><span style=\"color: red;\">* <\/span>ASPLOS: Short for Architectural Support for Programming Languages and Operating Systems, a major international conference covering computer architecture, programming languages, and operating systems<br \/>\n<span style=\"color: red;\">* <\/span>Yufeng Gu, Alireza Khadem, Sumanth Umesh, Ning Liang, Xavier Servot, Onur Mutlu, Ravi Iyer, and Reetuparna Das,\u00a0\u201c<em><a href=\"https:\/\/arxiv.org\/abs\/2502.07578\" target=\"_blank\" rel=\"noopener\">PIM Is All You Need: A CXL-Enabled GPU-Free System for Large Language Model Inference<\/a><\/em>,\u201d ASPLOS 2025<\/div>\n<div style=\"height: 16px; line-height: 16px;\"><\/div>\n<div style=\"text-align: center;\">\n<div style=\"display: inline-block; max-width: 748px; width: 100%; text-align: left;\">\n<div style=\"height: 2px; background: #666666; margin-bottom: 6px;\"><\/div>\n<h3 class=\"sub-title\" style=\"margin: 0; line-height: 1.4;\">Disaggregated architecture: Resource pooling and workload-optimized allocation<\/h3>\n<div style=\"height: 2px; background: #666666; margin-top: 6px;\"><\/div>\n<\/div>\n<\/div>\n<p>Disaggregated architecture<span style=\"color: #ff0000;\">*<\/span> starts from the same fundamental challenge but takes a somewhat different approach. Unlike near-memory accelerator architectures, which move computation closer to data, disaggregated architectures organize CPU, GPU\/NPU\/TPU, memory, storage, and network resources into a system-wide shared pool. In this structure, the necessary resources can then be flexibly combined according to workload requirements. Workloads that require large memory capacity can use an expanded memory pool, while workloads whose performance depends on bandwidth can be handled by near-memory accelerators. In essence, resources are deployed according to the workload\u2019s requirements rather than being fitted into a fixed server configuration. This is very much in-line with the \u201cAsymmetry Everywhere\u201d vision outlined in 2010 to enable customized, flexible, and reconfigurable processing to maximize energy efficiency and performance<span style=\"color: #ff0000;\">*<\/span>.<\/p>\n<div class=\"footnote\"><span style=\"color: red;\">* <\/span>Disaggregated architecture: An architecture that separates resources such as CPUs, GPUs, memory, and storage at the system level rather than fixing them within individual servers, allowing resources to be combined as needed<br \/>\n<span style=\"color: red;\">* <\/span>Onur Mutlu, \u201c<em><a href=\"https:\/\/people.inf.ethz.ch\/omutlu\/pub\/mutlu_asymmetry-everywhere-talk_nsf-acar10.pdf\" target=\"_blank\" rel=\"noopener\">Asymmetry Everywhere(with Automatic Resource Management<\/a>)<\/em>,\u201d CRA Workshop on Advancing Computer Architecture Research, 2010<\/div>\n<div style=\"height: 16px; line-height: 16px;\"><\/div>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-13468\" src=\"https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/30110747\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_03_01_graphic_2026.jpg\" alt=\"\" width=\"3334\" height=\"2293\" srcset=\"https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/30110747\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_03_01_graphic_2026.jpg 3334w, https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/30110747\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_03_01_graphic_2026-300x206.jpg 300w, https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/30110747\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_03_01_graphic_2026-1024x704.jpg 1024w, https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/30110747\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_03_01_graphic_2026-768x528.jpg 768w, https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/30110747\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_03_01_graphic_2026-1536x1056.jpg 1536w, https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/30110747\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_03_01_graphic_2026-2048x1409.jpg 2048w, https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/30110747\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_03_01_graphic_2026-1200x825.jpg 1200w, https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/30110747\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_03_01_graphic_2026-1980x1362.jpg 1980w\" sizes=\"auto, (max-width: 3334px) 100vw, 3334px\" \/><\/p>\n<p>Data that requires substantial memory capacity, such as the KV cache, and memory-intensive operations, such as the attention layer<span style=\"color: #ff0000;\">*<\/span>, are representative use cases for this type of architecture. With high-speed, flexible interconnects and flexible, compute-capable memory configurations, these workloads can be assigned to separate memory resources or near-memory acceleration domains. This reduces the burden on GPUs\/NPUs\/TPUs to retrieve and process all data directly, while allowing devices to locally and efficiently perform computations and to exchange results more efficiently.<\/p>\n<p>The effectiveness of a disaggregated architecture depends on balancing resource allocation with communication costs. How to minimize communication and how to perform efficient data flow within and across accelerators continue to be important questions because excessive communication can actually increase latency and energy consumption. It is therefore critical to determine which resources and workloads should be disaggregated, which should remain close to memory, and where exactly what computation should be performed.<\/p>\n<div class=\"footnote\" style=\"text-align: left; word-break: keep-all; overflow-wrap: break-word;\"><span style=\"color: red;\">* <\/span>Attention layer: A layer in an AI model that determines which information within a given context should receive greater attention. It plays a key role in the Transformer architecture underlying LLMs, with memory usage and data access requirements increasing as context length grows.<\/div>\n<div style=\"height: 16px; line-height: 16px;\"><\/div>\n<div style=\"text-align: center;\">\n<div style=\"display: inline-block; max-width: 748px; width: 100%; text-align: left;\">\n<div style=\"height: 2px; background: #666666; margin-bottom: 6px;\"><\/div>\n<h3 class=\"sub-title\" style=\"margin: 0; line-height: 1.4;\">Reconfigurable AI infrastructure<\/h3>\n<div style=\"height: 2px; background: #666666; margin-top: 6px;\"><\/div>\n<\/div>\n<\/div>\n<p>The competitiveness of AI infrastructure depends on how fast and efficient a path it can provide to the data required for computation. High-speed interconnects such as CXL and NVLink, near-memory accelerator architectures such as Tesseract, and disaggregated architectures that flexibly and efficiently combine computation and memory resources according to workload requirements, including CXL-based memory pooling and customized near-memory processing, are all moving in this direction. This is also why major Big Tech and cloud companies continue to invest heavily in AI infrastructure architecture. The economics of AI competition now extend beyond acquiring more compute devices to how efficiently one can design AI systems across the stack with memory as a first-class citizen: enabling efficient use of memory, power, networks, silicon, and minimizing data movement to greatly improve inference performance and efficiency.<\/p>\n<p>This shift is transforming AI infrastructure from static infrastructure into a dynamic system. As models grow and evolve, contexts lengthen, and inference calls increase, infrastructure must become more flexible and efficient. It must be able to select the appropriate processing path for each workload and model.<\/p>\n<p>As system architectures rapidly evolve, so does the role of memory. Traditionally, memory has primarily been treated as a device for storing and supplying data. With AI workloads, however, the distance between where data resides and where computation takes place directly affects performance and power efficiency. Accordingly, memory is quickly evolving to become an active first-class citizen to perform computation and improve overall system efficiency.<\/p>\n<p>Recent research results are particularly fascinating, showing that some computation is already possible within existing DRAM chips themselves<span style=\"color: #ff0000;\">*<\/span>. By changing how the memory controller<span style=\"color: #ff0000;\">*<\/span> accesses DRAM and activating multiple rows simultaneously, systems can perform large-scale(TeraOps per second<span style=\"color: #ff0000;\">*<\/span> level) bit-level operations or data copying &amp; initialization purely inside modern DRAM chips. These capabilities arise from the fundamental operational principles of DRAM circuitry itself. We call this approach processing using DRAM, since one \u201cuses\u201d (existing) DRAM chips to perform computation. This recent discovery on the extensive computational capabilities of real DRAM chips offers only a glimpse, yet an important one, of how memory could take on a more active role in future computing architectures and provides a foundation for building new processing-in-DRAM mechanisms into future DRAM chips and standards.<\/p>\n<p>AI performance cannot be fully explained simply by adding up the specifications of individual components. Future competitiveness will increasingly depend on how precisely systems are designed and how computation is performed around where data resides, which devices perform computation, and which paths are used to exchange results.<\/p>\n<div class=\"footnote\"><span style=\"color: red;\">* <\/span>Ismail Emir Yuksel, et al., \u201c<em><a href=\"https:\/\/arxiv.org\/abs\/2402.18736\" target=\"_blank\" rel=\"noopener\">Functionally-Complete Boolean Logic in Real DRAM Chips: Experimental Characterization and Analysis<\/a><\/em>,\u201d 30th International Symposium on High-Performance Computer Architecture(HPCA) (2024)<br \/>\n<span style=\"color: red;\">* <\/span>Memory controller: A device that controls data read and write commands and timing between the processor and memory<br \/>\n<span style=\"color: red;\">* <\/span>Tera Operations Per Second(TOPS): A throughput metric that measures the number of operations performed per second in trillions. One TOPS represents one trillion operations per second; here, it refers to the scale of bit-level operations performed within DRAM<\/div>\n<div style=\"height: 16px; line-height: 16px;\"><\/div>\n<div style=\"text-align: center;\">\n<div style=\"display: inline-block; max-width: 748px; width: 100%; text-align: left;\">\n<div style=\"height: 2px; background: #666666; margin-bottom: 6px;\"><\/div>\n<h3 class=\"sub-title\" style=\"margin: 0; line-height: 1.4;\">Next: Why is the semiconductor paradigm shifting toward memory?<\/h3>\n<div style=\"height: 2px; background: #666666; margin-top: 6px;\"><\/div>\n<\/div>\n<\/div>\n<p>Bottlenecks in AI infrastructure arise not only from individual components such as GPUs or memory, but especially from how computation is distributed and communication is performed, i.e., from how GPUs\/NPUs\/TPUs, HBM, storage, networks, and interconnects exchange data. The major bottleneck we face today is the huge amounts of data movement caused by the processor-centric paradigm. High-speed interconnects, near-memory accelerator architectures, and disaggregated architectures explored above are all system-level approaches to reducing the large burden of data movement.<\/p>\n<p>Now, the question turns to the architecture of semiconductors and memory themselves. Beyond simply placing data closer to compute devices, how much more data movement could be reduced if some computations were performed within or around memory itself?<\/p>\n<p>The next article will explore why the center of gravity in AI semiconductors is shifting toward memory, with a focus on HBM, advanced packaging, and processing-in-memory(PIM).<\/p>\n<p><img loading=\"lazy\" decoding=\"async\" class=\"alignnone size-full wp-image-13425\" src=\"https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/28155500\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_04_graphic_2026.jpg\" alt=\"\" width=\"1600\" height=\"484\" srcset=\"https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/28155500\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_04_graphic_2026.jpg 1600w, https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/28155500\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_04_graphic_2026-300x91.jpg 300w, https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/28155500\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_04_graphic_2026-1024x310.jpg 1024w, https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/28155500\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_04_graphic_2026-768x232.jpg 768w, https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/28155500\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_04_graphic_2026-1536x465.jpg 1536w, https:\/\/d18r0a86za96sg.cloudfront.net\/wp-content\/uploads\/2026\/09\/28155500\/AI-Ecosystem-Redesigning-infrastructure-Why-architecture-determines-performance_04_graphic_2026-1200x363.jpg 1200w\" sizes=\"auto, (max-width: 1600px) 100vw, 1600px\" \/><\/p>\n<p style=\"text-align: left;\"><em><strong>Disclaimer:<\/strong> <span style=\"color: #808080;\">The opinions expressed in this article are solely those of the author and do not necessarily reflect the official position of SK hynix.<\/span><\/em><\/p>\n","protected":false},"excerpt":{"rendered":"<p>As AI workloads evolve toward inference and agentic AI, the criteria for system performance are also changing. Computing power alone is no longer enough. Where data is stored, how quickly it moves, and how efficiently it is processed now have<\/p>\n","protected":false},"author":41,"featured_media":13430,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"_acf_changed":false,"_migrated_source_id":0,"footnotes":"","_members_access_role":[],"_members_access_error":""},"categories":[5],"tags":[1598,155,1597,13],"class_list":["post-13417","post","type-post","status-publish","format-standard","has-post-thumbnail","hentry","category-tech-and-ai","tag-agentic-ai","tag-cxl","tag-gpu","tag-hbm"],"acf":[],"_links":{"self":[{"href":"https:\/\/news.skhynix.com\/en\/wp-json\/wp\/v2\/posts\/13417","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/news.skhynix.com\/en\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/news.skhynix.com\/en\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/news.skhynix.com\/en\/wp-json\/wp\/v2\/users\/41"}],"replies":[{"embeddable":true,"href":"https:\/\/news.skhynix.com\/en\/wp-json\/wp\/v2\/comments?post=13417"}],"version-history":[{"count":13,"href":"https:\/\/news.skhynix.com\/en\/wp-json\/wp\/v2\/posts\/13417\/revisions"}],"predecessor-version":[{"id":13470,"href":"https:\/\/news.skhynix.com\/en\/wp-json\/wp\/v2\/posts\/13417\/revisions\/13470"}],"wp:featuredmedia":[{"embeddable":true,"href":"https:\/\/news.skhynix.com\/en\/wp-json\/wp\/v2\/media\/13430"}],"wp:attachment":[{"href":"https:\/\/news.skhynix.com\/en\/wp-json\/wp\/v2\/media?parent=13417"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/news.skhynix.com\/en\/wp-json\/wp\/v2\/categories?post=13417"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/news.skhynix.com\/en\/wp-json\/wp\/v2\/tags?post=13417"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}