AI is no longer defined by a single model or chip. For AI to operate effectively in real-world services and industrial applications, it takes faster compute, greater memory bandwidth, higher-performance networking, more efficient storage, and stable power and cooling working together. This is why competition in AI is shifting beyond individual technologies to the design and operation of the entire infrastructure.
In the AI Infrastructure Insight series, we explore how the architecture behind AI infrastructure is evolving — from compute, memory, storage, networking, and power and cooling to the way they come together as a unified system — and what this means for the industry.
[Series overview]
① What is changing inside AI data centers?
② Why faster GPUs alone can’t deliver AI performance
③ Why power and cooling have become the next challenge for AI infrastructure
④ How AI infrastructure will be designed for the future
When people think of AI, the first things that often come to mind are massive models and high-performance semiconductors. Training on ever-larger datasets, delivering faster inference, and handling increasingly complex workloads all require powerful compute capabilities. But as AI expands into real-world services and industrial applications, the focus of competition is shifting beyond the performance of individual chips to the architecture of the supporting infrastructure.
At the center of this shift is the data center. Once primarily designed to store and process data using servers and storage systems, today’s data centers have evolved into the core infrastructure powering AI services — training AI models, processing countless inference requests, and moving data in real time. Modern AI data centers can deliver reliable AI services only when compute*, memory, storage, networking, and power and cooling are seamlessly integrated into a single system.
Why AI data centers are being redefined
The growing importance of AI data centers reflects more than the need to deploy additional servers. It is being driven by a fundamental shift in AI. As AI models become larger, both computational demand and data movement increase. Training requires massive datasets to be read and processed repeatedly, while inference depends on retrieving the right information quickly in response to user requests. As AI applications expand across search, workplace productivity, content creation, manufacturing, finance, healthcare, and other industries, the volume and frequency of requests handled by data centers will continue to grow.
This shift is also changing how data centers are designed. In the past, the key priorities were maximizing server capacity and storage volume. In the AI era, however, performance and efficiency increasingly depend on how compute, memory, storage, networking, and power and cooling are architected and integrated. Even the most powerful processors cannot deliver their full potential if storage fails to supply data fast enough, memory cannot feed data efficiently, networks become bottlenecks, or power and cooling cannot keep pace.
As a result, today’s data centers have evolved beyond facilities for storing and processing data. They have become the production infrastructure that enables AI services to operate reliably in real-world environments. As AI continues to move from research labs into everyday life and industrial applications, data centers are becoming the foundation that turns AI’s potential into real-world services.
☑ How far can data centers go?
Today’s data centers are evolving beyond facilities housed in massive buildings, taking on a wide range of forms depending on their operating environments and intended purposes. Through Project Natick*, Microsoft demonstrated the feasibility of operating data centers underwater while validating the cooling efficiency of subsea environments. More recently, efforts to move AI computing beyond Earth have advanced from concept to demonstration. In November 2025, Google unveiled Project Suncatcher*, an initiative to perform AI computing in orbit using satellites equipped with its Tensor Processing Unit (TPU) chips, with two demonstration satellites scheduled for launch in early 2027. That same month, startup Starcloud* launched a satellite carrying NVIDIA GPUs and successfully trained an AI model in space for the first time in December 2025. These examples are less about the adoption of specific form factors than about a fundamental shift in how data centers are conceived.
The questions surrounding data centers in the AI era are changing. The focus of data center design is no longer simply how many servers can be deployed, but how to optimize power delivery, cooling, networking, and data movement as an integrated system.
As AI workloads become increasingly driven by reasoning and agentic AI, infrastructure performance depends not only on adding more compute, but also on creating an environment where data can be stored, transferred, and processed with maximum efficiency.
* Project Natick: An experimental project launched by Microsoft in 2014 to evaluate the feasibility of underwater data centers. In 2018, the company deployed a sealed vessel containing 864 servers off Scotland’s Orkney Islands, where it operated for nearly two years to test data center operations and cooling efficiency in an underwater environment before being retrieved in 2020.
* Project Suncatcher: A Google research initiative aimed at performing large-scale machine learning computation in Earth orbit. In collaboration with satellite company Planet, Google plans to launch two demonstration satellites in early 2027.
* Starcloud: A U.S. startup developing data centers in Earth orbit. Backed by NVIDIA, the company launched a satellite carrying an NVIDIA H100 GPU in November 2025, marking the first time a data-center-class GPU was deployed into orbit.▲ Microsoft’s Project Natick validated the feasibility of operating data centers in an underwater environment. (Source: Microsoft Research, Project Natick)
The 5 key factors driving AI data centers

To understand AI data centers, it is important to look at the core components that make them work. The key is not to view each component in isolation, but to understand how its role has evolved in the AI era and why all of them must work together as an integrated system.
The first is compute. GPUs, AI accelerators*, and CPUs serve as the engines that process large-scale AI training and inference. As AI models become more complex, compute performance becomes increasingly important. In AI data centers, however, compute no longer delivers performance independently. Instead, it is increasingly defined as a computing foundation that performs effectively only when it receives a stable supply of data and is connected to other infrastructure resources with minimal latency.
The second is memory. Memory is the layer that stores the data required for computation and delivers it to compute devices with sufficient bandwidth when needed, enabling processors to access and process data quickly and efficiently.
To maximize accelerator performance, the required data must be available close to the compute resources and delivered at the right time. That is why memory in AI data centers is not a single product, but a layered system. High Bandwidth Memory (HBM), which is located next to accelerators and handles the widest bandwidth, and server DRAM, which offers greater capacity and is shared across the entire server, each take on different roles.
As a result, memory has evolved into a critical layer that bridges compute and data, going beyond the role of a supporting component.
The third is storage. Storage provides the foundation for storing and retrieving massive volumes of training data, model data, user requests, and computational results. While the role of storage in the past was traditionally associated with long-term data retention, the AI era places increasing importance on how quickly data can be stored and retrieved, making storage performance a key determinant of system efficiency and application responsiveness.
The fourth is networking. Large-scale AI training and inference rely on multiple servers and accelerators working together as they exchange data while processing a single workload in parallel. Networking has therefore evolved beyond a communications layer into the foundation that enables distributed compute resources to operate as a unified system. As AI workloads continue to grow, latency*, bandwidth*, and data transfer efficiency have become critical factors in determining the scalability of AI data centers.
* Bandwidth: The amount of data that can be transmitted over a given period of time.
The fifth is power and cooling. As high-performance systems become more densely deployed, both power density* and heat generation increase. Power and cooling can no longer be considered operational factors addressed after deployment. Instead, they have become foundational design requirements that must be incorporated from the outset to ensure system stability and scalability.
How these five components are arranged and interconnected in a real-world environment defines the system architecture. Servers, racks*, and clusters* are the design units that integrate these core components into a functioning AI infrastructure. As AI workloads continue to expand, the focus of system design is shifting beyond individual servers toward rack- and cluster-scale architectures.
* Cluster: A group of interconnected servers or racks configured to operate as a single computing system.
The primary unit of system design is expanding from servers to racks and clusters
Another major shift is the growing scale of system design. Traditional data centers focused primarily on the performance and operational efficiency of individual servers. In AI data centers, however, the design focus is expanding beyond servers to the rack and cluster levels.
Large-scale AI training and inference can no longer be handled by a single server. Multiple AI accelerators and servers must exchange data, coordinate computation, and divide massive workloads across the system.
As a result, data center architecture is evolving from server-centric designs to rack-scale and ultimately cluster-scale architectures.
At the rack level, compute, memory, power delivery, cooling, and networking must be designed together as an integrated system. At the cluster level, multiple racks must operate as a single computing system, making network architecture and data movement efficiency even more critical. This article provides a high-level overview of that evolution, while later installments in the series will explore in greater depth how AI system architectures are continuing to evolve.

Market data also reinforces this trend. According to Omdia*, the AI data center chip market is projected to grow from $123 billion in 2024 to $207 billion in 2025, reaching $286 billion by 2030. As investment in AI infrastructure continues to accelerate, data centers are evolving beyond facilities that simply house more chips. They are becoming highly integrated systems in which chips, servers, racks, and clusters operate as a unified whole.
The growing scale of modern AI data centers reflects this transformation. Microsoft began operating its Fairwater AI data center in Wisconsin in June 2026. Earlier, when introducing the facility, Microsoft noted that the data center designed to store and process the data generated and consumed by its AI cluster stretches the length of five football fields. The example illustrates how compute, data, and storage infrastructure are increasingly being designed and deployed at massive scale.
Data center design as a strategic investment priority
Such trends are also expanding the scope of data center investment. Data center investment is expanding beyond servers and IT equipment to encompass the broader AI ecosystem, including AI semiconductors, networking, power and cooling, construction, and infrastructure. For cloud providers, designing data centers and securing sufficient power have become critical to delivering reliable AI services. Server and networking companies are focused on improving performance and connectivity at the rack and cluster levels. Power and cooling providers are emerging as essential infrastructure partners in enabling the expansion of AI data centers. Semiconductor companies, meanwhile, must look beyond chip performance alone and consider how their devices perform within real-world data center environments.
This transformation is also elevating the role of memory. Memory serves as the critical layer that enables AI systems to access data quickly while ensuring a seamless flow of data between compute resources and data. By keeping data moving efficiently, it helps maximize overall system performance. As memory takes on a greater role across AI infrastructure, the perspective of memory companies carries greater weight as well.

The relationship between data centers and everyday life is evolving as well. In Devon, England, waste heat generated by a compact data center about the size of a washing machine has been used to heat a public swimming pool. Other projects have demonstrated how small data centers installed in backyard sheds can provide home heating, while high-performance GPU servers placed beneath office desks have been used to help heat workspaces. These examples show that heat generated by servers does not have to be viewed simply as waste. Instead, it can be recycled and repurposed as an energy resource for local facilities. In this regard, the role of data centers is expanding beyond supporting AI services to become industrial infrastructure closely connected with power systems, cooling technologies, and local infrastructure.
Next: Following the flow of data
The message behind these changes is clear. The competitiveness of AI infrastructure depends not on the performance of individual components, but on the ability to connect and optimize multiple technologies into a unified system. Only when compute, memory, storage, networking, and power and cooling work together seamlessly can a data center provide the foundation for running AI services reliably in real-world environments. The next step is to examine how data moves through that system. Understanding AI data centers as multilayered infrastructure is only the starting point. In practice, AI performance depends on where data is stored, the path it takes through the system, and how quickly it reaches the compute resources that need it. In the next article, we will explore why faster GPUs alone are not enough to deliver AI performance — and how data movement and system bottlenecks have become the next critical challenge.
<References>
- Image Source: Microsoft Research, Project Natick
- Data Source: Omdia, AI Processors for Cloud and Data Center Forecast (2025)
- Omdia Press Release: “AI Data Center Chip Market to Hit $286bn; Growth Likely Peaking as Custom ASICs Gain Ground,” Aug. 28, 2025.
- Quote Source: Microsoft, Microsoft completes construction on first datacenter facility in Mount Pleasant, Wisconsin, June 23, 2026.
