Council Post: The AI Infrastructure Stack Is Being Rewritten For The Agentic Era

2026/08/27

Categories: business-finance

As CTO, Sven Oehme drives DDN’s AI data platform strategy, powering the world’s most demanding AI and HPC environments.

getty

Nearly every AI leader I have spoken with has reached the same conclusion: The model is no longer the bottleneck. The next phase of AI will be determined by whether the infrastructure can deliver intelligence reliably, economically and at scale.

The challenge is becoming clear. McKinsey projects that power demand from data centers supporting AI inference could triple, from approximately 31 gigawatts in 2025 to 93 gigawatts by 2030. The International Energy Agency reported that electricity consumption by AI-focused data centers increased by 50% in 2025 alone.

Three trends are reshaping that infrastructure: Agentic inference is driving specialization across the technology stack, storage is becoming an active extension of AI memory, and capital and sovereignty are becoming fundamental architectural considerations.

These shifts point to one conclusion. The architecture that supported the first wave of generative AI will not be sufficient for the agentic era.

Agentic AI Is Driving Heterogeneous Infrastructure​

The industry measured progress largely through the performance of individual processors. Agentic AI changes that equation.

A traditional chatbot may process a request and return a response. An AI agent can plan, retrieve information, call tools, collaborate with other agents, evaluate results and revise its work. A single request can initiate many interconnected model calls, searches, data transfers and computational operations.

These workflows are too diverse to be served optimally by one processor architecture. Training, prefill, decoding, retrieval, reasoning and data preparation each have different computational profiles. Forcing every stage onto one type of processor can increase costs, waste power and leave expensive resources underutilized.

AI infrastructure is becoming heterogeneous by design. CPUs, GPUs and purpose-built accelerators will increasingly operate within the same environment, with each matched to the workloads it can execute most efficiently.

This places greater importance on data. Heterogeneity creates value only when data can move across processors, memory tiers and locations without introducing latency or operational complexity. The most effective architecture will not necessarily contain the fastest individual chip. It will be the one that keeps every component productively engaged.

Storage Is Becoming Part Of The AI Memory Hierarchy​

The second major shift is the changing role of storage.

Storage has been treated as a repository where information is written, retained and retrieved. That definition is no longer sufficient for production AI. Storage is evolving into something more strategic: a data intelligence layer that actively organizes, moves, protects and delivers the data and context on which AI depends.

Agentic applications maintain longer-running sessions, operate across expanding context windows and repeatedly access embeddings, model parameters, retrieval indexes and key-value caches. They must preserve intermediate state so workflows can pause, resume and coordinate across agents and compute resources.

Keeping all that information in accelerator memory is neither technically practical nor economically sustainable. In one long-context inference study using an 8-billion-parameter model and a 1-million-token sequence, the unoptimized key-value cache required 128 gigabytes of GPU memory. Researchers reduced that requirement to one gigabyte by offloading and selectively managing the cache.

When GPUs wait for data, utilization declines while infrastructure costs remain unchanged. Response times increase, throughput falls and the cost of each workload rises. At scale, even brief periods of idle compute create significant capital and energy inefficiency.

The business impact is measurable. In Salesforce’s deployment of EXAScaler, reducing I/O bottlenecks resulted in a "75% reduction in I/O latency, 1.5x faster model training ... [and] a 42% reduction in overall training costs." Although every environment is different, these results demonstrate the relationship among data performance, accelerator productivity and AI economics.

This is not simply a storage-performance problem. It is an AI economics problem, and it is why the data layer has become one of the most strategic components of the AI stack.

Distributed key-value caching, high-performance object access and direct data paths are transforming the data platform into an extension of the AI memory hierarchy. Instead of repeatedly recomputing or reloading context, AI systems can preserve and reuse it across sessions, servers and accelerators.

Data intelligence extends beyond performance. The platform understands how information is organized, governed, protected and made available to models and agents. It delivers the right data and context to the right workload at the right time while maintaining the durability, scale, security and governance required for enterprise operations.

Rethinking How AI Infrastructure Is Measured​

Leaders must look beyond capacity and raw throughput. Organizations should measure time to first token, context-loading performance, accelerator utilization, data movement, energy consumption and cost per completed workload. The question is no longer “How much data can we store?” It is “How effectively can our data intelligence platform keep the entire AI system working?”

Capital And Sovereignty Become Architectural Considerations​

Capital adds another dimension. McKinsey estimates that data centers will require approximately $6.7 trillion in global investment through 2030, with $5.2 trillion supporting AI-capable infrastructure. Organizations must deploy enough capacity to capture demand without overbuilding systems that become inefficient or obsolete.

Capital strategy and technical architecture are now inseparable. Financing influences when capacity can be deployed, where it can operate, how quickly it can expand and whether new technologies can be adopted without stranding earlier investments. The technical stack must be modular, adaptable and capable of scaling incrementally.

Sovereignty must be reconsidered. In the agentic era, it means maintaining control over the systems that direct intelligence and action, including data, models, context, policies, identities and infrastructure. Organizations must determine how models are trained, where inference occurs, what information agents can access, how decisions are audited and who can change system behavior.

The Agentic Era Requires System-Level Architecture​

Specialized processors create value only when data coordinate them efficiently. Accelerators cannot deliver their full economic potential while waiting for data. Persistent agents cannot operate without a scalable context layer. Sovereign control cannot be assured without visibility into where data resides, how it moves and who can access it.

The agentic era requires a system-level architecture in which compute, memory, data intelligence, networking, orchestration, security, governance and capital operate as a coordinated whole. The objective is not maximum performance from one component. It is sustained efficiency across the entire AI lifecycle.

The winners will stop treating infrastructure as a cost to contain and start treating it as the system that converts intelligence into measurable business outcomes.​


Forbes Technology Council is an invitation-only community for world-class CIOs, CTOs and technology executives. Do I qualify?


>> Home