Home Artificial intelligence Agentic AI is bringing CPUs back to the core of enterprise AI infrastructure says AMD’s Derek Dicker
Artificial intelligence

Agentic AI is bringing CPUs back to the core of enterprise AI infrastructure says AMD’s Derek Dicker

Share


Graphics processing units (GPUs) have dominated the artificial intelligence (AI) infrastructure conversation. But as enterprises move towards agentic workflows, the work surrounding the model is becoming harder to ignore: coordinating tasks, calling tools, moving data, and keeping expensive accelerators busy.

For Derek Dicker, Corporate Vice President, Enterprise Business Group at AMD, that shift gives central processing units (CPUs) a more meaningful role. It also changes how enterprises should approach infrastructure, from consolidating older servers to designing complete systems and managing the cost of AI usage.

AMD’s system-level strategy also has an India dimension. The company and Tata Consultancy Services (TCS), through TCS’s HyperVault subsidiary, have announced plans to co-develop Helios-based AI infrastructure, with a data centre blueprint supporting up to 200 MW of capacity.

In this conversation with Dataquest, Dicker discusses AMD’s silicon, open software, and rack-scale strategy; the practical barriers to enterprise AI adoption; and why technology leaders need to question their assumptions constantly.

GPUs have become central to the AI conversation. How do you see this transition, and where does it leave the CPU?

If we take a giant step back in the evolution of our industry, AMD has been involved in several key technology transitions. On the CPU side, that stretches from the client devices you and I are looking at today to the data centre infrastructure that underpins everything we are doing. It really started on the CPU side.

The company had the foresight to acquire GPU assets that started more on the consumer and client side. Building out the Instinct products allowed AMD to participate in the data centre evolution that has unfolded. Under Dr. Lisa Su’s leadership, AMD took on the overall mission to solve the world’s toughest problems, and some of those problems required GPU technology.

Having CPU technology, GPU technology, and the networking technology that can link it all together, particularly in data centre infrastructure, positions the company very well to help our customers usher in the next era of AI.

This is an ever-changing space. If you went back perhaps a year and a half, when we were talking about AI, it was a fully GPU world. What has become very clear with agentic AI is that the CPU takes on a very meaningful role in these workflows.

The reality is that we are all, as an industry, finding our way through what the future looks like. The companies that will be able to help customers the most are the ones that have a full portfolio of those technologies.

Will inference create a more heterogeneous computing environment, with workloads determining where the work runs?

If we look at the evolution of AI workloads, there was an intense focus on training. While training is a critical part of AI infrastructure, we think inference is probably now at roughly 50:50 with training as a percentage of compute utilisation. It is expanding dramatically towards becoming the dominant portion of where compute spends its time.

The compute required for inference will be heavily dependent on where the work is being done. There will be client instances, edge instances, and data centre infrastructure performing inference. Each will have a different requirement, and the optimisation points will be different.

If inference is being performed in physical AI infrastructure, there is probably one set of technology and capability. If a worker is on a laptop that can run an open-source model locally, that will be a different implementation. In the data centre, there will be another.

Each will have its own power and performance characteristics. That is the beauty of the technologies AMD possesses: we have technology for every one of them.

How do AMD’s processors, networking, and software come together as an AI infrastructure strategy?

I would take it to the strategy level and unpack the pieces.

The first is the company’s ability to build semiconductor products. That covers silicon design, package design, and how products are delivered across CPU, GPU, networking, and other technologies. Those are core and foundational. Alongside CPU, GPU, and networking, we are also seeing applications for field-programmable gate arrays (FPGAs). That is one level of the foundation: the products.

The next piece is the standards that interconnect these things. We do quite a bit of work in standards bodies to ensure interoperability among industry ecosystem players. We firmly believe that open is the way to go. That applies to standards and also to software. Open-source software is a very big part of the strategy, and for us that is the ROCm stack.

You will see a focused effort on the developer community and building out ROCm across the CPU and GPU segments of the market.

Lastly, combining open standards with the software and silicon technologies, we have embarked on a journey to build rack-scale systems. Helios is the first example, where we work with the industry to develop standards and specifications to deploy a full rack.

It is not just any rack. It is a 72-GPU, highly orchestrated system that uses ROCm and allows the ecosystem to play a role, including in networking and cabling, within a full rack specification.

That is how we are looking at it: an open focus on standards and software, and a set of core technologies interconnected into a full rack-scale implementation.

DATAQUEST | SYSTEM ARCHITECTURE

AMD’s AI strategy in three layers

Derek Dicker describes how silicon, openness, and rack-scale systems come together.

LAYER 01

Silicon

CPUs, GPUs, networking, and field-programmable gate arrays (FPGAs) provide the core technologies.

LAYER 02

Open standards and software

Industry standards support interoperability; open software, including ROCm, supports developers.

LAYER 03

Rack-scale systems

Helios brings silicon, software, and ecosystem partners together within a complete rack specification.

The infrastructure conversation is moving from individual processors to complete systems.

Source: Dataquest interview with Derek Dicker, AMD. Editorial synthesis.

How can enterprises make room for AI within existing data centres, and what problems do customers bring to you?

When we engage customers and discuss building AI infrastructure into their environments, one consistent theme is that the data centres of today are bursting at the seams. They are fully built out and fully populated. Bringing AI technology into those data centres requires freeing up space.

A lot of our focus, particularly on the EPYC side for which I am responsible, has been on helping customers free up space as they move to the next generation of technology. That involves advanced silicon design, how we design cores, and the power and performance of those cores.

With our fifth-generation products, we can take legacy infrastructure from four or five years ago and replace upwards of eight older servers with a single AMD EPYC-based server. That consolidation effect is often the first step for people who want to modernise their infrastructure and make room for AI.

The fifth generation offers 8:1 consolidation. The calculations we are running and the discussions we are having with customers suggest 13:1 consolidation with the sixth generation. That means taking a fairly large amount of legacy infrastructure and compressing it down to make room.

Not everything is going to be greenfield. We are spending quite a bit of time putting together solutions that fit within existing data centre infrastructure. Within EPYC-based servers, that includes the ability to use PCI Express (PCIe) GPUs, such as the AMD Instinct MI350P, and then build solutions with the ecosystem to help customers solve their problems.

One of the main customer challenges is a large data centre footprint filled with older technology, alongside a strong desire and pressure from management to build AI infrastructure into the environment. Helping customers realise the total cost of ownership (TCO) advantages is a huge part of that discussion.

On the AI rollout side, we also spend time looking at their options. We follow a proof-of-concept approach, often bringing together an original equipment manufacturer (OEM) partner and independent software vendors (ISVs) in a solution that addresses the customer’s particular needs.

What has underpinned AMD’s progress in enterprise computing, particularly its ability to turn innovation into products customers can depend on?

One tenet of AMD’s culture that resonates very well with me is a deep interest in solving the world’s toughest problems, alongside a deep interest in engaging customers and addressing their biggest pain points.

AMD found customers willing to provide challenges: the desire for single-socket systems, and for high-performance, low-power solutions that enable the consolidation I described.

Listening to customers was the first step: taking in the requirements, working closely with them to refine the product definition, and then executing.

Even as an outsider, before I joined two years ago, I saw that the company’s reputation since EPYC was introduced was for on-time delivery of products. That is something customers depend upon.

The combination of listening to customers, solving their biggest problems, and then executing seems to have been the foundation for the success so far.

How is the CPU’s role evolving, from conventional enterprise workloads to GPU clusters and agentic AI?

There is a class of applications that has existed for quite some time that we refer to as general compute. In that space, there is a deep interest in building CPU-based servers that offer better TCO to the end customer.

They can get more work done in a smaller space at a lower cost, with a payback over a certain period.

The next evolution we saw in AI was what the industry refers to as head nodes. These are servers that sit in GPU clusters to help orchestrate traffic. They tend to require higher-frequency CPUs that allow the GPUs to stay constantly utilised.

If you are going to spend money on a GPU cluster, which can be quite expensive, you want to make sure it is as highly utilised as possible. Architecting CPU silicon that can keep it busy is very important. That is the second layer.

The other is the world of agentic AI. If you look back at workloads before agentic AI, you would find a ratio of a single CPU to four or eight GPUs, particularly in some of the larger training clusters and in inference around chatbot architectures.

What we see moving forward is a combination of reasoning, which runs on the GPUs, and quite a bit of orchestration of tool calls and other things on the CPU side. We are getting to upwards of a one-to-one CPU-to-GPU ratio.

When we take these things together and look at what they are doing to the size of the market, we think tremendous growth is coming for the CPU side.

DATAQUEST | AI INFRASTRUCTURE

Agentic AI gives the CPU more work

Three roles that coexist across enterprise and AI infrastructure.

ENTERPRISE WORKLOADS

Run the business

Run existing applications while getting more work done in less space and at a lower cost.

Focus: total cost of ownership

GPU COORDINATION

Keep GPUs busy

Coordinate traffic and support the GPU cluster so expensive accelerators stay productive.

Focus: GPU utilisation

AGENT ORCHESTRATION

Coordinate the work

Orchestrate tool calls and the surrounding workflow, alongside reasoning on GPUs.

Focus: agent workflows

The CPU’s role expands as AI moves from generating answers to executing workflows.

Source: Dataquest interview with Derek Dicker, AMD. Editorial synthesis.

How does AMD’s partner ecosystem help turn those technologies into solutions for particular industries and customers?

When I talk about our industry, I often use the phrase “it takes a village”. There need to be several different entities involved to deliver a solution to an end customer.

The first is the platform side. We need servers built on EPYC processors as the starting point. A tremendous amount of activity has taken place in the OEM and original design manufacturer (ODM) communities since the first generation of EPYC. We now have platform coverage within each of the major OEMs and ODMs.

Here in India, we also have relationships on the Make in India side and are excited about some of the local providers of OEM systems.

The second piece is the software side. A significant amount of effort goes into partnering with software infrastructure providers such as Red Hat, VMware, and Nutanix, which supply some of the software infrastructure customers use.

Then there are the vertical markets, such as banking, financial services, and insurance, telecommunications, manufacturing, and retail. They all have preferred software that they want optimised.

AMD has been investing heavily in those relationships so that when the software runs on EPYC, the customer gets the best result, whether that is power, performance, or the combination of the two in overall TCO. Those are the areas of engagement across the ecosystem where we spend a lot of our time.

What is AMD doing to help enterprises move AI from pilots into production, while managing deployment complexity and cost?

One piece of feedback we receive from customers is: “We have a particular set of work we need to do, and we are not exactly sure how to get started within our infrastructure to build a solution.”

That insight led us to the concept of building appliances. An appliance is essentially an OEM or ODM system with software on top. We validate it and enable a very specific offering that the customer can order and bring into their environment.

By packaging these solutions, we enable customers to get started very quickly. AMD Instinct Coder is available now as the first of a series of appliances we will deploy over time to help customers build out their infrastructure quickly.

We also spend a lot of time telling stories from our own IT infrastructure and our internal use of AI. What we found is that token consumption can sometimes surpass expectations and drive cost.

A key insight from that led to the development of what we call an intelligent token router. From an appliance perspective, it is a similar approach: building hardware with token-routing software, partnering with ISVs to deliver it, and putting those appliances into the market.

You can expect to see more of these types of appliances being used over time to help take friction out of AI adoption.

Does AI require a different mindset for managing technology and responding to change?

As a general principle, we have to assume that change is going to come faster than we have seen historically. The teams that can be nimble are going to have the most success in the new environment.

Organisations that embrace AI and do not wait to do it will benefit greatly. Because the rate and pace of change are so rapid, it is hard to predict exactly where to go next.

Those that stand the best chance will be the ones that are the most agile, nimble, and AI-first.

Andy Grove argued that only the paranoid survive. Which management principle has stayed with you as a technology leader?

I had the privilege of working at Intel when Andy was there, and I think the management principles he espoused have stood the test of time.

For driving a culture in the new AI era, probably the most important thing is to be hypervigilant and very customer-oriented.

What I have seen happen in organisations is that things sometimes get missed because people assume that what happened in the past will always happen in the future. With the rate and pace of change we see due to AI, that is a very dangerous place to live.

We spend a lot of time talking about listening to the voice of the customer and being very focused on questioning all of our assumptions constantly. That needs to continue moving forward.





Source link

Leave a comment

Leave a Reply

Your email address will not be published. Required fields are marked *

Related Articles
Artificial intelligence

This $30 tool lets you compare 50+ AI models without paying for each one

TL;DR: Harness the power of over 50 AI models with this lifetime...

Artificial intelligence

‘That’s so AI!’ What gen Alpha’s biggest insult tells us | Young people

Name: “That’s AI!”Age: Brand new.Appearance: Primary schools everywhere.What’s AI? It’s just a...

Artificial intelligence

UK government rejects ‘kill switch’ idea for dangerous AI

The UK government has rejected the idea of creating a so-called "kill...

Artificial intelligence

Agentic AI: ICO signals its data protection priorities

Stephanie Lees of Pinsent Masons, who specialises in AI and data...