AI GuidesExplainer

Where Does AI Actually Run?

6 minutes to completeEvergreen guide — kept up to date

AI services feel instant and digital, but every prompt is processed by physical computing hardware in a data centre somewhere. This guide explains what happens between submitting a prompt and receiving a response — the infrastructure, electricity, cooling and network requirements — and why the gap between visible AI and invisible infrastructure matters.

When you use an AI service — asking a question, generating an image, asking Copilot to draft an email — the interaction feels instant and weightless. You type or speak. A response appears. There is no obvious sign of anything physical happening.

But the computing behind that response happens in a real building, on real hardware, drawing real electricity and producing real heat. This guide explains what that chain looks like.

What Happens When You Submit a Prompt

The sequence of events behind an AI response is roughly as follows:

  • Your device — laptop, phone or desktop — sends the request over an internet connection
  • The request travels through internet infrastructure to a cloud region operated by the AI provider
  • Within that region, the request reaches a data centre — a large physical building housing computing hardware
  • Inside the data centre, AI accelerators (typically GPUs or custom processors) process the request
  • The model — a large mathematical system — generates a response token by token
  • The response travels back across the network to your device

This sequence may take a few seconds. The computing infrastructure behind it can be on the other side of the world.

What a Data Centre Actually Contains

A data centre is a purpose-built facility housing large numbers of computer servers, networking equipment and storage systems. It requires:

  • Electricity — continuous, reliable electrical supply, often measured in megawatts for large facilities
  • Cooling — the computing hardware generates substantial heat that must be removed continuously
  • Network connectivity — high-speed fibre connections to the internet and between data centres
  • Physical security — access controls, surveillance and perimeter security
  • Backup power — uninterruptible power supplies and diesel or gas generators for resilience
  • Land and buildings — the physical space to house all of the above

The cloud is just somebody else's infrastructure.

AI Hardware: GPUs and Accelerators

Most computing in traditional data centres uses standard server processors (CPUs). AI workloads increasingly rely on specialised hardware.

GPUs — Graphics Processing Units — were originally designed for graphics rendering but are well-suited to AI because they can perform many mathematical operations simultaneously. For large language models, image generation and other AI tasks, GPUs can process workloads much faster than standard processors.

Technology companies have also developed custom AI accelerators — for example, Google's TPUs and AWS's Trainium chips — purpose-designed to run specific AI workloads efficiently.

These systems consume substantially more electricity per server rack than conventional hardware, and generate proportionally more heat. An AI data centre looks different from a conventional cloud data centre — higher-density racks, more sophisticated cooling, higher electrical supply per building.

Training Versus Inference

AI data centres serve two fundamentally different types of workload:

Training

Training is the process of building an AI model. It involves processing vast amounts of data — text, images, code, audio — through the model repeatedly, adjusting the model's parameters to improve its outputs. Training runs can take weeks or months using thousands of GPUs and consume very large amounts of electricity. A major model training run is a computationally intensive event.

Inference

Inference is what happens every time you use a trained model — when you ask a question, generate an image or use an AI feature. Each request requires computation to produce a response. At the scale of millions of users, inference can account for a very large share of ongoing AI electricity demand.

Training creates the model. Inference serves the users. Both require infrastructure at very different scales.

Cloud Regions and Availability Zones

Major cloud providers operate in geographic regions — clusters of data centres in a specific location, such as Ireland, northern Virginia or Singapore. Within each region, there are typically multiple data centres called availability zones. This structure means that if one building has a power or hardware failure, the service can continue from another.

When you use an AI service based in the United States, your request may be served from a US data centre even if you are in the UK, depending on how the service is deployed. Data residency — where your data actually sits — is a meaningful compliance consideration for businesses handling sensitive information.

Electricity: The Critical Dependency

A data centre cannot operate without electricity. Large AI facilities may require hundreds of megawatts — a level of demand that requires dedicated grid connections, substations and in some cases new electricity generation capacity nearby.

Power Usage Effectiveness (PUE) is the standard measure of data-centre electrical efficiency. A PUE of 1.0 would mean all electricity goes to computing. Modern well-designed facilities aim for PUEs below 1.2, meaning less than 20% overhead for cooling, lighting and infrastructure. AI workloads push power density higher than conventional computing, making cooling efficiency increasingly important.

Cooling: Removing the Heat

Every watt of electricity consumed by computing hardware eventually becomes heat. That heat must be continuously removed to keep the hardware operating within safe temperature ranges.

Cooling approaches include: traditional air cooling using raised floors and computer room air conditioning; direct liquid cooling where cold plates are attached to processors; immersion cooling where hardware is submerged in dielectric fluid; and evaporative cooling for overall facility heat rejection, which uses water. High-density AI workloads are driving rapid adoption of liquid cooling for the processors themselves.

Water usage varies significantly by facility design, climate and cooling technology. Some facilities consume very little water. Others use evaporative cooling that requires continuous water supply. This variation means generalised claims about AI water use need careful qualification.

Why 'The Cloud' Obscures All of This

The term 'cloud' implies something weightless, ambient and digital. This framing was useful for explaining to non-technical users that they did not need to manage their own servers — but it has also created a widespread misunderstanding that digital services have no physical infrastructure.

AI is making this misunderstanding visible. As AI data centres require more electricity, more water, more land and more grid infrastructure, communities hosting those facilities experience direct physical consequences. The people using the AI may be thousands of kilometres away. The community hosting the infrastructure is not.

Every AI answer has a physical infrastructure cost somewhere, even when the user never sees it.

What This Means for Businesses

Businesses using AI services are dependent on this infrastructure, even though they never interact with it directly. Infrastructure conditions shape service pricing, availability, data location and resilience. As AI becomes more embedded in business operations, understanding the infrastructure dependency is part of sound technology governance.

  • Which cloud region is your AI provider using? Is your data staying within a jurisdiction that matters for your compliance requirements?
  • What are your provider's commitments on service availability and resilience?
  • What sustainability information does your provider publish about their infrastructure? Is this relevant to your ESG reporting?
  • Are you dependent on a single AI provider for critical business functions? What would you do if that provider changed pricing, discontinued a service or experienced an outage?

Every AI answer has a physical infrastructure cost somewhere, even when the user never sees it.

Plain-English Takeaway

AI services feel weightless and digital, but the computing happens in large physical data centres with real electricity requirements, real cooling systems and real infrastructure costs. Every AI request follows a physical chain from your device to a data centre, through AI accelerators drawing significant power, back across the network. Understanding this dependency is increasingly relevant for businesses as AI becomes embedded in operations, procurement, compliance and supplier risk management.

Downloadable guide

Where Does AI Actually Run? — Reference Guide

A one-page illustrated reference guide covering the physical infrastructure behind AI services: the prompt-to-response chain, data centre requirements, training versus inference, and questions worth asking your AI provider.

Download PDF

Free to download. No registration required.

Want the full business explanation?

The Technology Intelligence article covers why this matters, where it helps and what to watch out for.

Read the full Technology Intelligence article

Related Knowledge Centre resources

Is Your AI Use Ready for a Policy Question?

Coming Soon

AI Supplier Assessment Checklist

Coming Soon

Still unsure what applies to your business?

Ask the IT Club Advisor about Microsoft 365, browsers, cyber security, productivity or any everyday technology problem.

Free to ask. No credit card. No sales pressure. Fair usage applies.