AI infrastructure model

AI Capacity Calculator

Raw gigawatts say little by themselves. This model follows power through site, grid, accelerator, facility, hardware, system, and model constraints so every assumption can be inspected.

Illustrative baseline

The starter values demonstrate the mechanics. They are not a forecast and carry no empirical claim. Dated, sourced inputs are required before using an output as an estimate.

Editable scenario

United States

Binding constraint
Accelerators
Sites30.0 GW
Power24.0 GW
Accelerators13.7 GW

Capacity ceilings

Completed or buildable data-center shells, measured at the facility boundary.

GW

Power that can reach those sites after interconnection and generation constraints.

GW

Installed accelerators supported by available HBM, hosts, and deployment capacity.

million

Average device draw. CPU, networking, and cooling loads are handled separately below.

kW

Facility and system

Average electrical load divided by the usable facility-power ceiling across the year.

%

Total facility energy divided by IT equipment energy. Lower is better.

×

The balance covers CPUs, memory, storage, and networking inside the IT-power boundary.

%

Delivered accelerator math as a share of peak while the facility is loaded.

%

Losses from communication, storage, scheduling, parallelism, and inference software.

%

Hardware and workload

Use the same numerical precision and sparsity convention for both scenarios.

TFLOP/s/W

Parameters actually used for each generated token; MoE models can be far below total size.

billion

A transparent allowance for attention, context, routing, batching, and other work.

×
Editable scenario

China

Binding constraint
Accelerators
Sites26.0 GW
Power28.0 GW
Accelerators9.4 GW

Capacity ceilings

Completed or buildable data-center shells, measured at the facility boundary.

GW

Power that can reach those sites after interconnection and generation constraints.

GW

Installed accelerators supported by available HBM, hosts, and deployment capacity.

million

Average device draw. CPU, networking, and cooling loads are handled separately below.

kW

Facility and system

Average electrical load divided by the usable facility-power ceiling across the year.

%

Total facility energy divided by IT equipment energy. Lower is better.

×

The balance covers CPUs, memory, storage, and networking inside the IT-power boundary.

%

Delivered accelerator math as a share of peak while the facility is loaded.

%

Losses from communication, storage, scheduling, parallelism, and inference software.

%

Hardware and workload

Use the same numerical precision and sparsity convention for both scenarios.

TFLOP/s/W

Parameters actually used for each generated token; MoE models can be far below total size.

billion

A transparent allowance for attention, context, routing, batching, and other work.

×
Calculated outputs

What the assumptions produce

Outputs change immediately. The ratio is descriptive of the selected scenario, not a forecast.

Usable facility power

The smallest of the site, power, and accelerator ceilings

US
13.7 GW
China
9.4 GW
US 1.46× China

Annual electricity

Usable facility power multiplied by annual load factor

US
96.1 TWh
China
65.6 TWh
US 1.46× China

Useful compute

Sustained throughput after utilization and system losses

US
10.08 ZFLOP/s
China
4.14 ZFLOP/s
US 2.44× China

Annual inference tokens

Useful compute divided by modeled FLOPs per generated token

US
2.59 quintillion
China
1.07 quintillion
US 2.44× China

Inference tokens per facility kWh

A power-normalized view of the full modeled system

US
27.0 million
China
16.2 million
US 1.66× China
Sensitivity test

US sensitivity

10% increase
in each input
Accelerator inventory
+10.0%
Compute utilization
+10.0%
Hardware efficiency
+10.0%
Load factor
+10.0%
System delivery
+10.0%
Model overhead
-9.1%
Active parameters
-9.1%
Site capacity
0.0%
Deliverable power
0.0%
PUE
0.0%
Accelerator power share
0.0%

A zero means that another ceiling remains binding. PUE, active parameters, and model overhead generally run in the opposite direction; their token impact can also be zero while another constraint sets the ceiling.

Sensitivity test

China sensitivity

10% increase
in each input
Accelerator inventory
+10.0%
Compute utilization
+10.0%
Hardware efficiency
+10.0%
System delivery
+10.0%
Load factor
+10.0%
Active parameters
-9.1%
Model overhead
-9.1%
Accelerator power share
0.0%
Site capacity
0.0%
Deliverable power
0.0%
PUE
0.0%

A zero means that another ceiling remains binding. PUE, active parameters, and model overhead generally run in the opposite direction; their token impact can also be zero while another constraint sets the ceiling.

Formula trail

Every multiplier stays visible

1

Find the binding ceiling

Usable facility GW = min(site GW, deliverable power GW, accelerator-supported facility GW)

Accelerator-supported facility GW converts inventory and device draw through accelerator share and PUE.

2

Convert power into annual energy

TWh = usable GW × 8,760 hours × facility load factor ÷ 1,000

Installed capacity and annual electricity are shown separately.

3

Calculate useful compute

ZFLOP/s = active accelerator GW × TFLOP/s/W × compute utilization × system delivery

Precision, sparsity, utilization, networking, and software conventions must match across scenarios.

4

Apply the workload

TFLOPs/token = 0.002 × active parameters (billions) × model overhead

This is a simplified autoregressive decoding approximation. The overhead input carries attention, context, routing, and batching costs.

The token output covers autoregressive inference. The useful-compute result can support a separate training analysis, but the inference token conversion should not be read as training tokens.

PUE follows the standard facility-energy-to-IT-energy definition. Measured wall-power results, such as MLPerf Inference Power, are preferable to unmatched vendor peak specifications when they cover the relevant workload.