The ecological cost of a token: cloud inference vs. your laptop
Every prompt you send carries an invisible bill. Not the one you pay in dollars — the one the planet pays in electricity, water and raw materials. When a large language model generates a token in a data centre, a chain of physical costs is triggered: the GPU draws power, that power had to be generated, the building had to be cooled and the hardware itself had to be manufactured. None of this shows up in the API pricing page. All of it has an ecological impact.
The natural question for anyone running local AI is therefore: how does one token generated on my laptop compare to the same token generated in the cloud? We looked at the different existing studies on the subject and we built an interactive calculator so you can estimate your own footprint.
The numbers below are orders of magnitude, not precise invoices. Real-world footprints vary with the model size, the data centre's cooling technology, the local electricity mix and the utilization rate of the hardware. Treat them as a way to compare scales, not to audit a specific provider.
1. What a token actually costs in electricity
The most cited peer-reviewed baseline comes from the MIT Lincoln Laboratory study "From Words to Watts", which benchmarked LLM inference on server-grade GPUs and reported energy costs in the range of 0.3 to 3 watt-hours per 1,000 tokens, depending on model size and hardware (for a 70-billion-parameter model on NVIDIA A100 GPUs, generation lands around 1–2 Wh per 1,000 tokens).
Cloud inference rarely stops at the GPU's own draw. The server's other components, the power supply losses, the network and above all the data centre overhead (air conditioning, power distribution, lighting) multiply the on-chip figure. The industry measures this with the PUE (Power Usage Effectiveness): a modern facility runs a PUE around 1.1–1.2, meaning roughly 10–20% extra energy on top of the IT load — and older or hotter-region facilities can run much higher. Air conditioning plays a special role here. Rejecting the heat produced by the GPUs consumes electricity for pumps and fans, and that is exactly what the PUE captures.
There is also a batching subtlety: a data-centre GPU serves many users at once, which amortizes its idle draw, but a frontier model with hundreds of billions of parameters still burns far more joules per token than a 7B local model ever will.
On your laptop, the picture is simpler and harsher at the same time. A laptop GPU or NPU draws its full power for every token with no batching to amortize it, but the model is also typically 10 to 50 times smaller than a frontier cloud model. Measured on real hardware (Apple Silicon M-series and similar), a 7–8B quantized model generating at 20–40 tokens per second draws on the order of 0.3 to 0.8 Wh per 1,000 tokens including the whole machine.
Per token, a small local model on a laptop is in the same ballpark as cloud inference — sometimes even better, sometimes worse. The local advantage is not in the raw joules of a single token, but in everything that comes with those joules. That is where the gap opens up: on the cloud side, those joules come bundled with the water evaporated to cool the servers, the data centre's extra energy overhead (distribution, cooling, lighting), the manufacturing and frequent replacement of dedicated accelerators, and the critical minerals and electronic waste they generate. On your laptop, those side costs either do not exist or are already amortized.
2. Water: the footprint nobody prices
Electricity is only the visible part of the bill. A landmark study from UC Riverside, "Making AI Less Thirsty" (published in Communications of the ACM), estimated that training GPT-3 in Microsoft's US data centres directly evaporated around 700,000 litres of clean freshwater — and that global AI demand could withdraw 4.2 to 6.6 billion cubic metres of water per year by 2027, more than the total annual water withdrawal of four to six Denmarks.
Water enters the equation twice:
- Onsite, evaporation-based cooling towers reject the heat produced by the GPUs. The rough industry figure is 1 to 2 litres evaporated per kWh of heat removed and the World Economic Forum reports a 1-megawatt data centre can consume up to 25.5 million litres of water per year just for cooling.
- Offsite, thermoelectric power generation itself is water-hungry — roughly 1 to 2 litres consumed per kWh generated, before the electricity even reaches the data centre.
Combine these and a commonly cited estimate for GPT-3-style inference lands around 500 mL of total water per 1,000 tokens in water-stressed US regions, with wide variation. The UNU-INWEH report "Environmental Cost of AI's Energy Use" (2026) confirms the scale: the electricity-associated water footprint of one AI-generated image is about 29 mL, and a short complex AI video about 4.1 litres — nearly two days of one person's drinking water.
And your laptop? Zero litres of cooling water. A laptop is passively or air-cooled; no cooling tower, no evaporation, no draw on a municipal supply in a drought-stricken region. The offsite water cost of generating your household electricity still applies to both scenarios — but the entire onsite cooling burden of the cloud simply disappears on your desk.
A lesser-known shadow hangs over data centre cooling: PFAS. To save water, part of the industry is turning to two-phase liquid cooling, which uses no water but a fluorinated fluid — in other words PFAS, the "forever chemicals" that do not degrade and accumulate in the environment and in the human body. A ChemSec report published in September 2026 shows that most of the world's ten largest PFAS producers are expanding capacity, partly to meet demand from data centres and semiconductors. In July 2026, 17 environmental organisations asked the US EPA to reject the fast-track approval of Opteon 2P50, a new PFAS intended for data centre cooling, judging the toxicological assessment insufficient. These systems are presented as closed loops, but leaks ("fugitive emissions") and the end-of-life disposal of the fluids release these compounds, which break down in the air into TFA, classified as a hazardous substance by the European Chemicals Agency. According to ChemSec, the societal cost of PFAS reaches €440 billion in Europe, and a European Commission study puts annual health costs from PFAS alone at €39.5 billion. Contamination is already widespread: according to NHANES data, 97% of Americans have PFAS in their blood. Waterless cooling is therefore not pollution-free cooling.
3. Natural resources: minerals, land and e-waste
The third footprint is the least discussed and the one whose effects take the longest to materialize. AI hardware is built from critical minerals — cobalt, tungsten, tantalum, rare earths — often extracted in regions with weak environmental oversight. The UNU-INWEH report projects that AI infrastructure could generate up to 2.5 million tonnes of electronic waste per year by 2030, much of it processed in lower-income countries with limited safeguards. The same report estimates the land footprint of AI's 2030 electricity consumption at over 14,500 square kilometres — twice the Jakarta metropolitan area — and its carbon footprint at 399 million tonnes of CO2, which would need 6.7 billion trees grown for ten years to offset.
Here the local story is nuanced. Your laptop was also manufactured with critical minerals and manufacturing emissions are real: by most lifecycle analyses, 70 to 85% of a computer's lifetime carbon footprint is already spent before you first boot it. But there is a decisive difference in amortization: a cloud data centre's accelerators run a two-to-five-year refresh cycle, because efficiency gains per generation are brutal competitive weapons. Your laptop runs five to ten years and every hour of inference on it costs no additional manufacturing — the hardware already exists and would be powered on anyway.
The greenest token is the one generated on hardware you already own, with a model small enough for the task. Squeezing 5 million tokens out of your existing laptop instead of a 70B cloud model avoids the water and grid burden of someone else's data centre without adding any new hardware to the planet.
4. The scale problem: inference dominates
Public debate tends to focus on the spectacular training runs — GPT-3's 1.3 GWh, GPT-4's estimated 50–70 GWh. But the UNU-INWEH report confirms what operators already knew: once a model is deployed, inference accounts for 80 to 90% of its total energy use. ChatGPT alone is estimated to process around 2.5 billion prompts per day — roughly 383 GWh of electricity per year for a single product.
This is where your personal arithmetic matters. The IEA estimates global data centres consumed about 448 TWh in 2025 (more than most countries) and projects 945 TWh by 2030. You cannot change where providers site their data centres, but you do choose where your tokens are generated. Moving your routine workload — the summarization, the drafting, the everyday coding help — to a local model that is fit for purpose removes your share from that curve. Our cost comparison showed the financial break-even; the ecological logic follows the same slope, with water and minerals added on top.
5. Try it yourself
We turned these published figures into an interactive tool: enter how many tokens you use (or plan to use), pick a cloud model size, a local machine and an electricity mix and the calculator estimates the full lifecycle footprint — manufacturing and operation — for both paths side by side: electricity, carbon, water and mineral resource depletion.
The model applies a PUE of 1.2 to the cloud's operational energy, adds an embodied manufacturing share of about 25% on top (LLMCarbon), then converts that lifecycle electricity to carbon using the selected grid mix. Water follows Li et al.: 0.5 L per kWh of onsite cooling plus 1.5 L per kWh of offsite generation for the cloud, and only the offsite share for an air-cooled local machine. Mineral depletion is amortized from GPU production in milligrams of antimony equivalent, expressed per token served (Morand et al.). A local machine you already own carries no incremental manufacturing — only the electricity it draws.
Sources
- From Words to Watts: Benchmarking the Energy Costs of Large Language Model Inference — Samsi et al., MIT Lincoln Laboratory (arXiv:2310.03003)
- Making AI Less "Thirsty": Uncovering and Addressing the Secret Water Footprint of AI Models — Li, Yang, Islam, Ren, UC Riverside, Communications of the ACM (arXiv:2304.03271)
- LLMCarbon: Modeling the end-to-end Carbon Footprint of Large Language Models — Faiz et al., ICLR 2024 (arXiv:2309.14393, embodied carbon and PUE modelling)
- The Rising Unsustainability of AI Graphics Cards Production — Morand, Névéol, Ligozat, LIMITS 2026 (arXiv:2607.01258, abiotic depletion of GPU production)
- Environmental Cost of AI's Energy Use: Carbon, Water and Land Footprints — Aczel et al., UNU-INWEH, United Nations University, 2026
- How to make AI data centres more sustainable — UNEP technical highlight, 2026 (415 TWh estimate for 2024, doubling by 2030)
- AI's environmental costs threaten water, land and climate — UN News, June 2026
- UNEP releases guidelines to curb the environmental impact of data centres — UNEP, June 2025 (World Economic Forum cooling-water figure)
- Energy and AI — IEA flagship report, April 2025
- LLMCO2: Advancing Accurate Carbon Footprint Prediction for LLM Inferences — Fu et al. (arXiv:2410.02950)
- The world's top 10 PFAS producers — most are expanding production — ChemSec (International Chemical Secretariat), September 2026 (€440 billion societal cost in Europe)
- US environmental groups urge EPA to reject new PFAS to cool datacenters — The Guardian, July 2026 (Earthjustice comments to the EPA on Opteon 2P50)
- The cost of PFAS pollution for our society — European Commission, DG Environment, 2026 (€39.5 billion in annual health costs)