Storia.Impact · by AustereBack to Storia

Environmental impact

We know what our AI costs.

Most AI products can’t tell you how much water or power they use. We measured ours, and we publish it.

Last updated

July 2026

Figures

Estimates · production logging live

Scope

Water consumed by inference compute

Next revision

Quarterly

01 · The problem

Start with what the problem actually is.

AI runs on data centers. Data centers run hot, and cooling them evaporates fresh water out of the watershed they sit in.

The water isn’t contaminated. It’s removed. In a region that already has enough, the effect is small. In a region under drought stress, it isn’t.

So the question worth asking isn’t how much water the whole industry uses. It’s how much yours uses, where it goes, and whether you’ve done anything about it.

Here’s ours.

02 · The number

One marketing plan uses less water than a toilet flush.

Every Storia plan is a measured build. We know the exact token count of every step in the pipeline, so we can estimate the compute behind it and the water that cooling that compute consumes.

1–3 liters

Estimated water per marketing plan

~1.8 million

Tokens processed per plan

12 minutes

Time to a finished plan

A token is the unit of work an AI model does. Yours is a big number because your plan is long and specific. The water behind it is still smaller than one flush.

One Storia plan
1–3 L
One toilet flush
≈6 L

Flush figure is the 1.6-gallon US standard, for scale

Scale check · one football field, watered once = 15,000 gal

Storia, 12,000 plans a year · low estimate · 3,200 gal
High estimate · 9,500 gal
One watering of one field · 15,000 gal

A full year of Storia uses less water than a single sprinkler cycle on one field. Not once a week. Once.

03 · Efficiency

The efficiency came first.

We ran an independent audit of every API call in the plan engine and asked one question: where are we spending compute we don’t need to spend?

The answer was nowhere. Every step already runs the smallest model that can do the job well. Heavy reasoning goes to the model that writes your plan. Formatting and structure go to lighter ones. Everything that repeats between steps is cached instead of recomputed, which cuts the work by more than an order of magnitude.

The audit recommended zero changes to reduce compute, because there was nothing left to cut without making your plan worse.

Model mix, per plan

Story + archetypeHeavy reasoning
Plan writingHeavy reasoning
Structure + formattingLight model
Repeated contextCached, not recomputed

Storia · Plan engine

Model Efficiency & Resource Audit

An internal review of the compute behind every Storia marketing plan, run to confirm the engine uses the least energy and water needed to hold output quality.

Conducted July 2026 · grounded in production token counts & two full bench generations · figures marked estimate pending full production logging

Verdict

Zero compute reductions recommended. Every step already runs the smallest model that holds quality.

What one plan costs the planet

1–3 L

Estimated water per plan. Less than a single toilet flush.

~1.8M

Tokens processed per plan, the large majority served from cache.

>90%

Of input served from cache instead of recomputed each step.

At 12,000 plans a year: roughly 3,200–9,500 gallons total, less than one sprinkler cycle on a football field.

How the engine is tiered

Pipeline stageModelWork classFinding
Strategy sections (your plan)Opus 4.8Strategic reasoning, voiceKeep
Plan → dashboard conversionSonnet 4.6Structured synthesisKeep
Production briefs (per post)Sonnet 4.6Client-facing writingKeep
Repeated context across stepsCacheReuse, not recomputeOptimized

What the audit found

Nothing to cut. Heavy reasoning runs on the strongest model only where your plan is actually written. Everything else already runs lighter. No downgrade was recommended because none could be made without lowering output quality.
Caching does the heavy lifting. The repeated context behind each plan is written once and read back for every following step, cutting the real compute by more than an order of magnitude.
Now self-measuring. The engine records real usage on every call, so these figures move from estimate to measured production data, refreshed on this page each quarter.
StoriaEnvironmental impact audit · summary · bystoria.com/impact

“Under your constraints, the model mix is already right.”

Storia plan engine efficiency audit · July 2026

04 · Measurement

Every plan logs its own usage.

Estimates are a starting point. Our engine now records real token usage on every call it makes, which means the numbers on this page come from production, not from a spreadsheet.

We update this page quarterly with actual figures. If the number goes up, it goes up here too.

Revision log

Jul 2026Page published · estimates
Oct 2026First production figures · scheduled
Jan 2027Q4 actuals · scheduled

05 · Next

Where we’re pushing.

01

Regional inference data

We’ve asked our infrastructure provider for regional data on where inference runs and when workload-level environmental reporting will be available. Most of the industry doesn’t publish this yet. Customers asking is how that changes.

02

Region-pinned deployment

We’re also evaluating region-pinned deployment, which would let us choose the exact grid and watershed our compute runs on and match restoration to it precisely.

This page will keep getting more specific. That’s the point of it.

Questions about any number on this page? impact@bystoria.com