Open source ยท Apache-2.059 factors ยท 8 domains 69 data sources

Can this land hostan AI data center?

An agent harness that evaluates a place on Earth across power, water, climate, land, connectivity, regulation, community and supply chain โ€” and reports how well it actually knows each answer.

In plain terms: picking where to build an AI data center normally takes a six-month consulting engagement. Most of the information needed is public, but it is scattered across grid queues, county planning portals, satellite data and local news in a hundred incompatible formats.

This project gathers it automatically and scores a location. The unusual part is that it labels every single number with how it was obtained โ€” measured from a live scientific API, read off a government file, researched from news reports, or estimated. A confident score built on guesses is worse than no score, so the system refuses to hide the difference.

01

Sites analyzed

Seven locations across the United States, China and India โ€” chosen because they are mostly known outcomes. If the model cannot recognize sites that were actually built, it is not measuring anything real.

02

The clearest result: cooling climate

Wet-bulb temperature โ€” not ordinary temperature โ€” decides how hard a data center is to cool. It is one of the few things this system can measure precisely anywhere on Earth, from public weather reanalysis, with no API key.

Why this matters, and why intuition fails here. Abilene, Texas is 5.2 ยฐC hotter than Ashburn, Virginia on a design summer day. Yet Abilene's design wet-bulb is cooler, and it spends less than half as many hours above the threshold where evaporative cooling stops working well. West Texas is the better cooling climate. "Hotter is worse" gets it backwards, because dry heat sheds easily and humid heat does not.

Ulanqab in Inner Mongolia records zero hours above that threshold across a decade of hourly data. China designated it a national computing hub, and this harness surfaced the reason from open data alone โ€” it was told nothing about the designation.

03

Every number says how well it is known

This is the core idea. A score of 82 built on satellite measurements is a different asset from a score of 82 built on news articles, and most tools present them identically.

Tier A

A live scientific API returned the number. Reproducible today and next year.

Tier B

An official published dataset or government file was parsed.

Tier C

Researched on the web and synthesized, with links. Real evidence, weaker footing.

Tier D

Modeled, estimated, or supplied as an assumption. Labeled loudly as such.

Unknown

Could not be measured. Recorded with a reason โ€” never quietly filled with an average.

Most of the picture is still missing, and the site says so. Across these seven analyses, most factors are unmeasured. That is not a bug being hidden โ€” it is the reason every score carries a ยฑ band. Two sites whose bands overlap are reported as not separable rather than ranked, because ranking them would be a fiction.

04

What this cannot do

Stated plainly, because a tool that oversells itself gets trusted once.

It is not due diligence

It tells you which twenty sites out of two thousand deserve a real investigation, and what to ask when you get there. It does not replace the investigation.

It cannot predict a utility or a county board

It reports published queue data and historical timelines. Whether a utility will agree to serve 500 MW is a negotiation, not a measurement.

Land ownership is largely invisible

Cadastral data is openly published in about a dozen countries. Elsewhere the system resolves to administrative boundaries and says so rather than inventing parcels.

Coverage is uneven by design

Climate and terrain are strong everywhere. Grid capacity, land price and local politics are weak nearly everywhere. The tiers make that visible per factor.

Fiber route data is genuinely poor

Carriers treat routes as confidential. Long-haul proximity is often inferred from rights-of-way and marked as an inference.

A paired control did not separate

Two sites in one county โ€” one approved, one rejected โ€” diverge sharply on the community domain but overlap on total score. The tool declines to rank them. That is the design working in the case where it costs us the satisfying answer.