26,751 reachable nodes rendered as a cloud of dots collapsing into a row of ten: the 10 nodes that receive more than half of monthly economic flow

research · bitcoin

What Is an Economic Node?

A post-mortem of BIP110 & an attempt to answer the question “what is an economic node?” Many decentralization metrics were brought up prior to activation. This piece defines the object those metrics missed, reviews various measurement frameworks, & conducts our own analysis. While the reachable independent node population hovers around ~27k clearnet nodes, the research indicates that the number of “economically relevant” nodes in a fork war might be closer to 10.

Jesus Najera

Jesus Najera

26,751

reachable nodes

10

received >50% of monthly flow

7

clustered entities to a value majority

3

mining pools to a hash majority

~0

economic nodes that enforced BIP110

01BIP110 asked the network a question

In August 2026 the BIP110 soft-fork proposal tried to activate. It is the cleanest governance experiment Bitcoin has held since Taproot in 2021. Four different “decentralization” metrics looked at the same event & gave four different answers:

  • Node count said it was live: Knots/BIP110-capable clients were 7–15% of reachable nodes, thousands of machines signalling support.
  • Hashpower said it was marginal: miner signalling peaked around 2.5%, one small pool (OCEAN) carrying it.
  • Markets said it was dead: the one prediction venue priced activation at ~2% (98% failure) on trivial volume.
  • Economic acceptance said it was stillborn: no major exchange, custodian, or payment processor committed to enforce it.

The chain sided with the last one. The enforcing branch mined exactly two blocks (heights 961,632–961,633), then stalled while the main chain walked away. No separately-traded asset emerged. The market didn’t blink. Node count & hashpower were both loud, & both wrong about what mattered. The metric that actually called it was “who, weighted by economic acceptance, would enforce the new rule.” & that metric has no agreed name, no agreed unit, no published measurement. It is the economic majority. The economic node. The thing every fork post-mortem leans on & none of them defines.

So we set out to do three things: define it, list every way it could be measured, & actually measure it. The punchline is one chart. As you walk down from the number everyone quotes toward the number that decides forks, the count of actors who actually control each layer collapses by four orders of magnitude.

Log-scale bar chart of actors controlling a majority of each Bitcoin layer: 26,751 reachable nodes, 20 relay ASNs, 10 economic, 3 miners, ~0 BIP110 enforcers

Fig. 1 · THE CONCENTRATION LADDER

The number of actors who must collude, be compelled, or fail to control each layer of Bitcoin. Node count (26,751) is the illusion. Every layer beneath it is single-to-low-double digits. The economic layer, the one BIP110 actually turned on, is around ten. This chart is the whole argument.

02What we should be measuring

Before counting anything, fix the object. The reason node count fails is simple: it counts the wrong layer of a five-layer stack. A running Bitcoin node is a stack of distinct things that casual measurement collapses into one:

Socket
a listening process; what a crawler counts
Operator
the party that controls that process's software & keys
Infrastructure
the network/hosting it depends on (ASN, cloud, Tor)
Economic entity
the business or person whose value flows through it
Value gated
the settlement that is accepted or rejected on its verdict

Node count measures the first line. Economic weight lives in the last two. An economic node is not a socket, an IP, a machine, or an ASN. It is an independently-governed validation boundary whose accept/reject decision gates economically meaningful value. One exchange running 500 processes is one economic node with a lot of infrastructure. One operator spread across five clouds is one economic node. Five hundred unrelated people on one ISP are five hundred economic nodes that happen to share a network dependency. The unit is the decision boundary, & its weight is fork-specific & horizon-specific:

That is the quantity BIP110 measured for real: the sum of Wᵢ over entities willing to enforce it was near zero, so the fork had no economic weight regardless of its node or hashpower signal.

Reading the concentration numbers

Four measures recur throughout. Each turns a skewed distribution into a single honest figure, & each answers a slightly different question.

Nakamoto
the fewest actors whose combined share crosses 50%. A capture / veto count: how many must collude to control the layer. Sharp, but fragile, since it turns on the exact shape at the halfway line.
Gini
inequality on a 0-to-1 scale: 0 is perfectly even, 1 is one actor holding everything. It reads the whole distribution, so it is far more stable than Nakamoto.
HHI
the Herfindahl-Hirschman index, the sum of squared shares. The concentration measure antitrust regulators use; higher means more concentrated.
Hill numbers
effective counts borrowed from ecology: D₀ is raw richness, D₁ is exp(Shannon entropy), D₂ = 1/HHI (inverse Simpson). Read them as “how many equally-sized actors would produce this much diversity.”

03Everything we could compute

There is a whole design space here, not one metric. Here it is, sorted by the one thing that decides whether you can actually run it: what data it needs. Across the prior literature & our own work, twelve constructions span the range from useless to ideal. The honest axis isn’t elegance. It’s data availability, because the good measures all need who-gates-what information that is mostly private.

C0 · Node census
measures relay presence · needs a crawl · runs today: yes, §6
C1 · Symmetric centrality
measures graph structure · needs a crawl · runs today: yes, but gameable
C2 · Cost-anchored throughput
measures value in, cost-capped · needs a clustered chain · partial, §4
C3 · Hitting-time centrality
measures reachability from real value · needs a clustered flow graph · partial, §4
C4 · Fork-game power index
measures pivotality to viability · needs entity weights + quorum · needs C2/C3
C5 · Chain-split market pricing
measures market weight of a ruleset · needs a live fork market · at forks, §1
C6 · Validation-gating estimator
measures who self-validates (v,g) · needs network probing · frontier
C7 · Hodge circulation diagnostic
measures wash vs settlement · needs a clustered flow graph · frontier
C8 · Interdiction / economic min-cut
measures settlement stranded by refusal · needs an entity flow graph · our measure, §5
C9 · Proof-of-Validated-Flow
measures attested weight + ruleset · needs voluntary attestation · proposal
C10 · Miner concentration
measures hashpower coalition · needs pool attribution · yes, §4
C11 · Decentralization functionals
(Nakamoto, HHI, Hill, Gini) · any weight vector · yes, everywhere

The pattern is stark. The measures that need only a crawl (C0, C1) are the gameable ones. The measures that answer the real question (C2, C3, C8) need a clustered entity value-flow graph, & building that at full-chain scale is genuine research infrastructure: a synced node plus hours of graph compute, or paid data. That difficulty isn’t incidental. It is the reason node count survives as the default proxy: it is the only economic-adjacent number that is cheap to produce. The rest of this piece is what happens when you pay the cost anyway.

04Making them work on real data

The economic layer: clustered value-flow (C2/C3)

We pulled a real slice of the chain, 30,006 transactions (7 recent blocks, heights 962,803–962,809) with fully-resolved input addresses & values, built the directed value-flow graph, & ran the standard common-input-ownership clustering to collapse addresses into entities. 26,643 addresses became 18,095 entities. The concentration of inter-entity value received is brutal: Nakamoto 18, effective entities D₂ = 31, Gini 0.978, top-10 entities absorbing 40.9% of all value moved between entities. Eighteen entities gated a majority of the economic value in this window, which is the high end: across 14 sampled windows the median is 7 (range 2–18), unpacked in the robustness study below.

A single window under-clusters. An exchange’s true entity spans years & thousands of addresses, so we catch only fragments of it here, which means these slice figures are over-counts, not under-counts. To pin the number down we ran the full public chain in BigQuery. Over 30 days, 9.38 million addresses received value, & the concentration by receiving address (deliberately unclustered, so an upper bound on the entity count) came back at Nakamoto 10. Ten addresses received more than half of all on-chain value, with the top 100 taking 73.7% & the top 1,000 taking 88.9%. That full-chain 10 is the stable anchor; the per-window clustered figure wanders between 2 & 18 (median 7) depending on the window, which the robustness study below turns into a feature rather than a caveat.

Line chart of cumulative value received by top-N addresses over 30 days: top 10 hold 50.1%, top 100 hold 73.72%, top 1,000 hold 88.91%

Fig. 2 · FULL-CHAIN VALUE CONCENTRATION (30 DAYS, UNCLUSTERED)

Ten addresses received half of all on-chain value; a thousand received nearly nine-tenths. Because clustering merges addresses into fewer entities, the true economic-entity concentration is tighter than this. The clustered per-window slices (Nakamoto 2 to 18, median 7 across 14 windows) & the unclustered 30-day chain (Nakamoto 10) all land in the same place: single digits to low tens.

What the clustering actually looks like

Concentration numbers are abstract, so here is the graph itself. Each colored blob below is one actor, recovered purely from the fact that its addresses were co-spent as inputs to the same transactions. Nothing else. No labels, no external data.

Cluster graph of the ten largest entities; each colored blob is one actor spending from many addresses, the largest a single entity of 386 addresses

Fig. 3 · THE CLUSTERING, DRAWN

The ten largest entities in the slice. Every dot is an address, every edge a co-spend, every color one actor. The biggest blob is a single entity of 386 addresses. This is the operator layer of §2, rebuilt from nothing but the chain.

Now route value between those entities. A few hubs pull most of the flow; the rest barely participate. This is the economic layer becoming visible.

Entity value-flow network of the top 70 entities; red marks the 18 entities holding a majority of flow

Fig. 4 · THE VALUE-FLOW NETWORK

Top 70 entities by value received, edges weighted by value moved, node size by value in. Red marks the 18 entities that together took a majority of inter-entity value in this window. The hubs are visible by eye.

Two log-log charts showing entity size and value received per entity both follow heavy-tailed power-law distributions

Fig. 5 · BOTH DISTRIBUTIONS ARE HEAVY-TAILED

Entity size & value received both follow a power-law shape: a handful of entities dominate, the long tail is inert. The dotted line is where cumulative value crosses 50% (Nakamoto 18 in this window).

Robustness: 14 windows, two weeks of chain

One slice is an anecdote, so we ran the same pipeline on 14 independent windows: 3 blocks each, sampled roughly daily across two weeks of chain (heights 960,857–962,807), clustered & measured identically. The point estimate wobbles. The concentration does not.

Economic Nakamoto across 14 windows (median 7) beside economic Gini holding at 0.97 to 0.99

Fig. 6 · THE ESTIMATE WOBBLES, THE CONCENTRATION DOESN’T

Economic Nakamoto ranges 2–18 (median 7) across the 14 windows; economic Gini sits at 0.97–0.99 (median 0.984) in every single one. The Nakamoto coefficient is fragile because it turns on the exact shape at the 50% cutoff; Gini reads the whole distribution & barely moves.

Then we pushed it back a full year. Eight more windows around height 910,000 (roughly August 2025) tell the same story: the economic layer looked exactly this concentrated a year ago.

Economic Gini and Nakamoto compared across windows one year apart; concentration unchanged

Fig. 7 · A YEAR APART, THE SAME CONCENTRATION

Eight windows from ~August 2025 against the fourteen from now. Economic Gini sits at 0.98 in both cohorts (median 0.988 then, 0.984 now); Nakamoto is single-to-low-double digits in both (median 12 then, 7 now). Whatever moved on-chain in a year, it wasn’t the concentration.

Two honest limits: velocity & stock

Value received is a proxy, & two objections to it are worth taking seriously because they bound the result in opposite directions. We can now answer two of them with data, & only frame the third (elasticity).

Velocity inflates gross flow, but the concentration survives stripping it. An exchange hot wallet that turns the same coins over fifty times looks fifty times heavier than a custodian sitting on the same stock, & intra-entity churn gates nobody else’s settlement. So we recomputed the slice on net inter-entity accumulation, received minus sent per entity, which cancels pass-through. That net measure also sits closer to bcap’s “receive and send substantial payments” definition than gross receipts do. The churn is real: 32% of gross inter-entity flow is pass-through, & 4 of the top-10 gross entities are effectively net-zero conduits. But the headline barely moves. Net-accumulation concentration is Nakamoto 16 against the gross 18, with Gini unchanged at 0.98, because the entities that actually accumulate are themselves concentrated. Velocity was a fair objection; it does not rescue decentralization.

Substitution is unmeasured, so read this as an upper bound. A static max-flow on an observed month of value does not model users rerouting to a competing exchange in days. Concentration of flow is therefore an upper bound on concentration of veto, & how loose that bound is depends on substitution elasticity, which we have not measured. Concentration alone does not prove the gap is narrow. The honest reading of the clustered number is a ceiling on veto concentration, not a point estimate of it.

Flow is blind to stock, & the two disagree. A fork reprices holdings, not just throughput, so we ran the balance-weighted version as well (appendix Query 3), & the result cuts against the flow story, which we will report straight. At the address level, stock is far more distributed than flow: it takes 6,651 addresses to hold a majority of supply (against ~10 for flow), & the top 100 addresses hold just 16% (against 74% for flow). Who holds bitcoin is broad; who moves it is narrow. Two things stop this from settling the fork question either way. First, these are unclustered addresses, & exchanges & ETF custodians deliberately fragment cold storage across thousands of them, so the very entities that would concentrate stock are hidden here; clustering would tighten this number, by how much we cannot say. Second, a fork is gated by who accepts coins, not who passively holds them, which is the flow layer, not this one. So the flow figure stays the fork-relevant one, & stock is a genuinely more-distributed, separate story that entity clustering could yet sharpen.

The miner layer (C10)

Hashpower is the one economic-adjacent layer with clean public attribution. Over the last week & month, three pools (Foundry USA ~25%, AntPool ~18%, F2Pool ~17%) mined a majority of blocks. That’s Nakamoto 3. The full-chain BigQuery cross-check on coinbase-receiving addresses agrees: Nakamoto 3, top-5 addresses taking 77.8% of the subsidy over 90 days.

The market layer (C5)

When a fork exists, its chain-split market prices the economic weight directly. The historical record is thin but consistent: Bitfinex’s 2017 SegWit2x token peaked near 13% of BTC before collapsing; BCH pre-fork futures traded 0.1–0.2 BTC; BIP110’s prediction market sat at ~2%. In every case the market called the outcome the node count missed.

What stayed blocked

The validation-gating estimator (C6, who actually self-validates versus delegates to an API) & full-chain entity clustering (C2/C3 at scale) need infrastructure we couldn’t stand up here: network-wide probing & a synced-node graph pass. We report them as unrun rather than approximate them & pretend. That’s the honest boundary of the exercise, &, again, the reason the cheap proxy persists.

05Our measure: the economic min-cut

Now that we know what we’re looking for, the right measure isn’t a centrality score. It’s a cut. & BIP110 already ran it for us. The fork question is not “who is important” but “whose refusal strands the money.” That is a max-flow/min-cut problem on the entity value-flow graph. Attach a virtual sink to every entity that gates validated acceptance, weighted by the value it accepts; a coalition S that refuses to recognize a fork deletes its acceptance edges. The interdiction centrality of S is the fraction of settlement stranded:

The smallest S with I(S) past the viability threshold is the economic min-cut: the fewest economic nodes whose “no” kills a fork. It’s what people mean by the economic majority, finally made computable. & it answers the title question precisely. An economic node is a member of the min-cut set, an entity whose acceptance decision is load-bearing for settlement. Its weight is its marginal contribution to the cut. Entities with perfect substitutes carry near-zero weight no matter how big they are; a sole fiat off-ramp in a jurisdiction carries large weight at modest volume.

The §4 numbers are the min-cut’s raw material. We stop short of computing I(S) directly here, since that needs the full clustered-&-labeled graph, but the inference is hard to escape. A fork needs assent from whoever gates the majority of settlement. Value acceptance is that concentrated (full-chain Nakamoto 10; median 7 across 14 clustered windows; Gini 0.98 in every one). So the flow-gating min-cut can only be single digits to low tens of entities (a stock-weighted min-cut is a separate, unrun measure; see the limits in §4). That is not a flaw to panic about in isolation. It is why Bitcoin shrugged off a contentious change: a small, aligned economic core said no, & no is all it takes. The same concentration is a strength against capture-by-fork & a risk against capture-of-the-core. Which one it is depends entirely on who those ~10 entities actually are, & that is the measurement the industry still won’t publish.

06Why the node count is the floor, not the answer

The census still matters, as the proof that the popular metric measures the wrong layer. Here’s that case, with its own overclaims sanded off. A live Bitnodes snapshot (2026-08-10, 26,751 reachable nodes) enriched with offline IP→ASN tables shows the relay layer is already far more concentrated than its count. Two-thirds of nodes (65.5%) run on Tor or I2P with no verifiable operator identity. Of the 9,220 clearnet nodes, effective diversity by network is D₂ ≈ 36 (Shannon D₁ ≈ 176). That’s an AS-level infrastructure diversity, not an operator count. The ASN Nakamoto is 20, & three clouds (Akamai/Linode, Hetzner, Google) host over half of all cloud-based nodes.

Stacked bar of reachable nodes by network: Tor 49.7%, I2P 15.8%, IPv4 29.6%, IPv6 4.9%

Fig. 8 · THE RELAY LAYER IS MOSTLY UNAUDITABLE

Two-thirds of the reachable count is on privacy networks whose operator independence cannot be verified. Good for privacy, fatal for a headcount decentralization claim.

& the count is cheap to move. At $4/node/month, a numerical majority of clearnet nodes can be manufactured for roughly $37k/month. Bitcoin Core’s /16 netgroup bucketing raises the true price by forcing address diversity, so read that as a soft floor. Either way, manufacturing sockets doesn’t manufacture validation, acceptance, or economic weight. It manufactures entries in a census. Look at the September 2025 Knots-versus-Core dispute: a ~39% sybil accusation, largely retracted after ~1,000 flagged nodes turned out to be real storefront devices. Nobody could settle it, because the metric under dispute was that cheap.

The historical trend is the sharpest part of the census. Reachable clearnet nodes peaked around 11,442 in 2018 & have been flat-to-declining since (7,512 in 2025; 9,220 here). Every bit of headline growth from ~7k to ~27k is Tor & I2P. The verifiable node surface hasn’t grown in eight years. The census is a genuine result. It proves node count is not a decentralization metric. But it is the floor of the stack in §2, three layers below the one BIP110 turned on.

07So, what is an economic node?

An economic node is an independently-governed validation boundary whose acceptance or rejection of a ruleset gates economically meaningful Bitcoin value. It is a member of the economic min-cut. It is not a socket, an IP, a machine, or an ASN, & there are not 26,751 of them.

Measured three independent ways on real data, the set of entities that move a majority of Bitcoin’s value, a flow-gating proxy for who can veto a fork, is single digits to low tens. Nakamoto 10 by unclustered full-chain receipts, the stable anchor. Median 7 (range 2–18) across 14 clustered windows, with Gini pinned at 0.98 in every one, & unchanged in windows sampled a full year earlier. 3 at the mining layer. & for the specific fork that started all this, effectively 0 willing to enforce, which is why it died in two blocks. Value flow is a proxy for gating power, but a tight one when value is this concentrated. The concentration ladder (Fig. 1) is the whole finding: each step from the quoted number toward the deciding number sheds an order of magnitude.

This reframes every governance argument that leans on node counts. The number that decides forks is not the number of machines; it is the small set of entities whose refusal strands settlement.

What we’ve shown is that the object is real, the unit is a decision boundary, the measure is a cut, & the answer is small. bcap named this object & listed its measurement as unsolved; this is a first number for it. BIP110 didn’t fail because 27,000 nodes voted it down. It failed because the roughly ten that mattered never said yes.

08Methods & reproducibility

  • Node census: Bitnodes latest-snapshot API, 2026-08-10, height 961,942, 26,751 nodes; IP→ASN via offline iptoasn tables, joined by binary search. Effective diversity = Hill numbers (D₂ = 1/HHI). Listening-nodes only, a lower bound on network size; concentration figures are themselves lower bounds on infrastructure centralization.
  • Value-flow slice: 30,006 txs, 7 blocks (heights 962,803–962,809) from mempool.space with resolved prevouts; common-input-ownership union-find clustering; economic weight = inter-entity value received (Nakamoto 18, D₂ 31, Gini 0.978 for this drawn window). Robustness: 14 windows of 3 blocks sampled roughly daily over heights 960,857–962,807 give median Nakamoto 7 (range 2–18) & Gini 0.984 (range 0.969–0.993); a further 8 windows around height 910,000 (~1 year prior) give median Nakamoto 12 & Gini 0.988. A single window under-clusters, so per-window entity concentration is a conservative over-count. Graphs rendered with networkx. Velocity check: recomputing on net inter-entity accumulation gives Nakamoto 16 (vs gross 18), Gini unchanged, with 32% of gross flow being pass-through. Stock concentration (a balance-weighted min-cut) is specified as Query 3 but not run.
  • Full-chain economic weight: Google BigQuery bigquery-public-data.crypto_bitcoin, 30-day window, coinbase excluded; concentration by receiving address (unclustered = upper bound on entity count). Miner cross-check over 90 days of coinbase outputs.
  • Miners: mempool.space pool attribution, 1-week & 1-month windows.
  • BIP110: block heights & signalling figures from public fork-monitoring & post-mortems; the 2-block outcome may partly reflect insufficient hashpower; it demonstrates the layers are separable, not the full economic-weight distribution.
  • Caveats: single node-snapshot; slice not full history; hosting classification is keyword-heuristic (± a few points); the “$4/node” attack cost is a soft floor that ignores /16-diversity costs; BigQuery value-received includes self-churn (which is itself part of the finding).

09Appendix: the full-chain queries

The two BigQuery queries behind §4’s economic-weight & miner numbers, verbatim. Run against bigquery-public-data.crypto_bitcoin in any GCP project; the 30-day window scans a few GB (well inside the 1 TB free tier). Each returns one row.

Query 1 · Economic weight by receiving address (returned Nakamoto 10, top-10 = 50.1%)

-- coins excluded from issuance; concentration by receiving address (unclustered = upper bound)
DECLARE window_days INT64 DEFAULT 30;
WITH recv AS (
  SELECT addr, SUM(o.value) AS v
  FROM `bigquery-public-data.crypto_bitcoin.transactions` AS t,
       UNNEST(t.outputs) AS o, UNNEST(o.addresses) AS addr
  WHERE t.block_timestamp >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL window_days DAY)
    AND NOT t.is_coinbase
  GROUP BY addr ),
ranked AS (
  SELECT v,
    SUM(v) OVER ()                                         AS total,
    SUM(v) OVER (ORDER BY v DESC ROWS UNBOUNDED PRECEDING) AS cum,
    ROW_NUMBER() OVER (ORDER BY v DESC)                    AS rnk,
    COUNT(*) OVER ()                                       AS n_addr
  FROM recv )
SELECT MIN(n_addr) AS addresses_receiving,
       ROUND(MIN(total)/1e8, 1) AS total_btc,
       MIN(IF(cum > total*0.5, rnk, NULL)) AS nakamoto_addresses,
       ROUND(MAX(IF(rnk<=10,  cum,0))/MIN(total)*100,2) AS top10_pct,
       ROUND(MAX(IF(rnk<=100, cum,0))/MIN(total)*100,2) AS top100_pct,
       ROUND(MAX(IF(rnk<=1000,cum,0))/MIN(total)*100,2) AS top1000_pct
FROM ranked;
-- → 9,378,527 addresses · 10 · 50.1% · 73.72% · 88.91% (30-day window)

Query 2 · Miner concentration by coinbase address (returned Nakamoto 3, top-5 = 77.8%)

DECLARE miner_days INT64 DEFAULT 90;
WITH cb AS (
  SELECT addr, SUM(o.value) AS v
  FROM `bigquery-public-data.crypto_bitcoin.transactions` AS t,
       UNNEST(t.outputs) AS o, UNNEST(o.addresses) AS addr
  WHERE t.is_coinbase
    AND t.block_timestamp >= TIMESTAMP_SUB(CURRENT_TIMESTAMP(), INTERVAL miner_days DAY)
  GROUP BY addr ),
r AS ( SELECT v, SUM(v) OVER() total,
    SUM(v) OVER(ORDER BY v DESC ROWS UNBOUNDED PRECEDING) cum,
    ROW_NUMBER() OVER(ORDER BY v DESC) rnk, COUNT(*) OVER() n FROM cb )
SELECT MIN(n) AS coinbase_addresses, ROUND(MIN(total)/1e8,1) AS subsidy_btc,
       MIN(IF(cum > total*0.5, rnk, NULL)) AS miner_addr_nakamoto,
       ROUND(MAX(IF(rnk<=5, cum,0))/MIN(total)*100,2) AS top5_pct
FROM r;
-- → 27,505 coinbase addresses · 3 · 77.76% (90-day window; pools rotate addrs, so this over-counts pools)

Query 3 (proposed, heavy) · Stock concentration by address, the balance-weighted measure the flow number misses (ran on-demand)

-- current balance = value of unspent outputs. Scans the full outputs+inputs tables;
-- heavy (scans outputs+inputs) but ran fine on-demand. Result at the bottom.
WITH bal AS (
  SELECT o.addresses[SAFE_OFFSET(0)] AS addr, SUM(o.value) AS sats
  FROM `bigquery-public-data.crypto_bitcoin.outputs` AS o
  LEFT JOIN `bigquery-public-data.crypto_bitcoin.inputs` AS i
    ON o.transaction_hash = i.spent_transaction_hash AND o.index = i.spent_output_index
  WHERE i.spent_transaction_hash IS NULL          -- unspent = still in the UTXO set
    AND ARRAY_LENGTH(o.addresses) = 1
  GROUP BY addr ),
r AS ( SELECT sats, SUM(sats) OVER() tot,
    SUM(sats) OVER(ORDER BY sats DESC ROWS UNBOUNDED PRECEDING) cum,
    ROW_NUMBER() OVER(ORDER BY sats DESC) rnk, COUNT(*) OVER() n FROM bal )
SELECT MIN(n) AS balance_addresses,
       MIN(IF(cum > tot*0.5, rnk, NULL)) AS stock_nakamoto,   -- min addresses holding >50% of supply
       ROUND(MAX(IF(rnk<=100, cum,0))/MIN(tot)*100,2) AS top100_pct
FROM r;
-- addresses, not entities (clustering tightens it); includes exchange cold wallets and provably-lost coins.
-- → 86,238,708 addresses · Nakamoto 6,651 · top-100 = 16.06% (address-level: far MORE distributed than flow)

The clustering slice (fetch_txs.pycluster_viz.py), the node-census pipeline (analyze.py), & the chart generators are self-contained Python over the same public sources; the union-find common-input-ownership clustering is ~15 lines & reproduces the entity graphs (Figs. 3–5) & the Nakamoto figure for the shipped slice exactly. The robustness studies (Figs. 6–7) are multi_slice.py & multi_slice_yearago.py, checkpointed per window.

10Sources

  • Bitnodes · iptoasn · mempool.space API · BigQuery crypto_bitcoin · primary data
  • Cheng & Friedman, Sybilproof Reputation Mechanisms (2005) · the governing constraint
  • Hill, Diversity and Evenness (1973); Jost, Entropy and Diversity (2006) · effective-number framework
  • Srinivasan & Lee, Quantifying Decentralization (2017) · Nakamoto coefficient
  • Meiklejohn et al., A Fistful of Bitcoins (2013) · common-input-ownership clustering
  • Wood, Deterministic network interdiction (1993); Borgatti, Identifying key players (2006) · the min-cut measure
  • Gencer et al., Decentralization in Bitcoin and Ethereum Networks (FC 2018); Wu & Neumueller, Bitcoin Under Stress (2026) · node-topology baselines & the Tor-resilience counterpoint
  • Ren Crypto Fish, Lee & Alden, Bitcoin Consensus Analysis Project (bcap) (2024) · coined “Economic Nodes”, the flow-weighting axiom & the stakeholder / power-over-time framework this paper measures
  • Lopp, A Layman’s Guide to BIP-110 (2026) · the natural experiment
  • Companion: Measuring the Economic Weight of a Bitcoin Node · the full C0–C11 construction kit

Node count says 26,751. Clustered value flow says ~7. Full-chain receipts say ~10. Miners say 3. BIP110 said ~0, & the chain agreed, in two blocks. An economic node is a member of the set whose refusal strands settlement; there are single digits of them, & the only thing still hidden is their names.

Ready to get started?

Multisig-secured treasury, payroll, and payments for your business.

Login | Register