onchain data bank

how the numbers on this site are made

Methodology

onchain data bank · this edition: 24 September 2026

This page is the counterpart of a claim we make everywhere else – that every answer here shows its work. It says what is collected, from whom, how often, over which windows, what quality means in this product and what happens to a number that cannot be trusted. Nothing on this page has to be taken on faith: every market state names its own contract, and the shape of the current one is public at /api/market/plate.

What a market state is

One object per market, rebuilt every minute from the data collected up to that minute. It is not a dashboard export and not a summary: it is the same artefact the pages of this site read, the same one an agent receives over MCP, and the same one the model is given when it writes an opinion.

Every state carries, in its own meta: a fixed as_of (the minute it describes, never the minute it was built), a revision, the schema it obeys and the digest of the reading contract that defines its fields. Beside it stands a quality block: whether validation passed, the status of the build and how many published numbers carry a warning.

The body is eleven sections, and they are always the same eleven: price, price structure, order flow, price–flow response, derivatives, positioning, volatility, options, ETF, on-chain and the peer asset. A section that has nothing to say says so and never disappears.

Every state is read against the state of the world of the same minute – rates, global markets, the US session, economic indicators, the macro calendar, the newswire and earnings – a document of its own, kept by its publisher under its publisher's contract. The two travel together wherever the market does, and a kept market snapshot names the world it was read against in world_snapshot_id.

Where the numbers come from

Public exchange feeds and commercial data providers, including on-chain, ETF, macro data and the general financial newswire. The composition is fixed in code, not chosen at runtime: a venue that falls over is a gap in the data and a drop in completeness, never a silent substitution by another venue.

The venues are named, the commercial providers are not. A venue is part of the meaning of a number – "spot flow" says nothing until you know which books it was taken from – so every one of them is named here and in the data itself. Which vendor we buy the daily series from changes nothing about how a number is read, so the middle column carries a dash there. Commercial source identities and the contractual provenance behind them are available to qualified institutional customers under NDA. What a reader needs in order to read a number stands beside the dash: what is taken, and how fresh it is.

layersourcewhat is taken, and how often
Spot order flowCoinbase, Binance public trade streams, read without a break and sealed into one-minute bars on the exchange's own event time – CVD, volume, OHLC and the tick count
Perpetual order flowBybit, Binance Futures trades and liquidations over the stream, open interest and funding on the grid the venue itself publishes them on
Canonical priceCoinbase closed candles – the minute is closed by the stream, the history is backfilled over REST. There is no failover to another venue, because two prices on one screen are one price too many
OptionsDeribit, Bybit the chain in slices through the day and trades over the stream, kept as an hourly history of every expiry
Implied volatilityDeribit DVOL hourly bars of the index, and the running hour kept current
On-chain– daily NUPL and MVRV Z-score, the coins held on exchanges with the flows in and out of them, the supply of the USD-pegged stablecoins and, for ether, the staking queues of the beacon chain, refreshed several times a day. The providers publish a day or two behind the chain, and the whole history arrives with every reading
ETF flows– daily net flow, value traded and net assets of the US spot ETFs of this market, in total and fund by fund, followed through the publication evening in New York so that an issuer's revision lands the same night
Macro– the economic calendar of high-impact US, EU, JP and UK releases, treasury yields, cross-asset returns, economic indicators and earnings. A published figure reaches the market state within seconds of appearing
NewsWSJ, Reuters, CNBC, Market Watch, Bloomberg Markets and Finance, Yahoo Finance, Barrons, NYTimes, 24/7 Wall Street, PYMNTS, The Guardian, New York Post the general financial wire of these outlets, read once a minute. Every story is judged on its own against a fixed rubric for whether it can move large-cap crypto at all, and only what carries weight reaches the market state – a handful a week, not a feed. The outlets are named for the same reason the venues are: who wrote a story is part of reading it. The headline and the abstract are theirs, and every row links back to them

The windows

Every window is anchored on a shared watermark – the last minute finalized by all venues of the scope – and never on "now". That is what makes a comparison between venues honest: both are asked about exactly the same minutes.

familywindowsnotes
Order flow1h, 4h, 24h, 3d, 7d, 14d split by venue on 1h, 4h and 24h, and the equal interval before the window travels with the same three
Open interest1h, 4h, 24h, 3d, 7d, 14d a delta only holds when the set of venues is the same at both ends of the window. The previous equal interval on 4h and 24h
Liquidations1h, 4h, 24h Bybit publishes its full feed, Binance Futures a sampled one, so a liquidation total is a floor, not a census, and it says so
Daily seriesas the source publishes ETF flows, on-chain metrics and macro releases keep the calendar of their publisher. Nothing is resampled to look denser than it is

What quality means here

Two questions are asked in a strict order, and they are different questions.

First, does the period exist at all? A 7-day window on a market four days old is not "a 7-day window at 57%": there is no such week. Until the history reaches the start of the window the metric is warming up – its value is null and it says when it will exist.

Then, how much of it was observed? Inside a period that does exist, completeness is the share of expected source-time cells actually seen. The value published beside it is the sum or the average of what was observed only – nothing is scaled up to pretend the gap was not there.

statewhat it means
livethe value is still forming. It may change, and it exists only in live views and in an interval that has not closed
completethe value was accepted by the validator: a normal feed, a normal poll, a closed bucket of the source, or an exact late backfill – a delayed delivery is a delivery, not a lesser number
unavailablethere is no reliable value: it stays null, the aggregates that depend on it stay null, the screen shows a dash, and the point takes no part in statistics, baselines or autoscaling

Three rules follow from this, and they are the ones worth checking us against:

Revisions

Nothing is rewritten in place. The market as it stands is rebuilt every minute and carries no id. When it matters, it is kept: on every event the world's publisher keeps its world on, and for every question a model is asked, the market and its world are stored for good under an id, with the event that caused them, byte for byte as a reader received them. A kept snapshot never changes, and a new question reads the market as it stands unless a kept snapshot is named.

Providers do revise their own numbers: ETF flows are refined in the evening in New York, and a macro figure can be corrected after its release. Those corrections arrive in the next minute of the market, and never retroactively inside a snapshot that has already been kept.

Method changes are versioned too. Each state names the reading contract its numbers are read by – by the digest of its text – and the schema they are shaped by. When a definition changes, the contract changes with it rather than the meaning of an unchanged field.

How far back the data goes

A stream has no past: it begins the day its collector was started, and nothing can backfill trades nobody was listening to. A REST-backfilled series carries whatever the venue keeps, and a provider that hands over its whole history hands over all of it. So the depth is per source, and the figures below are what stood on the day of this edition.

seriesbeginswhy there
Order flow, open interest, funding, liquidations 6 July 2026 one-minute bars, from the day the collectors were started
Price, daily candles11 August 2024 two years of REST backfill from the venue
Price, minute candlesrolling ~90 days as deep as the venue serves over REST
Implied volatility (DVOL)24 March 2021 the whole history of the index, as the venue keeps it
On-chain (NUPL, MVRV-Z)8 August 2015 the provider returns the full series on every poll
Exchange holdings and flows30 July 2015 the provider returns the full series on every poll
Stablecoin supply29 November 2017 the whole series, as the provider keeps it
Staking queues (ETH)24 September 2026 a reading of the beacon chain from the day this series was started. The chain keeps no history of its queues
ETF flows9 June 2026 from the day this contour was started, and deeper here than at the provider itself
ETF flows by fund25 August 2026 a month back from the day this series was started, as deep as the provider serves a fund
Macro calendar21 July 2026 and roughly a month forward, as the provider schedules
Treasury yields8 July 2026 from the day the macro contour was started

The market states themselves are not kept minute by minute: the latest ones stay, and every state that was actually read by a model is pinned and kept for good – so an opinion can always be checked against the exact data it was written from. What is archived in full is the data underneath, at the depths above.

What this is not

Checking any of it

Changes

The date at the top is the edition. When a source, a window or a rule changes, it changes here and the date moves with it.