← All research
Tilting proof-of-work toward the CPU
Decentralisation is a hardware policy. We sized the proof-of-work so its working set lives in CPU cache — turning a GPU's throughput advantage into a memory-latency penalty. Measured on our hardware: 26–40× faster on a CPU than a GPU.
Who can mine decides who controls a chain. If a proof-of-work rewards raw parallel throughput, GPUs win, then ASICs win, and a handful of foundries own the network. TSN's goal is the opposite: keep block production on the commodity CPUs that ordinary people already own. That is not a slogan you can wish into being — it has to be built into the hash function.
The lever: memory latency, not compute
A GPU is thousands of arithmetic units fed by a wide but high-latency memory system. It is magnificent at problems that are large, parallel, and predictable in their memory access. It is poor at problems that are small, serial-ish, and random in their access, because then the arithmetic units spend their lives stalled, waiting on memory.
A modern CPU core is the mirror image: fewer units, but a deep, low-latency cache hierarchy sitting inches from the ALU. If the entire working set of the hash fits in that fast cache and every step chases a pointer to an unpredictable place inside it, the CPU runs at full speed while the GPU chokes. So we made the proof-of-work cache-resident: the working set is tuned to fit a CPU's fast cache, and the access pattern is latency-bound and data-dependent by construction.
The measured result
On the hardware we tested, the design lands the intended inversion: a CPU core out-hashes a GPU by a factor between 26× and 40×, depending on the specific CPU and card. The range is honest — it is not one number, because it depends on cache size and memory system, which is exactly the point: the advantage tracks the CPU's cache, not its clock.
| Workload shape | Favours | CPU / GPU |
|---|---|---|
| Large, parallel, predictable | GPU | ≪ 1× |
| Cache-resident, random, latency-bound | CPU | 26–40× |
A related hardening
The same rework that crossed the CPU/GPU line also tightened an adjacent surface: the number of work attestations a block may carry is now capped at 64. Uncapped, that field is an amplification lever; bounded, it is a predictable, verifiable quantity. Small change, and it belongs to the same "make the cost structure legible" family as the cache-residency itself.
Honesty: this is a truce, not a victory
Cache-hardness is not a law of physics; it is a bet about today's memory hierarchies. A future device — a GPU with a radically different on-package memory, or an ASIC purpose-built around this exact working-set size — could narrow the gap. That is the permanent condition of anti-specialisation PoW: you are buying time and raising the cost of centralisation, not ending it. The measured 26–40× is real and it is deployed; we treat it as a lead to defend, not a moat to retire behind.
Filed under engineering · deployed at the cache-hard hardfork. Numbers are from our own benches; your ratio will move with your cache.