A processor fortified as a castle of cache memory ← All research
Engineering

Tilting proof-of-work toward the CPU

Decentralisation is a hardware policy. We sized the proof-of-work so its working set lives in CPU cache — turning a GPU's throughput advantage into a memory-latency penalty. Measured on our hardware: 26–40× faster on a CPU than a GPU.

10 min TSN core Deployed

Who can mine decides who controls a chain. If a proof-of-work rewards raw parallel throughput, GPUs win, then ASICs win, and a handful of foundries own the network. TSN's goal is the opposite: keep block production on the commodity CPUs that ordinary people already own. That is not a slogan you can wish into being — it has to be built into the hash function.

The lever: memory latency, not compute

A GPU is thousands of arithmetic units fed by a wide but high-latency memory system. It is magnificent at problems that are large, parallel, and predictable in their memory access. It is poor at problems that are small, serial-ish, and random in their access, because then the arithmetic units spend their lives stalled, waiting on memory.

A modern CPU core is the mirror image: fewer units, but a deep, low-latency cache hierarchy sitting inches from the ALU. If the entire working set of the hash fits in that fast cache and every step chases a pointer to an unpredictable place inside it, the CPU runs at full speed while the GPU chokes. So we made the proof-of-work cache-resident: the working set is tuned to fit a CPU's fast cache, and the access pattern is latency-bound and data-dependent by construction.

CPU CORE ALU L2/L3 working set ~1–4 ns hit · full speed GPU thousands of cores global memory (far) ~200–400 ns miss · cores stall the set doesn't fit the GPU's fast path, so throughput idles.
Fig 1. When the working set is cache-sized and access is random, the GPU's parallelism can't hide its memory latency. The CPU's cache does.

The measured result

On the hardware we tested, the design lands the intended inversion: a CPU core out-hashes a GPU by a factor between 26× and 40×, depending on the specific CPU and card. The range is honest — it is not one number, because it depends on cache size and memory system, which is exactly the point: the advantage tracks the CPU's cache, not its clock.

Workload shapeFavoursCPU / GPU
Large, parallel, predictableGPU≪ 1×
Cache-resident, random, latency-boundCPU26–40×
GPU 1× (baseline) CPU 26× — 40× relative hashrate on the cache-hard PoW (higher = better for decentralisation)
Fig 2. The inversion, drawn to scale. The bar is a range, not a point, because the edge is a property of the CPU's cache.

A related hardening

The same rework that crossed the CPU/GPU line also tightened an adjacent surface: the number of work attestations a block may carry is now capped at 64. Uncapped, that field is an amplification lever; bounded, it is a predictable, verifiable quantity. Small change, and it belongs to the same "make the cost structure legible" family as the cache-residency itself.

Honesty: this is a truce, not a victory

Cache-hardness is not a law of physics; it is a bet about today's memory hierarchies. A future device — a GPU with a radically different on-package memory, or an ASIC purpose-built around this exact working-set size — could narrow the gap. That is the permanent condition of anti-specialisation PoW: you are buying time and raising the cost of centralisation, not ending it. The measured 26–40× is real and it is deployed; we treat it as a lead to defend, not a moat to retire behind.

Filed under engineering · deployed at the cache-hard hardfork. Numbers are from our own benches; your ratio will move with your cache.