Computers are faster than you think

Computers are incredibly fast, and you can almost always make your code faster. When optimizing, you should be able to keep optimizing until you’re hitting these limits. Analyze your code and figure out what the fundamental limits are.

Handy references

Latency Comparison Numbers (~2012)
----------------------------------
L1 cache reference                           0.5 ns
Branch mispredict                            5   ns
L2 cache reference                           7   ns                      14x L1 cache
Mutex lock/unlock                           25   ns
Main memory reference                      100   ns                      20x L2 cache, 200x L1 cache
Compress 1K bytes with Zippy             3,000   ns        3 us
Send 1K bytes over 1 Gbps network       10,000   ns       10 us
Read 4K randomly from SSD*             150,000   ns      150 us          ~1GB/sec SSD
Read 1 MB sequentially from memory     250,000   ns      250 us
Round trip within same datacenter      500,000   ns      500 us
Read 1 MB sequentially from SSD*     1,000,000   ns    1,000 us    1 ms  ~1GB/sec SSD, 4X memory
Postgres indexed read, local         2,000,000   ns    2,000 us    2 ms  `SELECT ... WHERE id = ?`
Postgres indexed read, same DC       4,000,000   ns    4,000 us    4 ms  App server calls DB server
Postgres write, durable row          6,000,000   ns    6,000 us    6 ms  `INSERT` one row and commit
HDD disk seek                       10,000,000   ns   10,000 us   10 ms  20x datacenter roundtrip
Local LLM, generate 1 token         15,000,000   ns   15,000 us   15 ms  Small model on consumer GPU (2026)
Read 1 MB sequentially from disk    20,000,000   ns   20,000 us   20 ms  80x memory, 20X SSD
Frontier LLM, generate 1 token      20,000,000   ns   20,000 us   20 ms  Hosted model output (2026)
Local LLM, time to first token      75,000,000   ns   75,000 us   75 ms  Small model, short prompt (2026)
Local LLM (CPU), generate 1 token  100,000,000   ns  100,000 us  100 ms  Small model, no GPU (2026)
Send packet CA->Netherlands->CA    150,000,000   ns  150,000 us  150 ms
Fast LLM, time to first token      250,000,000   ns  250,000 us  250 ms  Specialized inference hardware (2026)
Frontier LLM, time to first token                              1,000 ms  Short prompt, no cache (2026)
Frontier LLM, short response                                   3,000 ms  ~100 output tokens (2026)
Frontier LLM, long context prefill                            10,000 ms  ~100K input tokens, no cache (2026)
Frontier LLM, reasoning response                              30,000 ms  Single call with thinking (2026)

Credit
------
Stolen (minor edits) from:  https://gist.github.com/jboner/2841832
By Jeff Dean:               http://research.google.com/people/jeff/
Originally by Peter Norvig: http://norvig.com/21-days.html#answers

Always benchmark

You can’t optimize what you can’t measure. Establish a baseline before you touch any code, then re-run the benchmark each time you make changes. Always write these notes down, ideally in a markdown file next to the code so that it’s easy to reference.

End-to-end benchmarks are a must. You MUST have end to end benchmarks. Inner benchmarks can also be useful, especially at the beginning when you’re figuring out your initial approach, but be wary of the work just moving from one place to another rather than actually getting faster.