Computers are faster than you think
Computers are incredibly fast, and you can almost always make your code faster. When optimizing, you should be able to keep optimizing until you’re hitting these limits. Analyze your code and figure out what the fundamental limits are.
Handy references
Latency Comparison Numbers (~2012)
----------------------------------
L1 cache reference 0.5 ns
Branch mispredict 5 ns
L2 cache reference 7 ns 14x L1 cache
Mutex lock/unlock 25 ns
Main memory reference 100 ns 20x L2 cache, 200x L1 cache
Compress 1K bytes with Zippy 3,000 ns 3 us
Send 1K bytes over 1 Gbps network 10,000 ns 10 us
Read 4K randomly from SSD* 150,000 ns 150 us ~1GB/sec SSD
Read 1 MB sequentially from memory 250,000 ns 250 us
Round trip within same datacenter 500,000 ns 500 us
Read 1 MB sequentially from SSD* 1,000,000 ns 1,000 us 1 ms ~1GB/sec SSD, 4X memory
Postgres indexed read, local 2,000,000 ns 2,000 us 2 ms `SELECT ... WHERE id = ?`
Postgres indexed read, same DC 4,000,000 ns 4,000 us 4 ms App server calls DB server
Postgres write, durable row 6,000,000 ns 6,000 us 6 ms `INSERT` one row and commit
HDD disk seek 10,000,000 ns 10,000 us 10 ms 20x datacenter roundtrip
Local LLM, generate 1 token 15,000,000 ns 15,000 us 15 ms Small model on consumer GPU (2026)
Read 1 MB sequentially from disk 20,000,000 ns 20,000 us 20 ms 80x memory, 20X SSD
Frontier LLM, generate 1 token 20,000,000 ns 20,000 us 20 ms Hosted model output (2026)
Local LLM, time to first token 75,000,000 ns 75,000 us 75 ms Small model, short prompt (2026)
Local LLM (CPU), generate 1 token 100,000,000 ns 100,000 us 100 ms Small model, no GPU (2026)
Send packet CA->Netherlands->CA 150,000,000 ns 150,000 us 150 ms
Fast LLM, time to first token 250,000,000 ns 250,000 us 250 ms Specialized inference hardware (2026)
Frontier LLM, time to first token 1,000 ms Short prompt, no cache (2026)
Frontier LLM, short response 3,000 ms ~100 output tokens (2026)
Frontier LLM, long context prefill 10,000 ms ~100K input tokens, no cache (2026)
Frontier LLM, reasoning response 30,000 ms Single call with thinking (2026)
Credit
------
Stolen (minor edits) from: https://gist.github.com/jboner/2841832
By Jeff Dean: http://research.google.com/people/jeff/
Originally by Peter Norvig: http://norvig.com/21-days.html#answers
Always benchmark
You can’t optimize what you can’t measure. Establish a baseline before you touch any code, then re-run the benchmark each time you make changes. Always write these notes down, ideally in a markdown file next to the code so that it’s easy to reference.
End-to-end benchmarks are a must. You MUST have end to end benchmarks. Inner benchmarks can also be useful, especially at the beginning when you’re figuring out your initial approach, but be wary of the work just moving from one place to another rather than actually getting faster.