π» Tom Tunguz: Birds Don't Fly Like Planes β Why Local AI Models Win on Time-to-Answer, Not Tokens/Sec
Benchmarking DeepSeek V4 against two local Ollama models on 25 VC tasks, Tunguz finds token speed is the wrong metric β time to answer is what matters.
β’ A local Qwen3.6-35B model generates 2.2x faster than its Qwen3.8-27B peer, yet finishes slower because it "thinks" 3.1x longer β quality across all three models ties at 8/9.
β’ Smaller local models reason more to close the knowledge gap with bigger cloud models, like a bumblebee taking a longer flight path to the same destination.
β’ Takeaway: your laptop can now run a model as capable as most cloud options β just budget for latency, not throughput.
π Read more
#tomtunguz