💻 Tom Tunguz: Birds Don't Fly Like Planes — Why Local AI Models Win on Time-to-Answer, Not Tokens/Sec
Benchmarking DeepSeek V4 against two local Ollama models on 25 VC tasks, Tunguz finds token speed is the wrong metric — time to answer is what matters.
• A local Qwen3.6-35B model generates 2.2x faster than its Qwen3.8-27B peer, yet finishes slower because it "thinks" 3.1x longer — quality across all three models ties at 8/9.
• Smaller local models reason more to close the knowledge gap with bigger cloud models, like a bumblebee taking a longer flight path to the same destination.
• Takeaway: your laptop can now run a model as capable as most cloud options — just budget for latency, not throughput.
🔗 Read more
#tomtunguz