Gemini vs OpenAI vs Claude, tokens per second and latency measured in September 2026

96 streaming runs across 16 current models on one morning, six per model, with time to first token, output tokens per second and the reasoning tokens that explain most of the wait.

All 96 runs were made between 01:30 and 01:44 UTC on 6 September 2026, sent through a Mullvad WireGuard VPN exit in Warsaw, Poland (AS212238), so every network hop to each provider includes the VPN leg. API speed moves hour to hour, so read each figure as a dated snapshot, not a permanent ranking.

GPT-5.4 Nano returned its first token in a median 2.00 seconds and Gemini 3.8 Flash in 2.03 seconds, the two fastest of 16 models timed through one API on 6 September 2026, while Gemini 3.5 Flash Lite produced the fastest output at a median 389 tokens per second, about 1.5 times the next model. The surprise is that first token time tracks reasoning tokens more than vendor. Grok 4.6 spent a median 504 tokens thinking and 12.02 seconds before its first word, and two Gemini models burned about 1,440 of a 1,500 token budget on reasoning at low effort, leaving under 60 tokens of visible answer.

2.00 sfastest median first token, GPT-5.4 Nano
389 tok/sfastest output, Gemini 3.5 Flash Lite
12.02 sslowest first token, Grok 4.6
r = 0.73reasoning tokens vs first token time

Every number here comes from 96 streaming requests sent from one machine, six per model, interleaved so no vendor got a quiet minute the others did not.

All 16 models sorted by time to first token

Median of six streaming runs per model through OpenRouter default provider routing, 6 September 2026. Prompt "Explain in about 200 words how TCP congestion control works. Plain prose, no lists.", max_tokens 1500, temperature 0, reasoning effort low. TTFT and total are seconds. Output price is USD per 1M output tokens from the OpenRouter models feed read at 01:29 UTC on 6 September 2026 (DeepSeek prices on that feed are exchange rate linked and drift within the day); GPT-5.6 Sol uses the official OpenAI Standard rate.

ModelRouted toRuns keptMedian TTFT s (min to max)Median output tok/s (min to max)Median reasoning tokensMedian total sOutput $ per 1M
GPT-5.4 NanoOpenAI6 of 62.00 (1.67 to 3.52)199 (134 to 216)03.48$1.25
Gemini 3.8 FlashGoogle6 of 62.03 (1.92 to 2.44)133 (92 to 170)03.82$3.75
GPT-5.4 MiniOpenAI6 of 62.04 (1.92 to 2.13)148 (140 to 184)133.80$4.50
Claude Sonnet 5Claude Platform on AWS6 of 62.35 (2.12 to 2.71)88 (86 to 96)06.92$10.00
Kimi K3Sail Research6 of 62.38 (1.99 to 2.96)82 (72 to 118)145.75$15.00
GPT-6 AstraOpenAI6 of 62.41 (2.16 to 2.89)53 (41 to 61)07.35$50.00
Claude Haiku 4.5Amazon Bedrock6 of 62.83 (2.58 to 3.11)98 (91 to 102)1155.78$5.00
Gemini 3.5 Flash LiteGoogle5 of 63.83 (3.58 to 8.41)389 (161 to 402)4764.64$2.50
Claude Fable 5.1Anthropic6 of 63.88 (3.56 to 5.00)65 (62 to 67)010.49$50.00
DeepSeek V4 FlashDigitalOcean (2), GMICloud (1), Novita (1), Parasail (1), Baidu (1)6 of 63.89 (2.32 to 9.01)77 (9 to 156)746.50$0.16
GPT-5.6 SolOpenAI6 of 64.07 (3.18 to 6.83)45 (43 to 78)4410.26$20.00 (official; OpenRouter lists $10)
Claude Opus 5Claude Platform on AWS6 of 64.66 (4.35 to 4.88)89 (87 to 109)249.45$25.00
DeepSeek V4 ProSiliconFlow (1), StreamLake (5)6 of 65.29 (4.44 to 7.52)50 (37 to 62)12511.08$1.56
Gemini 3.5 FlashGoogle4 of 66.75 (5.99 to 8.95)268 (248 to 293)8057.59$9.00
Gemini 3.1 Pro PreviewGoogle3 of 67.92 (7.81 to 7.92)186 (184 to 197)5879.33$12.00
Grok 4.6xAI6 of 612.02 (10.37 to 14.37)53 (47 to 93)50416.56$6.00

Six runs were dropped by one rule, applied to every model alike. A run counts only if it produced at least 80 content tokens and generated for more than 0.5 seconds. Two Gemini 3.5 Flash runs and three Gemini 3.1 Pro Preview runs hit the 1,500 token cap after 1,437 to 1,440 reasoning tokens, then returned 56 to 59 content tokens (about 340 characters) before being cut off. That is a real finding, not a glitch.

At reasoning effort low, on this prompt, those two models sometimes spend the whole output budget on reasoning tokens, and a caller with max_tokens set for the answer alone gets a truncated reply. The sixth exclusion is a Gemini 3.5 Flash Lite run that returned 198 tokens in 0.489 seconds, a hair under the half second floor. Had it counted, its output rate would have been 405 tokens per second, above the model's kept median.

On the price column, OpenRouter lists openai/gpt-5.6-sol at $2 input and $10 output per 1M tokens while the official OpenAI pricing page lists $4 and $20 for Standard processing; the OpenRouter figure equals OpenAI's Batch and Flex rate. The table uses the official $20.

Gemini vs OpenAI vs Claude, head to head

Best and worst model per vendor from the same 6 September 2026 run. Medians of the kept runs.

VendorFastest first tokenFastest outputSlowest first tokenModels timed
GoogleGemini 3.8 Flash, 2.03 sGemini 3.5 Flash Lite, 389 tok/sGemini 3.1 Pro Preview, 7.92 s4
OpenAIGPT-5.4 Nano, 2.00 sGPT-5.4 Nano, 199 tok/sGPT-5.6 Sol, 4.07 s4
AnthropicClaude Sonnet 5, 2.35 sClaude Haiku 4.5, 98 tok/sClaude Opus 5, 4.66 s4
DeepSeekDeepSeek V4 Flash, 3.89 sDeepSeek V4 Flash, 77 tok/sDeepSeek V4 Pro, 5.29 s2
MoonshotKimi K3, 2.38 sKimi K3, 82 tok/sKimi K3, 2.38 s1
xAIGrok 4.6, 12.02 sGrok 4.6, 53 tok/sGrok 4.6, 12.02 s1

On first token, Google and OpenAI tie at the small end. GPT-5.4 Nano at 2.00 seconds and Gemini 3.8 Flash at 2.03 seconds are 30 milliseconds apart, well inside the run to run spread of either. On output speed the two Gemini 3.5 Flash models are in a different league, 389 and 268 tokens per second against 199 for Nano, but both made the reader wait 3.83 and 6.75 seconds for the first word, and Gemini 3.8 Flash at 133 tokens per second was slower than both Nano and GPT-5.4 Mini at 148.

Anthropic's fastest first token, Sonnet 5 at 2.35 seconds, sits a third of a second behind the leaders, and its decode speeds cluster between 65 and 98 tokens per second. The Claude flagships beat the OpenAI flagships on decode, Opus 5 at 89 and Fable 5.1 at 65 against GPT-6 Astra at 53 and GPT-5.6 Sol at 45, and Astra beat Fable to the first token, 2.41 seconds against 3.88.

What the numbers show

Reasoning tokens explain about half of the first token variance

The five models whose median reasoning count was zero (GPT-5.4 Nano, Gemini 3.8 Flash, Claude Sonnet 5, GPT-6 Astra, Claude Fable 5.1) had a median first token of 2.35 seconds. The four models with 400 or more median reasoning tokens (Gemini 3.5 Flash Lite, Gemini 3.5 Flash, Gemini 3.1 Pro Preview, Grok 4.6) had a median of 7.34 seconds. Across all 16 models the correlation between median reasoning tokens and median TTFT is r 0.73 (r squared 0.53), so reasoning explains about half of the variance in first token time and this run leaves the other half unexplained.

Grok 4.6 is the extreme case, 504 reasoning tokens and 12.02 seconds, six times Nano's wait. Gemini 3.5 Flash follows at 805 tokens and 6.75 seconds. The request passed reasoning effort low to OpenRouter, and the mapping of that setting to each vendor's own parameter is OpenRouter's; on this prompt the result ranged from zero reasoning tokens to the entire budget.

Gemini Flash decodes fastest but starts late

Gemini 3.5 Flash Lite streamed at 389 tokens per second and Gemini 3.5 Flash at 268, the two highest rates in the run. Nothing else broke 200. Yet the model that finished a 200 word answer soonest inside Google's own lineup was Gemini 3.8 Flash, which reasoned for zero tokens, decoded at a slower 133 tokens per second, and was done in a median 3.82 seconds. Gemini 3.5 Flash needed 7.59 seconds for the same job, 3.77 seconds longer, because its head start on reasoning ate the decode advantage twice over. For a reply of this length, first token time decides who feels fast.

Provider routing moved DeepSeek V4 Flash more than any model choice

DeepSeek V4 Flash was routed to five different hosts in six runs. Baidu answered in 2.32 seconds and streamed at 156 tokens per second. DigitalOcean, which took two of the six, answered in 8.07 and 9.01 seconds and streamed at 11.4 and 8.7 tokens per second, and its slower run took 37.47 seconds end to end against 4.07 seconds for the Baidu run.

That is a 17.9 times spread in output rate inside one model, wider than the 8.6 times spread between the fastest and slowest model medians across the whole table. DeepSeek V4 Pro went to StreamLake five times and SiliconFlow once and was steadier, 36.8 to 62.1 tokens per second. Buy an open weight model through a router without pinning the provider and the median tells you less than the range does.

GPT-6 Astra and GPT-5.6 Sol decoded at about a quarter of Nano's rate

GPT-6 Astra had a first token at 2.41 seconds, sixth fastest, then streamed at 53 tokens per second, 3.7 times slower than GPT-5.4 Nano, and needed 7.35 seconds for about 250 tokens. GPT-5.6 Sol was slower on both counts at 4.07 seconds and 45 tokens per second, the lowest decode rate in the table. Other large models sat nearby, DeepSeek V4 Pro at 50, Grok 4.6 at 53 and Claude Fable 5.1 at 65, while Claude Opus 5 reached 89. At $50 per 1M output tokens, the 429 token Fable answer cost about $0.0215, roughly 70 times the $0.0003 that GPT-5.4 Nano charged for its 244 tokens, and it took three times as long to arrive.

Claude at low effort mostly skipped reasoning

Claude Fable 5.1 and Claude Sonnet 5 emitted zero reasoning tokens in all six runs each. Claude Opus 5 emitted 22 to 24 per run and Claude Haiku 4.5 between 67 and 136. Fable's first token still came at a median 3.88 seconds with no reasoning tokens at all, and Opus's at 4.66. In exchange the two Claude flagships wrote the longest answers of the run, medians of 429 and 441 content tokens against 404 for Sonnet 5 and 218 to 282 for the other 13, which is why their total times run past 9 seconds despite decent decode rates.

What this does and does not measure

One 14 word prompt asking for a 200 word explanation of TCP congestion control, sent 96 times; providers counted it as anywhere from 20 to 224 prompt tokens with their own tokenizers. Each request used the OpenRouter chat completions endpoint with stream true, usage included, max_tokens 1500, temperature 0 and reasoning effort low, from Python urllib with no SDK. Sixteen models were cycled six times in a fixed order, so consecutive runs of the same model were separated by about two minutes.

Time to first token is wall time from the moment the request was sent to the first content delta, so it includes DNS, TLS, OpenRouter's own routing hop and the provider's queue. Output tokens per second is content tokens (completion tokens minus reasoning tokens, both from the usage object OpenRouter returns at the end of the stream) divided by the time between the first and last content delta.

What it does not measure. A single egress point; every request left through a Mullvad WireGuard VPN exit in Warsaw, Poland (AS212238), so the network hops to each provider include the VPN leg. One time window, 01:30 to 01:44 UTC on a Sunday. Provider routing was left on OpenRouter's default, not pinned, so DeepSeek numbers in particular describe the router's choice on the day. Six runs per model is enough to show the range, not enough for a confidence interval, which is why every cell prints min to max. Reasoning effort low was passed to OpenRouter, and the mapping of that setting to each vendor's own parameter is OpenRouter's.

A 12 run pilot at max_tokens 300 (published at runs-v1-maxtok300.jsonl, with the measuring script and aggregate summary) had five of twelve models stop at exactly 300 completion tokens with reasoning models exhausting the budget, which is why the reported run used 1,500; the pilot is not in the tables. Direct vendor endpoints will likely differ from these figures by the router overhead and by region. The older page AI API latency comparison carries April 2026 figures for earlier models, including Gemini 2.5 Pro and Flash, and was measured with a different method, so do not read the two pages as one time series.

Raw data

All 96 runs, including the six excluded ones, are embedded in this page's source as a JSON array inside the script element with id runs, one object per request with provider, timestamps, usage, reasoning tokens, content tokens and the derived rate, so anyone can recompute every median above.

Questions this run can answer

Which AI API has the lowest latency right now?

In this run, GPT-5.4 Nano at a median 2.00 seconds to first token, with Gemini 3.8 Flash at 2.03 and GPT-5.4 Mini at 2.04. All three emitted zero or near zero reasoning tokens. The slowest was Grok 4.6 at 12.02 seconds, which reasoned for a median 504 tokens first. Figures include OpenRouter overhead and the VPN leg from Warsaw.

Is Gemini faster than OpenAI in tokens per second?

For two of the four Gemini models, yes. Gemini 3.5 Flash Lite decoded at 389 tokens per second and Gemini 3.5 Flash at 268, against 199 for GPT-5.4 Nano, OpenAI's fastest here. On time to first token the two vendors tied at the small end, 2.03 seconds against 2.00, but the two fast decoding Gemini models waited 3.83 and 6.75 seconds before their first word. Gemini 3.8 Flash, at 133 tokens per second, was slower than Nano.

What is Claude API latency in September 2026?

Median time to first token was 2.35 seconds for Claude Sonnet 5, 2.83 for Claude Haiku 4.5, 3.88 for Claude Fable 5.1 and 4.66 for Claude Opus 5. Output speeds were 88, 98, 65 and 89 tokens per second in the same order. Fable 5.1 and Sonnet 5 emitted zero reasoning tokens in every run at low effort. Current prices for the flagship are on the Claude Fable 5.1 pricing page.

What is the Gemini Pro time to first token?

We measured Gemini 3.1 Pro Preview, not Gemini 2.5 Pro. Its median time to first token was 7.92 seconds across the three runs that produced a full answer, with a median 587 reasoning tokens. The other three runs spent about 1,440 of the 1,500 token budget on reasoning and returned under 60 tokens of answer. The April 2026 page linked above has the older Pro figures.

Why were six of the 96 runs excluded?

A run counts only if it produced at least 80 content tokens and generated for more than 0.5 seconds. Five runs of two Gemini models ran out of token budget while reasoning and returned 56 to 59 content tokens, and one Gemini 3.5 Flash Lite run returned 198 tokens in 0.489 seconds, just under the floor. All six are in the embedded data with their timings.

KickLLM Margin Studio · Offline analysis app

Take the benchmark into your own workload.

Review measured cost, acceptance and latency from your own trials. Use Margin Studio to compare the evidence, model contribution margins and export an explainable report.

Get Margin Studio — $39 Try the interactive preview → One-time purchase · Downloadable ZIP
By the same builder: GitHub · theluckystrike BeLikeNative · Grammar AI EarlyThunder · Dev Blog Bug Bounty Reality Zovo · AI Dev Tools