FIELD NOTE · A THOUGHT EXPERIMENT ABOUT SPEED AND JUDGMENT
What If AI Could Write
a Million Tokens a Second?
Picture one capable AI agent writing a million output tokens every second. Drafts would arrive almost instantly. Knowing which one to trust would not.
Watch the film, or read the interactive guide. Same idea, your pace.
One Million Tokens a Second
The original animated film, with narration and music. Tap the picture to play. This is a thought experiment, not a measurement of any current model. The reading version below adds two interactive calculators.
“A million a second” can mean four things.
- CONTEXT CAPACITYHow much fits
- The text a model can hold in one request, including its reply. A size, not a speed.
- INPUT PROCESSINGHow fast it reads
- Prompt tokens can be processed largely in parallel. Huge inputs still take time, and the input rate is different from the output rate.
- AGGREGATE THROUGHPUTHow much a system serves
- 10,000 streams × 100 tokens/sec = 1,000,000 tokens/sec in total. Toy arithmetic: it says nothing about when any one stream finishes.
- SINGLE-AGENT OUTPUT · OUR DIALHow fast one agent writes
- One agent writing a million tokens a second. In standard generation, each token depends on the ones before it.
Anthropic’s Claude Opus 5.5 overview, for example, lists a 1M-token context window, a much smaller output cap per request and “moderate” comparative latency. Memory size is not writing speed, and it does not mean a million-token reply fits in one request.
Engineering helps, within limits. Services reach big totals by batching many requests, as the 2023 vLLM paper describes, and speculative decoding can accept several drafted tokens in one pass. Neither removes a genuine chain in which step two needs the result of step one.
A token is a chunk of text, often part of a word, so token counts are not word counts.
Turn the dial.
Pick a workload, then slide the imagined speed from 100 to 1,000,000 tokens per second.
At 1,000,000 tokens/sec, 1,000 story endings (1,000,000 output tokens) take about 1 second to write.
All four budgets at both ends of the dial
| Workload | Output tokens | 100 tokens/sec | 1,000,000 tokens/sec |
|---|---|---|---|
| 1,000 story endings × 1,000 | 1,000,000 | 2 h 46 min 40 s | 1 second |
| 40 app drafts × 20,000 | 800,000 | 2 h 13 min 20 s | 0.8 seconds |
| 10,000 critiqued candidates × 300 | 3,000,000 | 8 h 20 min | 3 seconds |
| 10,000 rehearsals × 1,000 | 10,000,000 | 27 h 46 min 40 s | 10 seconds |
Imagined comparisons, not measurements of any model. Times count generated output only: no reading, ranking, tool calls, permissions, tests, deployment, people or experiments. A budget is a writing allowance, not a finished product.
Parallel streams can accelerate these independent jobs too. But each draft still needs time on its own stream: matching aggregate throughput does not match completion time. One fast agent matters most when each step waits on the last.
Cheap drafts move the hard part.
Each budget counts written output only. None is a finished product, a verified discovery or a real result.
A map of possible endings
1,000 × 1,000 = 1,000,000 tokens · 1 second of writing
Ask for an ending to your story and get a thousand: hopeful, dark, strange. Nobody reads a thousand, so the useful version maps them for you to explore, then blends the two you like.
Still scarce: taste. Only you know which ending is yours.
Software for a neighborhood tool library
40 × 20,000 = 800,000 tokens · 0.8 seconds of writing
Describe a lending app for shared tools and forty drafts exist before you finish the sentence. Screens could adapt, from borrowing a ladder to a repair-day sign-up. Tests still run on their own clock, and if none checks the seven-day due date, all forty can pass while lending ladders for seventy.
Still scarce: a clear spec, and tests of what “working” means.
Ten real experiments from ten thousand ideas
10,000 × 300 = 3,000,000 tokens · 3 seconds of writing
Hunting for a better catalyst? An agent could propose and critique ten thousand candidates. Then everything waits at the bench: reactions may run for hours, cultures for days, field trials for a season. Ideas from one model can also share one blind spot.
Still scarce: physical evidence, and choosing which ten experiments earn lab time.
Rehearsing a hard conversation
10,000 × 1,000 = 10,000,000 tokens · 10 seconds of writing
Before talking to your landlord about the lease, an agent could play the conversation ten thousand ways: stubborn landlord, generous landlord, you when tired. Use it like a flight simulator, for practice and blind spots. Ten thousand rehearsals of the wrong person are a confident mistake, not a prophecy.
Still scarce: fidelity to the real person, and your own practice.
10,000× faster writing is not 10,000× faster work.
Take the forty app drafts: 800,000 output tokens. Compare an illustrative 100 tokens per second with the imagined million, then add one fixed check after writing that speed does not touch, such as a test run.
With 60 seconds of checking, the batch takes 2 h 14 min 20 s at 100 tokens/sec and 60.8 seconds at 1,000,000 tokens/sec. That is 132.57× faster end to end, and checking is 98.7% of the faster total.
The same 800,000 tokens with four check times
| Fixed check | 100 tokens/sec | 1,000,000 tokens/sec | End to end |
|---|---|---|---|
| None | 2 h 13 min 20 s | 0.8 seconds | 10,000× |
| 60 seconds | 2 h 14 min 20 s | 60.8 seconds | 132.57× |
| 15 min | 2 h 28 min 20 s | 15 min 1 s | 9.88× |
| 24 h | 26 h 13 min 20 s | 24 h 1 s | 1.09× |
Toy model: total time = output tokens ÷ speed + one fixed check, counted once for the whole batch (not per draft) after writing ends. It does not simulate parallel tests, cost, energy or draft quality.
With a one-minute check, the job drops from 8,060 to 60.8 seconds: about 132.57 times faster, not 10,000. At zero the full 10,000× returns; at a day the gain nearly vanishes. Whatever you do not speed up becomes nearly all the remaining time, the logic of Amdahl’s law. Real requests have more such steps: OpenAI’s latency guide notes that very large prompts, tool calls and network trips add delays of their own.
When generation gets cheap, judgment gets precious.
Every experiment above ends in the same place. What stays scarce is judgment, wearing different hats:
- TasteWhich option is yours.
- Specs and testsWhat “working” means.
- PracticeSpeed cannot learn it for you.
- Physical evidenceThe world answers at its own pace.
- FidelityA simulation is only as good as its model.
- Cost and energyIs this worth running at all?
Fast does not mean cheap: every token runs on hardware someone pays to power. More options do not guarantee a good one either. Drafts from one model and one set of assumptions can be correlated, or all wrong together. Volume measures output, not understanding.
Before you ask for more, decide how you will choose.
No imaginary dial required. Next time you hand work to an AI agent:
- Write down what “done” means.One testable sentence, like “ladders are due back in seven days,” not “handles loans.”
- Set criteria before reading options.Two or three, so the most fluent draft does not win by default.
- Find the slow step.Name the check that speed will not touch. Shorten or automate it where possible, without skipping the validation each result needs.
- Ask for disagreement, not volume.Request options built on different assumptions, then ask what would make all of them wrong.
If your slow step is deciding what to measure or which bets to make, that is strategy more than tooling. Private consulting helps clarify the larger system and where your next move matters ↗
Sources, dates and boundaries.
This Field Note adapts ideas from echohive’s film One Million Tokens a Second; it is an edited companion, not a transcript. The speed is imagined, and no source below measures, claims or predicts it. They support only the distinctions between capacity, reading, serving and writing.
Open the four sources
- Anthropic, Claude Opus 5.5 overview. Official documentation. Lists a 1M-token context window, a separate maximum output per request and “moderate” comparative latency. It does not describe a million output tokens per second.
- Anthropic, Context windows. Official documentation. Describes the context window as working memory that includes the generated response: a capacity, not a speed.
- OpenAI, Latency optimization. Official API guide. Output generation commonly dominates latency; very large prompts still matter, and tool or network calls add their own delays.
- Agrawal et al., Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve. arXiv, 2024. A historical foundation, not a current benchmark: parallel prompt processing (prefill), token-by-token decoding and batching.
The 2023 vLLM paper linked in section 01 is also a historical foundation. All workloads, speeds, budgets and the bottleneck model are illustrative arithmetic, not benchmarks, forecasts or product specifications. Sources checked October 2, 2026.
KEEP THE CURIOSITY. STRENGTHEN THE PRACTICE.
Understand the shift.
Practice the judgment.
Learn at your own pace, think it through with others on Sundays, or get help with your own direction. The Architect membership includes the full Get Amplified collection plus the 1000x Lab: mostly live discussion with replays, not a coding class. Private consulting is separate and not included.
Ready to make this a practice? Compare the Get Amplified and 1000x Lab options on Patreon, where the current tier details are explained.
Explore membership on Patreon ↗














