Next token prediction & hallucination — hands-on lab

Next Token Lab

Three linked stations. Each one writes a prompt for you to run in any AI tool — this page doesn't call a model itself, it's the instrument that shapes what you ask it and helps you read what comes back.

The four ideas this lab is built on

Capability zone vs. limitation zone

Capability zone: tasks that resemble patterns seen many times — summarizing, reformatting, explaining common concepts. Limitation zone: novel or sparse territory, and anywhere the task requires telling "true" apart from "sounds true."

Fabrication concentrates in specificity

Names, dates, statistics, citations, URLs, quotes. The more precise a claim, the more it warrants verification — this is what Station 02's specificity dial is built to demonstrate.

What pushes the limitation further out

Citations, uncertainty signaling, constrained generation, and generator-verifier loops — the four product-side checks Station 02's highlighter and Station 03's refinement prompt are both built around.

4D connection — Discernment

Next-token prediction is the foundation of Discernment: knowing an output was generated word-by-word from likelihood, not looked up or verified, tells you exactly what kind of scrutiny to apply before you trust it.

01

Next-token predictor

Type a seed phrase — try "Mary had…" — pick which generation algorithm to simulate, and generate a prompt that asks an AI to rank its most likely next-word continuations with estimated probabilities as that algorithm would produce them. This is the mechanism from the lesson made visible: one continuation dominating, or the field staying wide open — and how much context the algorithm actually used to get there.

Algorithm — a century of next-word prediction

Seed text


02

Confidence & specificity blog lab

Give a topic, then set how confidently and how specifically the blog should be written — specificity is the dial that matters most: fabrication concentrates in precise claims (names, dates, statistics, citations), so the more precise a claim, the more it warrants verification. Generate the prompt, run it in an AI tool, then paste the result back in below to see citations, hedging, absolute language, and specific claims highlighted — the same product-side checks (citations, uncertainty signaling, constrained generation, generator-verifier loops) that exist to push this limitation further out.

Topic + dials

Confidence tone 5/10

Low = hedged and cautious. High = stated as flat, certain fact regardless of whether it's actually known.

Specificity 5/10

Low = general, safe statements. High = named studies, exact dates, precise statistics — where fabrication concentrates.

Paste the blog the AI wrote back here to inspect it

Source / citation language Hedged / uncertainty language Absolute / overconfident language Specific claim (verify this)
This is a keyword-pattern scan, not a fact-checker — it flags the shape of the language, not whether any given claim is true. A wall of "specific claim" highlights with no "source" highlights nearby is the pattern worth noticing.

03

Blog refinement prompt builder

Paste any blog (text only — leave out images) and generate a refinement prompt built from the six "what can you do" checks: confident tone isn't an accuracy signal, specificity is where fabrication concentrates, treat outputs as drafts, place the task in the capability or limitation zone, lean on the four product-side checks, and remember the model can't tell grounded from invented — this is the Discernment step: knowing it was generated tells you exactly what scrutiny to apply.

Blog to refine