Course 4 capstone · Next Token Prediction × Knowledge × Working Memory × Steerability

Calibration Map

This isn't a blog-writing tool. It doesn't make AI more accurate — it tells you where to focus your verification effort, based on your task's risk profile.

Why this exists, in plain terms: AI answers don't come with a built-in warning label for how much to trust them. A well-written, confident-sounding answer can be completely reliable, or completely wrong, and the writing style doesn't tell you which. What actually determines the risk is four separate things about your task: how common the request is, how well-known the topic is, how much context the task needs to hold onto, and how precisely you specified what you want. This tool locates your task on all four, in one place, and shows you exactly what changes when you compensate for each one.

Input: a description of your task, four quick honest self-assessments (this tool can't know these on its own — only you can), and optionally your actual draft instruction.
Outcome: a visual map of where your task sits on all four risk spectrums, the specific habit to apply for each, and — if you want it — a real side-by-side test: your task run as-is versus run with those habits actually applied.
Useful for: deciding, before you rely on an AI's answer, how much verification effort the task actually calls for — a quick factual lookup on a common topic needs almost none; a long, ambiguous, niche, context-heavy task needs a lot.
See a worked example first (photosynthesis blog, naive vs. calibrated)

This is what the tool produces once you've run both prompts externally and pasted the results back in — shown here with a topic already filled in, so you know what to expect before you start.

Task & settings used

Topic: "How photosynthesis works" for high-school readers · Next Token Prediction: novel/sparse (75%) · Knowledge: niche (70%) · Working Memory: long context (60%) · Steerability: well-structured instructions (40%, low risk)

Results, side by side (markdown already stripped)

Naive prompt result

Photosynthesis is the process by which plants convert light energy into chemical energy stored in glucose. It occurs mainly in the leaves and involves two stages: the light-dependent reactions and the light-independent reactions (Calvin cycle). In the light-dependent reactions, photons strike chlorophyll in the thylakoid membranes, exciting electrons and driving a proton gradient that powers ATP synthase. Photolysis of water releases oxygen while electrons pass to NADP+, forming NADPH. The Calvin cycle then uses that ATP and NADPH to fix CO2 into glucose in the stroma. Overall: 6CO2 + 6H2O + light → C6H12O6 + 6O2.

Calibrated prompt result

Part 1 of 3 — The plant's solar panel: sunlight hits chlorophyll inside a leaf's chloroplasts, and an electron jumps to a higher energy state — the first step of capturing that energy. Part 2 of 3 — Splitting water: the plant breaks H2O apart, releasing oxygen (the air we breathe) and using the hydrogen to charge two "batteries," ATP and NADPH. Part 3 of 3 — Building sugar from air: those batteries power the Calvin cycle, stitching CO2 into glucose in the stroma — a bean plant growing taller is literally made from air, water, and sunlight.

What actually changed (word-level diff)

Photosynthesis is the process by which plants convert light energy into chemical energy stored in glucose, described through the electron transport chain and proton gradients energy step by step — starting with an excited electron, then water-splitting, then sugar-building. Broken into three labeled parts so a reader can follow one idea at a time instead of one dense paragraph of biochemistry terms.

The comparative-analysis prompt this tool hands you

Compare these two blog posts on photosynthesis for a high-school audience: [paste Naive result] [paste Calibrated result] Assess: (1) clarity for the target audience, (2) whether one scaffolds understanding step-by-step better, (3) which is more likely to hold attention, (4) whether either sacrifices accuracy for simplicity, (5) structure and pacing. Give specific examples for each point, then say which you'd recommend for a high-school class and why.

Nothing here is faked — the tool only ever builds the two prompts and diffs the two real results you paste back. Judging which is "better" is still your call, or the analysis prompt's, run for real.

This tool never claims to write a better answer itself, and it never scores your task numerically. Three of the four placements below are your own honest self-assessment — no tool can know how novel your topic is or how much context you're really supplying except you. The only mechanically-checked one is the fourth, if you paste in your actual instruction text.
1 — Next Token Prediction: is this well-worn territory?
2 — Knowledge: how mainstream is the topic?
3 — Working Memory: how much context does this need?