← All projects Project / 04

Tooling · 2026

scriptgen

Study the style of the creators you love, then write and assemble video projects.

Interactive prototype · guided walkthrough

The redesigned ScriptGen app, driven by a scripted tour — each step first explains what the user is doing and why, then the cursor performs it in the real interface. Use ‹ ▶ › below to step at your own pace.

Step 01 / 15 Loading the prototype…
ScriptGen — the redesigned Study overview
ScriptGen — the redesigned Report
ScriptGen — the redesigned Watch-along
ScriptGen — the Calibration data tab

Desktop only — on phones this section is replaced by stills. Data: the real 66-question Hansa study; analyst scores from the trial ledger.

Photoshop mockup

ScriptGen on a laptop — the Study report for a true-crime script: timeline, open loops and evidence, mocked up in Photoshop
The app mocked up in Photoshop before it was built: the Study report, the shape of a video before a word of your own script is written.

The problem

Faster scripts should still sound like you.

Most people who try making YouTube video essays quit because of the time it takes. Finding a story worth telling, chasing down the sources, working out a structure, writing the script, recording the voice-over, and finally editing the video — every step is slow, and there are six of them.

Some creators have answered by automating the whole channel, a practice that's questionable at best. Others are looking for a middle ground: using AI to multiply their output without taking themselves out of the creative process.

But you can't hand a prompt to ChatGPT, walk away, and expect to be impressed by what you come back to. Even the brightest agents need direction — which is why ScriptGen focuses on learning about you before it tries to write for you.

ScriptGen — the Study report screen: the video's timeline, open loops and the evidence each claim rests on
The Study report as built — the shape of a video before a word of your own script is written.

The approach

Learn the creator before writing the script.

Before asking for a single script, I wanted a profile tailored to the individual. LLMs excel at writing — but who's to say they'll write something you like? ScriptGen keeps the ideal way to make a script in your hands. How much research is enough, how many sources you want in hand before a script feels solid — that depends on the person. So the engine learns it from you, in layers.

  1. Point it at videos you like. Reference videos by creators whose writing you admire.
  2. Let it fan out. A bench of model agents reads each script and video independently.
  3. It builds you an interview. Where the models' claims disagree, the engine cues those moments into a watch-along.
  4. You settle it. You tell the engine what's correct, moment by moment.
  5. It learns. Your rulings pick the best model for each analysis job — and sharpen its picture of the scripts you actually want.
The bench, live from the Nexpo watch-along — 19 seats across 8 model families voting on one moment (Q38, 23:48). Hover a seat for its read; it stays sealed until the human rules.

The hard part

Making long research runs reliable.

A pluggable harvester schedules each source under its own API budget, commits to SQLite incrementally so a long run survives cancellation, and scores convergence with a blend of a formula and an optional ONNX machine-learning model — degrading gracefully to the formula when no model is loaded.

Two scored calibration trials, straight from the seat ledgers. Mostly, effort doesn't buy accuracy — mid-effort seats match or beat their expensive siblings.

Outcome

Matching models to the work.

The app shipped — six Rust crates, one Tauri shell, built solo. But the real result is what the calibration runs revealed: 42 analyst seats across seven model families, each ruled on by a human, and the roles sorted themselves out.

Evidence hunter — codex 5.5
Caught the most on-screen proof in both studies: 30 of 37 moments on Nexpo, 10 of 13 on Fern — while staying 87–91% accurate. The seat that finds things.
Careful judge — codex 5.4 (xhigh)
20 for 20 on Fern, 90% on Nexpo. It abstains rather than guess — which is exactly what you want from the seat whose vote breaks ties.
Source librarian — sonnet 5
Named 92–100% of the real sources on Fern, best on Nexpo too. Weaker at spotting proof on screen, so it owns the "who is being cited" job instead.
Not worth a seat — haiku 4.5
40–72% accurate on Nexpo, 18 wrong calls at low effort, and the only family where the effort knob swung results by 30 points. Cheap, but dropped from the bench.

Opus 4.8 never raised a false alarm across either study — and missed half the evidence doing it. And four Fern seats asked for haiku silently ran sonnet 5; the scorer now labels every seat by the model that actually answered.

Fern · calibration run

23 seats · 50 rulings · 13 evidence moments · 13 sources

40%60%80%100%0%25%50%75%100%EVIDENCE MOMENTS CAUGHT →ACCURACY WHEN VOTING →codex54-high · codex 5.4 — 91% accurate on 33 votes · caught 7/13 evidence moments · named 46% of sourcescodex54-med · codex 5.4 — 91% accurate on 32 votes · caught 7/13 evidence moments · named 46% of sourcescodex54-xhigh · codex 5.4 — 100% accurate on 20 votes · caught 7/13 evidence moments · named 62% of sourcescodex54mini-high · codex 5.4-mini — 83% accurate on 6 votes · caught 2/13 evidence moments · named 38% of sourcescodex54mini-low · codex 5.4-mini — 71% accurate on 28 votes · caught 2/13 evidence moments · named 85% of sourcescodex54mini-med · codex 5.4-mini — 84% accurate on 31 votes · caught 6/13 evidence moments · named 85% of sourcescodex55-high · codex 5.5 — 91% accurate on 43 votes · caught 10/13 evidence moments · named 69% of sourcescodex55-med · codex 5.5 — 86% accurate on 35 votes · caught 9/13 evidence moments · named 46% of sourcescodex55-xhigh · codex 5.5 — 91% accurate on 44 votes · caught 10/13 evidence moments · named 69% of sourcessonnet-5 (asked haiku)-high · sonnet 5 — 88% accurate on 33 votes · caught 9/13 evidence moments · named 85% of sourcessonnet-5 (asked haiku)-low · sonnet 5 — 85% accurate on 20 votes · caught 8/13 evidence moments · named 77% of sourcessonnet-5 (asked haiku)-med · sonnet 5 — 94% accurate on 33 votes · caught 10/13 evidence moments · named 85% of sourcessonnet-5 (asked haiku)-xhigh · sonnet 5 — 87% accurate on 31 votes · caught 7/13 evidence moments · named 77% of sourceshaiku45-high · haiku 4.5 — 84% accurate on 19 votes · caught 5/13 evidence moments · named 69% of sourceshaiku45-low · haiku 4.5 — 71% accurate on 14 votes · caught 4/13 evidence moments · named 46% of sourceshaiku45-med · haiku 4.5 — 76% accurate on 17 votes · caught 3/13 evidence moments · named 54% of sourceshaiku45-xhigh · haiku 4.5 — 57% accurate on 14 votes · caught 1/13 evidence moments · named 54% of sourcesopus-high · opus 4.8 — 78% accurate on 32 votes · caught 4/13 evidence moments · named 69% of sourcesopus-xhigh · opus 4.8 — 80% accurate on 35 votes · caught 4/13 evidence moments · named 77% of sourcessonnet-high · sonnet 5 — 91% accurate on 33 votes · caught 9/13 evidence moments · named 100% of sourcessonnet-low · sonnet 5 — 77% accurate on 30 votes · caught 6/13 evidence moments · named 85% of sourcessonnet-med · sonnet 5 — 94% accurate on 33 votes · caught 9/13 evidence moments · named 92% of sourcessonnet-xhigh · sonnet 5 — 94% accurate on 31 votes · caught 8/13 evidence moments · named 100% of sourcescodex54-xhighcodex55-xhighhaiku45-xhighopus-xhighsonnet-xhigh

Nexpo · Deepest Rabbit Hole

19 seats · 43 rulings · 37 evidence moments · 33 sources

40%60%80%100%0%25%50%75%100%EVIDENCE MOMENTS CAUGHT →ACCURACY WHEN VOTING →codex54-high · codex 5.4 — 95% accurate on 19 votes · caught 14/37 evidence moments · named 70% of sourcescodex54-med · codex 5.4 — 89% accurate on 28 votes · caught 24/37 evidence moments · named 52% of sourcescodex54-xhigh · codex 5.4 — 90% accurate on 30 votes · caught 24/37 evidence moments · named 61% of sourcescodex55-high · codex 5.5 — 87% accurate on 38 votes · caught 30/37 evidence moments · named 82% of sourcescodex55-med · codex 5.5 — 82% accurate on 34 votes · caught 24/37 evidence moments · named 67% of sourcescodex55-xhigh · codex 5.5 — 80% accurate on 35 votes · caught 26/37 evidence moments · named 82% of sourcescodex56luna-xhigh · codex 5.6 — 81% accurate on 37 votes · caught 27/37 evidence moments · named 61% of sourcescodex56sol-xhigh · codex 5.6 — 80% accurate on 35 votes · caught 27/37 evidence moments · named 76% of sourcescodex56terra-xhigh · codex 5.6 — 79% accurate on 33 votes · caught 22/37 evidence moments · named 82% of sourceshaiku45-high · haiku 4.5 — 66% accurate on 29 votes · caught 16/37 evidence moments · named 55% of sourceshaiku45-low · haiku 4.5 — 40% accurate on 30 votes · caught 9/37 evidence moments · named 64% of sourceshaiku45-med · haiku 4.5 — 52% accurate on 29 votes · caught 12/37 evidence moments · named 52% of sourceshaiku45-xhigh · haiku 4.5 — 72% accurate on 25 votes · caught 17/37 evidence moments · named 48% of sourcesopus-high · opus 4.8 — 82% accurate on 28 votes · caught 18/37 evidence moments · named 73% of sourcesopus-xhigh · opus 4.8 — 77% accurate on 30 votes · caught 20/37 evidence moments · named 76% of sourcessonnet-high · sonnet 5 — 83% accurate on 29 votes · caught 23/37 evidence moments · named 64% of sourcessonnet-low · sonnet 5 — 75% accurate on 28 votes · caught 18/37 evidence moments · named 64% of sourcessonnet-med · sonnet 5 — 90% accurate on 31 votes · caught 26/37 evidence moments · named 70% of sourcessonnet-xhigh · sonnet 5 — 89% accurate on 28 votes · caught 20/37 evidence moments · named 64% of sourcescodex54-xhighcodex55-highhaiku45-lowopus-highsonnet-med
codex 5.4codex 5.5codex 5.6codex 5.4-minihaiku 4.5sonnet 5opus 4.8 dot size = share of sources correctly named · hover a dot
Every analyst seat from two calibration studies, scored against my own rulings. Up = it's right when it speaks; right = it catches the moments where proof is actually on screen. Scores from seat_scores.json; "unknown" votes count as abstentions, never errors.

Interface redesign · before and after

Same tool, plain English, data first.

The first UI spoke the engine's internal taxonomy — seats, bench, rulings. The redesign says analysts, what they saw, your call; leads with per-model accuracy; and gives every moment card a “wrong moment?” flag.

Before · Study overview ScriptGen before the redesign — Study overview
After · Study overview ScriptGen after the redesign — Study overview
Study overview — a status wall became one calibration meter and a resume-where-you-left-off row.
Before · Report ScriptGen before the redesign — Report
After · Report ScriptGen after the redesign — Report
Report — the timeline gained claim ticks, cuts and the sponsor break; evidence stats sit beside it instead of below.
Before · Watch-along ScriptGen before the redesign — Watch-along
After · Watch-along ScriptGen after the redesign — Watch-along
Watch-along — the interview moved into the app: cued video, pinned seekbar, a bench you reveal on demand.
Before · Style profile ScriptGen before the redesign — Style profile
After · Style profile ScriptGen after the redesign — Style profile
Style profile → Calibration data — the old profile page led with metric tables; the new build moves the numbers into a dedicated, sortable tab and locks the profile until the rulings earn it.
Next project Ophthalytics