Tooling · 2026
scriptgen
Study the style of the creators you love, then write and assemble video projects.
Interactive prototype · guided walkthrough
The redesigned ScriptGen app, driven by a scripted tour — each step first explains what the user is doing and why, then the cursor performs it in the real interface. Use ‹ ▶ › below to step at your own pace.




Desktop only — on phones this section is replaced by stills. Data: the real 66-question Hansa study; analyst scores from the trial ledger.
Photoshop mockup

The problem
Faster scripts should still sound like you.
Most people who try making YouTube video essays quit because of the time it takes. Finding a story worth telling, chasing down the sources, working out a structure, writing the script, recording the voice-over, and finally editing the video — every step is slow, and there are six of them.
Some creators have answered by automating the whole channel, a practice that's questionable at best. Others are looking for a middle ground: using AI to multiply their output without taking themselves out of the creative process.
But you can't hand a prompt to ChatGPT, walk away, and expect to be impressed by what you come back to. Even the brightest agents need direction — which is why ScriptGen focuses on learning about you before it tries to write for you.

The approach
Learn the creator before writing the script.
Before asking for a single script, I wanted a profile tailored to the individual. LLMs excel at writing — but who's to say they'll write something you like? ScriptGen keeps the ideal way to make a script in your hands. How much research is enough, how many sources you want in hand before a script feels solid — that depends on the person. So the engine learns it from you, in layers.
- Point it at videos you like. Reference videos by creators whose writing you admire.
- Let it fan out. A bench of model agents reads each script and video independently.
- It builds you an interview. Where the models' claims disagree, the engine cues those moments into a watch-along.
- You settle it. You tell the engine what's correct, moment by moment.
- It learns. Your rulings pick the best model for each analysis job — and sharpen its picture of the scripts you actually want.
The hard part
Making long research runs reliable.
A pluggable harvester schedules each source under its own API budget, commits to SQLite incrementally so a long run survives cancellation, and scores convergence with a blend of a formula and an optional ONNX machine-learning model — degrading gracefully to the formula when no model is loaded.
Outcome
Matching models to the work.
The app shipped — six Rust crates, one Tauri shell, built solo. But the real result is what the calibration runs revealed: 42 analyst seats across seven model families, each ruled on by a human, and the roles sorted themselves out.
- Evidence hunter — codex 5.5
- Caught the most on-screen proof in both studies: 30 of 37 moments on Nexpo, 10 of 13 on Fern — while staying 87–91% accurate. The seat that finds things.
- Careful judge — codex 5.4 (xhigh)
- 20 for 20 on Fern, 90% on Nexpo. It abstains rather than guess — which is exactly what you want from the seat whose vote breaks ties.
- Source librarian — sonnet 5
- Named 92–100% of the real sources on Fern, best on Nexpo too. Weaker at spotting proof on screen, so it owns the "who is being cited" job instead.
- Not worth a seat — haiku 4.5
- 40–72% accurate on Nexpo, 18 wrong calls at low effort, and the only family where the effort knob swung results by 30 points. Cheap, but dropped from the bench.
Opus 4.8 never raised a false alarm across either study — and missed half the evidence doing it. And four Fern seats asked for haiku silently ran sonnet 5; the scorer now labels every seat by the model that actually answered.
Fern · calibration run
23 seats · 50 rulings · 13 evidence moments · 13 sources
Nexpo · Deepest Rabbit Hole
19 seats · 43 rulings · 37 evidence moments · 33 sources
seat_scores.json; "unknown" votes count as abstentions, never errors.Interface redesign · before and after
Same tool, plain English, data first.
The first UI spoke the engine's internal taxonomy — seats, bench, rulings. The redesign says analysts, what they saw, your call; leads with per-model accuracy; and gives every moment card a “wrong moment?” flag.