Canon-first planning
Writing the synopsis (including the ending) before parallel drafting bounds plot drift across hundreds of pages.
Benchmarks
Wall-clock timings and stability findings from local runs on Apple Silicon with llama-3.1:8b, plus sample generated book pages (redacted previews only).
M2 Ultra · four Ollama servers · sixteen concurrent writers · Q6_K quantization
| Configuration | Director | Planner | Writer | Reviewer | Total |
|---|---|---|---|---|---|
| Fiction 90k | 25–30s | ~1 min | 45–90 min | 10–20 min | 1–2 h |
| Fiction 150k | ~30s | 85s | 95–130 min | 14–22 min | 110–155 min |
| Fiction 300k | ~45s | 130s | 190–245 min | 28–45 min | 220–295 min |
Writing the synopsis (including the ending) before parallel drafting bounds plot drift across hundreds of pages.
Reviewer concurrency stays at one quarter of writer slots, eliminating transient llama.cpp failures on long prompts.
Atomic checkpoints after every scene mean a multi-hour run survives interruption with at most one scene of rework.
Pictures mode calls Ollama image APIs; animated mode renders ManimRedux AVI→MP4; video mode reserves a text-to-video hook with ffmpeg stubs.
All modes emit fragment LaTeX—graphicx figures, media9 clips, hardback 7×10″ layout, XeLaTeX for Devanagari language courses.
Fill-in-the-blank, multiple choice, true/false, and mix-and-match exercises rotate per chapter with a separate answer key.
Preview pages from a Hindi language textbook run—vocabulary tables, lesson prose, and structured exercises. Full manuscripts and checkpoints are not published here.
Before retry-and-fallback logic, reviewer-stage KV-cache failures aborted roughly 24% of 90k-word runs. With tier-specific concurrency caps and transient-error retries, completion reached 100% across fifty trials, with graceful stitch fallback on at most one chapter per run in 18% of cases.