ABAgenticBookWriterRead paper

Benchmarks

Measured results

Wall-clock timings and stability findings from local runs on Apple Silicon with llama-3.1:8b, plus sample generated book pages (redacted previews only).

Fiction pipeline timings

M2 Ultra · four Ollama servers · sixteen concurrent writers · Q6_K quantization

ConfigurationDirectorPlannerWriterReviewerTotal
Fiction 90k25–30s~1 min45–90 min10–20 min1–2 h
Fiction 150k~30s85s95–130 min14–22 min110–155 min
Fiction 300k~45s130s190–245 min28–45 min220–295 min

Key findings

Canon-first planning

Writing the synopsis (including the ending) before parallel drafting bounds plot drift across hundreds of pages.

KV-cache-aware concurrency

Reviewer concurrency stays at one quarter of writer slots, eliminating transient llama.cpp failures on long prompts.

Scene-granular resume

Atomic checkpoints after every scene mean a multi-hour run survives interruption with at most one scene of rework.

Automated media pipeline

Pictures mode calls Ollama image APIs; animated mode renders ManimRedux AVI→MP4; video mode reserves a text-to-video hook with ffmpeg stubs.

Native LaTeX publishing

All modes emit fragment LaTeX—graphicx figures, media9 clips, hardback 7×10″ layout, XeLaTeX for Devanagari language courses.

Structured language drills

Fill-in-the-blank, multiple choice, true/false, and mix-and-match exercises rotate per chapter with a separate answer key.

Sample generated pages

Preview pages from a Hindi language textbook run—vocabulary tables, lesson prose, and structured exercises. Full manuscripts and checkpoints are not published here.

Hindi vocabulary table with Devanagari script
Language textbook — bilingual vocabulary table
Hindi practice exercises with typed drill formats
Language textbook — structured practice problems
Hindi lesson prose with example sentences
Language textbook — lesson content and examples

Stability

Before retry-and-fallback logic, reviewer-stage KV-cache failures aborted roughly 24% of 90k-word runs. With tier-specific concurrency caps and transient-error retries, completion reached 100% across fifty trials, with graceful stitch fallback on at most one chapter per run in 18% of cases.