Build narrated document experiences from PDF, EPUB, and plain-text sources. screencastgen can generate audio, synchronized highlighting, lip-synced presenters, browser reader bundles, and educational animations.
| Goal | Start here |
|---|---|
| Install and verify the project | Getting started |
| Generate narrated audio | Audio guide |
| Highlight words in sync with narration | Highlight guide |
| Add a lip-synced presenter | Lip-sync guide |
| Generate an educational animation | Visualization guide |
| Run the browser application | Web application guide |
| Split processing across CPU and GPU hosts | Remote GPU guide |
The CLI and web application call a shared set of pipelines. Provider registries select the TTS, alignment, lip-sync, and visualization implementations. A remote GPU server can run model-heavy operations while a client performs document and media processing.
flowchart LR
User[CLI or web application] --> Pipelines
Pipelines --> Core[Document and media core]
Core --> TTS[TTS providers]
Core --> Align[Alignment providers]
Core --> Lip[Lip-sync providers]
Pipelines --> Visual[Visualization providers]
Core <--> Remote[Remote GPU server]
Read the architecture overview for the complete design or browse the developer reference for individual modules and services.
Document pipelines share the same extraction-to-speech foundation:
Highlight and lip-sync outputs add word-level alignment after synthesis. PDF inputs use page images plus PyMuPDF word positions for precise highlighting; other document formats use a reflowed text renderer. Lip-sync outputs then animate the reference presenter with the selected face-animation provider and package the result as a LipSync Reader or EPUB depending on the requested format. The offline reader ZIP is the recommended portable output.
The same provider abstraction works locally or against a remote GPU server: client-side document processing stays on the CPU host, while model-heavy TTS, alignment, and lip-sync work can run remotely.