AI Podcast Clipper — Long-Form to Short-Form · beiryu
AI Podcast Clipper — Long-Form to Short-Form
Paused
An automated pipeline that turns full podcast episodes into publish-ready vertical clips — transcription, moment selection, speaker-tracked cropping, and burned-in captions — running on serverless GPU compute with durable job orchestration.
Every podcast produces hours of material that never reaches an audience beyond the people who already subscribe. Clipping it into short vertical video is the highest-return distribution work available, and it is also the most tedious — which is exactly the shape of problem worth automating.
This platform takes a full episode and returns finished clips: transcribed, selected for the most compelling moments, cropped to follow whoever is speaking, captioned, and exported at the aspect ratio of the destination platform.
Transcription — Whisper for accurate, word-timed transcripts across accents and languages, which everything downstream depends on
Moment selection — the transcript analysed for the segments that stand alone: a complete thought, a strong opening, a payoff at the end. A clip that starts mid-sentence is worthless regardless of how good the content is
Speaker detection and tracking — computer vision identifying the active speaker and following them through the frame
Dynamic cropping — reframing 16:9 to 9:16 around the tracked speaker, rather than a static centre crop that cuts people out of shot
Caption generation — word-timed captions burned in with styling, since most short-form video is watched muted
Multi-format export — 9:16, 1:1, and 16:9 from the same source
The engineering constraint is that video processing is expensive, slow, and fails in the middle. The architecture is built around that rather than around the happy path:
Serverless GPU compute through Modal — inference runs on GPUs that exist only for the duration of the job, so cost tracks usage rather than provisioned capacity sitting idle between episodes
Durable orchestration through Inngest — the pipeline is a sequence of steps with independent retries, so a transcription failure does not discard completed work and a crash mid-render resumes rather than restarting
Containerised processing for FFmpeg and model dependencies, keeping the environment identical between development and production
S3 for media with lifecycle policies, since raw episode uploads are large and short-lived
Redis caching on the paths that repeat
Parallel episode handling, so a user uploading a back catalogue is not processed one at a time
Drag-and-drop upload with support for common video and audio formats
Live processing status driven from real pipeline state, with per-step progress — the difference between a user waiting patiently and a user assuming it broke
Browser preview and adjustment — every generated clip reviewable and trimmable before export
Custom branding — logo placement, caption styling, and colour themes applied across a batch
Batch queueing for multiple episodes
Subscription billing through Stripe, metered against processing minutes
This project is about production AI infrastructure rather than model work: orchestrating expensive, failure-prone, long-running jobs reliably enough that a user trusts the platform with a two-hour upload.