Everything, documented.
The complete Caption Plug manual - every panel section, both output modes, and how the pieces fit together. Prefer step-by-step recipes? The tutorials walk the common tasks one at a time.
Overview
Caption Plug is a panel that lives inside Adobe Premiere Pro (2020/v14 or newer, Windows and macOS). It listens to your timeline audio, transcribes it with word-level timestamps (Whisper, through your own API key), and places fully animated, frame-accurate captions straight onto your sequence - 45 caption presets, 8 keyword-title presets, and 62 bundled fonts. No export-upload-reimport round trip, no cloud render queue: your footage never leaves your machine, only the audio you choose to transcribe does.
Getting started
- Install. Download the installer for your platform from your account page and click through - the right Premiere folder is preselected and no admin rights are needed. Restart Premiere and open Window ▸ Extensions ▸ Caption Plug. Full walkthrough: setup guide.
- Sign in.First launch asks for your captionplug.com email and password - that's the whole activation, no license keys. One purchase covers 3 machines.
- Add an API key. Paste an OpenAI or Groq key under Transcription(the panel's guided tutorial shows exactly where to get one - about two minutes, one time).
- Generate. Open a sequence, pick a preset, hit Generate. Captions land on a new track, synced to the word.
The panel at a glance
The panel is organised into collapsible sections - closed sections show their current setting on the right, and the live preview plus the Generate button stay visible at all times.
- Animation preset - the 45 caption presets and 8 keyword-title presets, grouped by family.
- Style - font, size, colors, words per caption, position. The color wells open a real picker: saturation/hue surface, hex field, preset swatches.
- Audio - where the audio comes from. By default Generate grabs it straight off your timeline; the 🎯 Use timeline audio button re-exports on demand after you move clips.
- SFW mode - censoring and bleeps (details below).
- Transcription - provider (OpenAI or Groq), API key, language.
- Output - Animated clips or Editable text, plus Export .srt only.
- Storage - where rendered media lives and the cleanup button.
- Account & updates- who's signed in, your activated machines, and Check for updates.
The status pill in the header shows which sequence the panel is locked onto, with its resolution and frame rate. The ? Tutorial button in the top bar reopens the guided first-run tour any time. A Show debug log toggle in the bottom dock reveals the detailed status log - it stays hidden by default and pops open automatically if something needs your attention.
Generating captions
- Open the sequence you want captioned and check the status pill picked it up.
- Pick an animation preset and tune the style - the live preview runs the actual render engine, so what you see is literally what gets rendered.
- Select just your dialogue clips for the cleanest transcript: the panel mutes unselected audio (music beds, SFX) during its export and offsets the captions so they land exactly where those clips sit. With nothing selected, the whole sequence is captioned.
- Hit Generate. Transcription runs through your key, then captions render and land on a fresh track at the exact frame.
Words per captionis the single most style-defining setting: 1 gives the one-word-at-a-time style, 2-4 reads best for the classic MrBeast look. Word timing comes straight from Whisper, so highlights follow the actual speech. Transcripts are cached per audio file - toggling settings and regenerating doesn't re-upload or re-pay.
Output modes
- Animated clips- transparent PNG image-sequence clips rendered at your sequence's exact resolution and frame rate, placed on a new video track. Because every frame is rasterized against the sequence fps, sync is frame-accurate by construction - no drifting keyframes. Frames where nothing moves are stored as hard links rather than copies, so disk use stays far below one-PNG-per-frame.
- Editable text- a native Premiere caption track instead of rendered media: text stays editable in the caption editor, styling lives in Essential Graphics (set it once, save a Track Style), and almost nothing touches the disk. Per-word animation presets don't apply in this mode - that's a Premiere limitation, not a setting.
- Export .srt only - transcribes and saves a standard SRT file wherever you choose (YouTube, other editors), and also attaches it as a caption track on Premiere versions that support it.
Presets & styling
The 45 caption presets are grouped into seven families - Essentials (Beast Pop, CapCut Bounce, Punch In, Snap Up), Hype & Impact (Impact Shake, Earthquake, Drop Bounce…), Comic & Cartoon (Jelly Squash, Sticker Pop, Letter Cascade…), Glitch & Retro (RGB Glitch, VHS Tape, Neon Sign, 8-Bit Build…), Karaoke & Flow (Karaoke Fill, Spotlight Focus, Lyric Bar…), Boxes & Highlights (Hormozi Box, TikTok Pill, Redacted…) and Cinematic & Per-letter (Clean Fade, Typewriter, MoGraph Stack…). Every preset is a mechanically different animation. Try them all live on the presets page.
The 62 bundled fonts are grouped the same way (Beast & Impact, Comic & Fun, Hand & Marker, Retro & Y2K, Pixel & Typewriter…) and render correctly in animated captions even on a machine where none of them are installed - the panel carries them internally. The installer also adds them to your per-user font folder so Editable text mode can use them in Essential Graphics. All 62 are SIL OFL / Apache licensed: safe for commercial and monetized videos.
Keyword titles
The second sub-category in the preset grid, built for long-form. Instead of captioning the whole speech, the panel picks the key momentsout of the transcript - names of people, places and brands, numbers and prices, short punchy statements - and drops a smooth, high-production title on each one (Lower Third, Stat Punch, Count Up, Search Bar, Glass Card…). Picking uses the same API key you pasted for transcription, with phrases matched back onto Whisper's word timestamps so timing stays exact; if the call fails, a free local fallback detects names and figures instead. Keyword titles render via the Animated clips output.
SFW mode (censoring & bleeps)
For monetized and brand-safe edits, SFW mode censors captions, bleeps the timeline - or both. See it in action in the censor demo.
- Censor inappropriate words - swearing, sexual, violent, weapon and drug words are masked in the captions (
F**kstyle: first and last letter kept). Matching is exact-word against curated lists, so clean words that merely contain a bad substring - class, Sussex, document - are never touched. Applies to animated clips, editable captions and SRT exports alike. - Add censor bleeps to the timeline- drops a 1 kHz bleep tone on a new audio track over swear and sexual words. Back-to-back swears merge into one bleep, and the tone is a tiny generated WAV stored with your project media so everything relinks fine.
- Cut the original audio under each bleep(on by default) - other audio is gated to silence across each bleep window with clip-volume keyframes and soft 40 ms ramps, leaving just the bleep audible. Fully non-destructive: delete the volume keyframes to undo.
- 🔇 Censor audio only - bleeps your timeline without generating captions at all: one button transcribes, finds the words, and places the bleeps and dialogue cuts, leaving your video tracks untouched.
Transcripts are cached uncensored, so toggling SFW and regenerating doesn't re-upload anything.
Transcription & API keys
Transcription runs through your own API key - OpenAI (whisper-1, platform.openai.com/api-keys) or Groq (whisper-large-v3, console.groq.com/keys). Groq is fast, currently very cheap, and noticeably better on names. A typical short-form clip costs well under a cent; the key is saved locally and never touches our servers. Whisper handles most major languages - leave the language on auto-detect or set it explicitly.
Uploads are capped at 25 MB per file by the providers (roughly 30-40 minutes of MP3 speech). With ffmpeg on your PATH the panel auto-compresses bigger files; otherwise export the audio as MP3 first and pick that file.
Storage & media
Rendered caption media is saved into a Caption Plug Media folder next to your project file, so projects relink cleanly on any machine. Identical hold frames are stored as hard links rather than copies, keeping real disk use at a fraction of the frame count. The 🧹 Clean up old caption media button in the Storage section deletes every render folder the current project no longer references.
Account, license & updates
- One purchase covers 3 machines you own - sign in on each, and free a slot from your account page any time you switch computers.
- Signing in needs internet once; after that the panel works offline. It re-checks silently in the background, and only an explicitly revoked license ever signs you out - being offline never punishes you.
- Your password is never stored on your machine - the panel keeps only a server-minted activation token.
- Updates are built in: the panel checks for new builds twice a day (or on demand via Check for updates). When one ships, an ⬆ Update button appears in the top bar - one click installs it in place, settings kept, then restart Premiere. See what shipped in the changelog.
Honest limits
- Transcription needs internet and a (cheap) API key; rendering and placement are fully local and nothing else leaves your machine.
- 25 MB per upload (≈30-40 min of MP3 speech). Longer? Split, or let ffmpeg auto-compress.
- Animated clips write real PNG files. Hard-linking keeps disk use far below one-PNG-per-frame, but heavily animated long videos still add up - the cleanup button reclaims it.
- Animated caption clips are rendered graphics: to change one word, fix the transcript and regenerate (fast - transcripts are cached), or use Editable text mode where every caption stays editable.
- Native caption tracks can't do per-word animation - a Premiere limitation, not a setting.
Getting help
The tutorials solve the common problems step by step, the FAQ answers the pre-purchase questions, and the changelog lists what shipped when. Anything else: support@captionplug.com - a human reads every message.
Read enough? $8.99 once.
Every preset, every font, the censor, and all v1.x updates - one payment, no subscription.