Built because typing subtitles at 1 a.m. is not content creation.
Most people who post talking videos are not editors. They record on a phone, they have twenty minutes before the next thing, and the captions are the part they skip when they are tired. PutCaption exists so that part takes two minutes and looks like it took an hour.
It is deliberately small. There is one job: put the right words on screen at the right moment, in a style that people keep watching. No stock library, no AI voiceovers, no clip-finder that decides what is interesting about your video. You know that already.
The engine that draws captions in the editor is the exact same code that draws them into your export. That sounds obvious, and yet it is the reason most caption tools produce a file that looks slightly different from the preview.

What we hold ourselves to
Your words, not ours
We do not rewrite what you said. The transcript is the transcript; you edit it, not a model.
Honest about the machine
Speech-to-text is done by a cloud engine. Only the audio leaves your server, and only to be transcribed.
No dark patterns
Free means free with no watermark. Cancelling is one click. Exports you already made stay yours.
Credits
The demo on the home page uses real footage so you can judge the transcription honestly.
- Demo footage: “The Mirror Effect – How to be confident on camera” by Girl Director TV, published under the Creative Commons Attribution licence. The captions on it were produced by PutCaption itself, untouched.
- Photography: Afffect, Vitaly Gariev, Detail .co, Vladislav Anchuk and Videodeck .co via Unsplash.
- Type: Space Grotesk and Instrument Serif, plus the caption fonts, all from Google Fonts under the Open Font Licence.