Word-by-Word Caption Generator
Upload a video and the word-by-word caption generator adds each word the instant you say it, with a quick bounce. The line builds as you talk.
Create an account to generate and export your captions.
How the Word Pop style plays
Word Pop is the Sycamore template for word-by-word captions. Each word stays hidden until you say it. At that moment it appears at full strength and bounces up from a smaller size, so the line fills in at the pace of your voice.
A few words share the screen at once, about four by default. When the line is done, the next one starts fresh. The text is set in Inter at a heavy weight, white by default, and the word you're saying right now can take a highlight color.
Viewers can't read ahead, because the rest of the sentence isn't there yet. Their eyes stay on the newest word, and that keeps them listening to you. It also helps anyone watching on mute follow along at your speed.
Full-line subtitles vs. word by word
Both show what you said. They feel very different on a phone screen.
When words appear
- Full-line subtitles
- The whole sentence shows at once
- Word Pop in Sycamore
- Each word shows the moment it's spoken
Where the eye goes
- Full-line subtitles
- Viewers read ahead, then wait
- Word Pop in Sycamore
- Viewers follow the newest word
Words on screen
- Full-line subtitles
- Often a full sentence or two lines
- Word Pop in Sycamore
- 1 to 8 words, your choice
Motion
- Full-line subtitles
- Text sits still
- Word Pop in Sycamore
- Each word bounces in, with an optional slide
| Full-line subtitles | Word Pop in Sycamore | |
|---|---|---|
| When words appear | The whole sentence shows at once | Each word shows the moment it's spoken |
| Where the eye goes | Viewers read ahead, then wait | Viewers follow the newest word |
| Words on screen | Often a full sentence or two lines | 1 to 8 words, your choice |
| Motion | Text sits still | Each word bounces in, with an optional slide |
Shape the pop to your voice
Words per line
Pick anywhere from 1 to 8. Smart chunking breaks lines at natural pauses, so a line rarely cuts off mid-thought.
Pop feel
Choose snappy, smooth or bouncy for how each word lands.
Slide direction
Words can slide in from the top or bottom, or just scale up in place.
Text and highlight color
Set the text color, and give the word you're saying its own color or turn that off.
Size, weight and case
Font size from 24 to 96 px, Regular, Bold or Black weight, and an all-caps option.
Tick sound per word
Add a soft tick as each word appears if you want the captions to be heard as well as seen.
Getting word-by-word captions to read well
Fast talkers should drop to two or three words per line. Each word needs a beat on screen before the line clears, or viewers miss it.
Captions sit in the center by default. If your face is in the middle of the frame, move them to the top or bottom and use the edge offset to keep them clear of the app's buttons.
Word Pop works on 9:16 vertical video for TikTok, Reels and Shorts, and on 16:9 and 1:1 too.
Let your next video pop word by word
Frequently
Asked Questions
What is a word-by-word caption generator?
It's a tool that transcribes your video with timing for every word and shows each word on screen as you say it. In Sycamore, the Word Pop style does this with a quick bounce on every word.
How do I make captions appear one word at a time?
Upload your video to Sycamore and pick the Word Pop style. Each word is timed to your audio, so it appears right when you say it.
Can I show just one word on screen at a time?
Yes. Set words per line to 1 and each word gets the screen to itself. You can go up to 8 if you'd rather show longer lines.
Can I turn off the highlight on the current word?
Yes. Set the active word option to none and every word stays the same color. Or keep it on and pick any highlight color you like.
What files can I upload for word-by-word subtitles?
MP4, MOV, WebM and MP3. You can click any word to fix a transcription mistake before you export.