Accurate captions depend on more than just recognizing words — you need to know exactly when each word starts and ends. Gemini's multimodal models can return structured JSON with word-level timestamps, which is exactly what caption animations need.
Structured output
By requesting a strict response schema, SnipCaptions gets back a clean list of words with start and end times in seconds. That data drives both the live preview and the exported video without any post-processing.
Bring your own key
Because you connect your own Gemini key, you control rate limits and cost, and you are not locked into a vendor's captioning pricing.