Speech → Subtitles · On device · iPhone + iPad

Turn speech into an editable subtitle file.

Choose Apple on-device system recognition or one of four downloadable local OpenASR models, recognize speech in an audio or video source, review the text, and export SRT, ASS, or TXT. Apple system recognition and Whisper Tiny are included without Pro; Whisper Base, X-ASR, and Whisper Small require ClipFlow Pro.

Recognition first, rendering optional

How to create subtitles from speech on iPhone.

This tool generates a subtitle document. Burning that text into video is a separate Pro workflow, so you can review the words before changing any frames.

Step 1

Select a source, method, and format

Choose one audio or video source with more than two seconds of audio and select SRT, ASS, or TXT. Choose system recognition, then select Chinese, English, Japanese, French, Korean, Spanish, or German from the language menu; alternatively, choose an installed OpenASR model. Apple system recognition and Whisper Tiny are free. Downloading or using Whisper Base, X-ASR, or Whisper Small requires an active ClipFlow Pro entitlement. ClipFlow remembers your recognition choice and preferred output format.

Step 2

Recognize and edit

System recognition requests speech-recognition permission and is required to remain on device; it does not request microphone permission or fall back to cloud recognition. OpenASR needs no system speech-recognition permission and automatically handles the selected model's supported languages. Whisper performs language identification; X-ASR uses a fixed Chinese-English bilingual mode. Review compact cue rows with the cue number before its text. Swipe left to edit or delete; touch and hold to see its timing or to copy, edit, or delete. Use your personal dictionary to apply saved source-to-replacement rules, or enable automatic replacement in Settings for future recognition jobs. Chinese and Japanese recognition removes extra inter-segment spaces, while languages that use word spacing preserve it.

Step 3

Export or embed

Save the subtitle file, or use the dedicated Embed in Video action to continue into Burn In Subtitles. Installed OpenASR models can be switched or deleted independently. If Pro expires, an already-downloaded Pro model remains available to delete for storage management but cannot be selected or used until Pro is active again.

One system option · Four audited downloads

Match the recognition method to your device and recording.

Apple's system option and Whisper Tiny are available without Pro. The other three downloadable OpenASR models require Pro. Downloads range from about 63 MB to 303 MB, and every downloadable model is below 1.1 GB in OpenASR's upstream isolated-process peak-memory reference. Their five-level speed ratings are relative comparisons derived from OpenASR's M1 CPU benchmark, not guaranteed timings for a particular iPhone or iPad.

No download · Free · 7 languages

Apple On-Device System Recognition

Select Chinese, English, Japanese, French, Korean, Spanish, or German before recognition. This option uses system speech-recognition permission, is required to keep audio on the device, and does not request microphone permission.

Free · Lightweight · 63 MB · Speed 5/5

Whisper Tiny Q8

About 275 MB reference peak memory. The smallest and fastest choice, with lower recognition strength. Supports Chinese, English, Japanese, and other languages.

Pro · Balanced · 108 MB · Speed 4/5

Whisper Base Q8 · Recommended

About 405 MB reference peak memory. The recommended everyday balance of footprint and recognition strength. Supports Chinese, English, Japanese, and other languages.

Pro · Chinese-English · 176 MB · Speed 4/5

X-ASR Chinese-English Q8

About 549 MB reference peak memory. A focused option for Chinese, English, and code-switched speech; it does not provide Japanese coverage.

Pro · High quality · 303 MB · Speed 3/5

Whisper Small Q8

About 1.01 GB reference peak memory. A stronger multilingual option when additional processing cost is acceptable. Supports Chinese, English, Japanese, and other languages.

Choose the output

SRT and ASS preserve timed cues for caption workflows; generated TXT results retain readable time ranges. Every recognized format can be reviewed row by row: swipe for edit and delete actions, or touch and hold for copy, edit, and delete. Recognition is only a first draft: background noise, names, accents, and specialist vocabulary can require corrections. Personal dictionary entries stay in local app preferences until you delete them in Settings or uninstall ClipFlow.

On iOS 17 or later, ClipFlow offers Apple on-device system recognition plus four downloadable OpenASR models. It can keep more than one verified OpenASR model installed, but it loads only the selected model for a recognition job. Downloads happen only after confirmation, installed models use local storage until you delete them, and both recognition methods process audio entirely on your device.

Apple system recognition and Whisper Tiny are free. Whisper Base, X-ASR Chinese-English, and Whisper Small require active Pro access. Users who installed Whisper Large V3 Turbo in an earlier version can continue using it with Pro or delete it to reclaim storage; it is no longer offered for download and disappears from model management after deletion. If Pro expires, installed Pro models cannot be selected or used, but they remain visible so you can delete them.

Open-source license and notices

ClipFlow uses the Apache-2.0 OpenASR engine. These versioned documents preserve the license terms, attribution, and ClipFlow modification notice for the OpenASR 0.1.23 / 7a5b3bd integration. The NOTICE also links to the complete native dependency license archive.

Related guides