Adding Captions Automatically¶
Acoustica can transcribe speech and turn it into captions for you, using the Whisper speech recognition engine. This is normally far quicker than typing captions by hand: let the engine produce a first pass, then correct it in the Caption Editor.
Before You Start¶
Speech recognition needs a model, and which model you use has a large effect on both the accuracy and the time the transcription takes. Open the Speech Recognition Settings in the Preferences and check that:
-
a model is selected — the installer includes the tiny models, and larger and more accurate ones can be downloaded from within the settings; and
-
the computing device is the one you want to use. Speech recognition is computationally heavy and a GPU is generally much faster than a CPU.
See Speech Recognition Settings for the full description of the model manager.
Transcribing a Recording¶
-
Select the part of the recording you want to transcribe. Transcription runs on the selected time range, so select the whole recording to caption all of it.
-
Choose Add Captions Automatically... from the Analysis menu, or from the Automatic Editing sub menu in the Edit menu.
-
Choose the Language of the speech in the dialog that appears, and confirm.

Acoustica remembers the language you chose and offers it again next time.
Transcription then runs as a background task and the recognized speech is added as captions, each with its own time region. Depending on the length of the selection, the model size and the computing device, this can take from a few seconds to considerably longer than real time.
Choosing the Language¶
Set the language to the language actually being spoken. If you work in several languages and have language-specific models installed, switch on Automatically select language-specific model if available in the Speech Recognition Settings: a model trained on a single language generally performs better than a multilingual model of the same size.
After Transcription¶

The result is a set of ordinary captions, so everything in Caption Editing applies to them:
-
correct the text of a caption by selecting it in the caption list and editing it;
-
assign an Actor to each caption, which is what makes a transcript of a multi-speaker recording readable;
-
adjust the time region of a caption by unlocking it and dragging its edges in the waveform; and
-
export the result as subtitles or as a transcript with the Export button, or with Export Captions or Transcription... from the Export sub menu in the File menu.
Automatic speech recognition is not perfect and its accuracy falls sharply with background noise, overlapping speakers and strong accents. On difficult material, cleaning the audio first — with DeNoise, Extract:Dialogue or DeVerberate — will usually improve the transcription more than switching to a larger model.