Transcription settings for a single interview

Tune the model, profanity filter, spoken punctuation and speaker detection for a single recording on iPhone, iPad or Mac, without touching your defaults.

Every interview carries its own copy of the transcription settings. Before you transcribe, you can adjust any of them for that recording alone — your defaults stay untouched.

Where to find them

They live in the Transcription setup box, which appears on any interview that has not been transcribed yet, just above the Transcribe button.

On iPhone and iPad, open the interview and scroll to the box. Tap the ⓘ button in its corner to show a short explanation under each setting, and tap again to hide them.

On Mac, open the interview; the box sits in the document window with the same controls.

The settings

Default Language — the language spoken in the audio. See setting the language for one interview.

Default Model — which recognition model to use. Matching the model to your material is the single biggest lever on quality:

Model Best for
Default General-purpose, high-fidelity audio that fits none of the cases below
Long form conversation Interviews, meetings, podcasts and other conversational recordings
Single shot directed speech Short utterances — a single command or a brief directed phrase
Voice commands or voice search Short voice queries
Audio from a phone call Call audio, including lower sample rates and call quality
Conversation between a medical provider and patient Clinical conversations
Originated from dictation notes by a medical provider Medical dictation

For most interview work, Long form conversation is the right choice.

Profanity Filter — when on, all but the first character of a filtered word becomes asterisks, so it reads as f***. When off, nothing is filtered.

Spoken Punctuation — when on, spoken punctuation becomes real symbols, so “how are you question mark” comes out as “how are you?”. Useful for dictation, usually unhelpful for natural conversation where people say those words for real.

Speaker Diarization — when on, Transcriber tags each recognised passage with a speaker number, shown as Speaker 1, Speaker 2 and so on beside the transcript.

Max speaker count — how many speakers to expect, from 1 to 12. Only adjustable while speaker diarization is on. Treat it as a ceiling rather than an exact figure; the system settles on the real number within the range you give it.

Then transcribe

Choose Transcribe when the settings look right. Changes here apply to this interview only — to change them for everything, see default transcription settings.

Related articles