Transcripts and speakers
How notes get transcripts with each voice attributed.
Every note you record or import is transcribed in the cloud once its audio finishes uploading. You don't pick a language or press anything. It runs automatically, and the finished transcript appears on the note.
Transcription runs after upload, not while you record. Live on-screen transcription is not available in this version.
What the engine detects
- Language. The spoken language is identified automatically; there is no setting to configure beforehand.
- Speakers. Different voices are separated so a conversation reads like a dialogue rather than one wall of text.
- Word timing. Every word is timestamped, which is what makes tap-to-jump playback and SRT subtitle export possible.
Correcting the transcript
The transcript can't be edited directly in the app right now. To fix a misheard name, a technical term, or a garbled phrase, ask Oakey on that note to correct the passage. Transcripts are versioned: each correction creates a new revision instead of overwriting the old text, so a background job that finishes late can never replace your corrections.
To get names and terms right the first time, add them under Settings › Dictionary. You can also search within a transcript.
Speakers
Speaker separation assigns generic labels: Speaker 1, Speaker 2, and so on. The engine knows the voices are different people; it doesn't know who they are. The phone app doesn't let you rename speakers right now; the macOS desktop preview can.
When you rename a speaker in the macOS preview, Oak Note keeps an encrypted voice template of that person for your account and can label the same person by name on later notes. It never puts a name on a voice nobody named. There is no screen to review or remove remembered people yet; deleting your account removes them.
You can optionally set up a voice profile in Settings › Voice profile. Future recordings can then compare each diarized speaker only with your account-scoped profile. A successful match keeps your account display name and adds a You marker; weak or unavailable matches stay anonymous. This convenience label is not voice authentication.
Named speakers, including You and anyone named in the macOS preview, make summaries and to-dos attribute things to real people, and the names appear in exports.
When results aren't perfect
Speaker separation works from audio alone, so crosstalk, very similar voices, and distant or noisy recordings all make it harder: segments can land under the wrong name, or one voice can split into two.
For future recordings, placing your Oak Note closer to the group helps more than anything else. If a note keeps failing to process, see Troubleshooting.