Skip to content

Audio Transcription

Dusky can transcribe your interview audio in real time, giving you a live text feed of the conversation. The transcript also serves as context for AI-generated answers.

Dusky transcribes both sides of the call — the interviewer through system audio, and you through your microphone. Lines are labelled Interviewer and You in the transcript panel.

Press Cmd/Ctrl+Shift+A to toggle audio transcription on or off. A recording indicator appears in the Dusky toolbar when audio is active.

  • Microphone permission must be granted. On macOS, check System Settings > Privacy & Security > Microphone. On Windows, check Settings > Privacy & Security > Microphone. See Permissions troubleshooting if you need help.
  • Signed in to your Dusky account (transcription uses a server-issued token).

Press Cmd/Ctrl+Shift+T to toggle the transcript panel. Live captions appear as the conversation goes on.

When audio is active, pressing Cmd/Ctrl+Enter sends the transcript to AI instead of taking a screenshot. This is ideal for spoken questions — behavioral interviews, system design prompts, or any question the interviewer says out loud rather than showing on screen.

You can use both features at the same time. Dusky decides what to send based on the microphone state:

Mic stateTranscript available?Cmd/Ctrl+Enter behavior
OffCaptures screenshot, sends to AI
OnYesSends transcript to AI
OnNoNothing happens — the press is ignored

Dusky transcribes six languages, detected automatically:

English · Spanish · German · French · Portuguese · Italian

Detection happens per turn, so a call that switches between any of these — say, a French question followed by an English follow-up — is followed automatically, with nothing to set. Languages beyond these six aren’t transcribed live yet.

This is deliberate, not a missing feature. On the streaming transcription API Dusky uses, the language is a property of the speech model and is detected automatically; the setting that names a language exists only for pre-recorded files, not for live audio. A dropdown would be a control that cannot do what it appears to do, so Dusky does not ship one.

Open the menu and expand Meeting Audio: Auto-detect. It shows “Detected automatically.” and the current list of languages.

This is an information panel, not a control — there is nothing to select. The list is sent by Dusky’s backend rather than baked into the app, so it stays truthful if the underlying model changes.

Answer Language — a separate row in the menu — controls the language Dusky writes its answers in. It offers 22 options, far more than the six it can hear, and the two settings are independent: wanting Spanish answers in an English-language interview is a perfectly reasonable thing to do.

If you pick an answer language that Dusky cannot transcribe, the menu says so inline — for example, choosing Japanese shows:

日本語 answers are supported. Spoken 日本語 isn’t transcribed yet.

That is a statement of fact, not an error, and it blocks nothing — 15 of the 22 options in that menu show it.

Audio is streamed to our transcription service and discarded immediately after processing. Transcripts are stored only on your device.

Be aware of what “both sides” means for you personally: your own microphone audio is transcribed too, not just the interviewer’s. Anything you say out loud during a session — including asides and things not directed at the interviewer — goes through the same transcription path and lands in the same local transcript and session history. Logging out wipes that history; see Signing In & Your Account.

See Data Handling for full details.