Skip to content

Recording & transcription

Boswell records two audio sources at once and transcribes them on your Mac using Apple’s on-device speech recognition. Nothing is uploaded unless you explicitly configure a cloud LLM.

When you start a recording, Boswell listens to:

  • Microphone — your voice and any in-room audio picked up by your mic.
  • System audio — whatever is playing through your Mac’s speakers: Zoom, Meet, Teams, a podcast, a voice memo, or any other source.

Both are captured simultaneously, with no virtual audio driver required. If you’re in a call, the other participants see nothing — no bot, no extra guest in the meeting, no screen-sharing artifact.

Transcription runs on your Mac using Apple’s built-in, on-device speech recognition — the same speech engine macOS uses system-wide. There’s no model to choose or download: macOS manages the on-device speech model for you, and your audio becomes text without ever leaving your Mac.

Transcription starts immediately — you don’t wait until the end. While you’re in the meeting, Boswell is already turning the audio into text in the background. There’s no progress bar to watch.

Boswell records your microphone and the system audio as two separate sides, so the transcript keeps your side of the conversation distinct from the other side — everyone coming through your speakers. The two streams are captured and transcribed independently, then shown together in order.

Click Stop in the Boswell window (or use the menu-bar icon’s Stop item) and recording ends. Boswell then turns the recording into a note. See Enhanced Notes for what it contains.

By default, no audio or text leaves your Mac — recording and transcription run on-device, with no network connection required. Optional AI features (like connecting your own cloud provider for summaries) reach the network only if you turn them on, and then go straight to that provider — never through Boswell’s servers.

See What stays on-device for the full picture.