BaseToolbox LogoBaseToolbox
Blog

© 2025 BaseToolbox. All rights reserved.

Privacy PolicyAboutContact Us
  1. Home
  2. /
  3. Local Audio to Text

Local Audio to Text

Transcribe office recordings, voice memos, interviews, or microphone audio entirely in your browser with a shared lazy-loaded Whisper model. Review timestamped segments and export TXT, Markdown, SRT, or VTT.

Saved sessions

Add an office recording

Up to 250 MiB and 30 minutes on this device. Actual codec support depends on the browser.

Detect language
Whisper loads only when transcription starts. This MVP is not live transcription and does not promise speaker diarization. Review names, numbers, dates, and technical terms against the recording.
Local AI runtime
Checking this device
BackendWASM
Device profile—
Adaptive limits—
Writing model
Qwen 0.5B · standard
Transcription model
Whisper Tiny · fast
Files, prompts, and results stay in this browser. Browser storage can still be cleared by you, private mode, or browser cleanup.
Private browser-local transcription

Audio decoding, resampling, Whisper inference, timestamps, saved sessions, and exports run in your browser. Recordings and transcript text are excluded from BaseToolbox analytics.

Current limitations
  • The browser must be able to decode the file's audio codec; support varies for MP3, WAV, M4A, AAC, OGG, WebM, and MP4 containers.
  • The MVP is batch transcription, not guaranteed real-time transcription, and does not promise speaker diarization.
  • Recognition accuracy varies with noise, accents, overlapping speakers, names, numbers, and specialist vocabulary.
  • Verified model weights normally stay cached after refresh, but clearing site data, private browsing, model-version changes, integrity failures, or browser eviction can require another download.

How local recording transcription works

1

Upload or record audio

Choose a browser-decodable file or grant microphone permission and record an office voice note.

2

Decode and transcribe locally

The browser mixes channels, resamples to 16 kHz mono, and lazily loads the selected shared Whisper tier through the global inference queue.

3

Review, export, or create notes

Check timestamped text, export TXT, Markdown, SRT, or VTT, or hand the transcript to a separate meeting-notes session.

Local audio transcription FAQ

Is my recording uploaded?
Is this the same as FilmJoy's subtitle or lyrics workflow?
Does the stronger transcription model belong only to this page?
Will refreshing download Whisper again?

Related office writing tools

Create decisions and action items from a meeting transcriptSummarize text with source references