Skip to content
Browse the docs

Transcription Accuracy

Improve poor accuracy, wrong-language detection, or missing words.

If your transcript has misheard words, the wrong language, missing names, or jumbled speakers, accuracy almost always comes down to a few fixable things: the audio quality reaching the engine, the hints you give it, and how speech is layered in the room. This page walks through each cause and the quickest fix.

Most accuracy problems are solved before you press record, by choosing the right audio mode, adding expected-language hints, and giving Kalima a little context about names and jargon. Use the checklist below to diagnose what you are seeing.

Start here: a quick checklist

Clean audio
Pick the right quality mode and a good mic so the engine hears clear speech.
Language hints
Tell Kalima which languages to expect so it stops guessing.
Session context
Add names and jargon up front so specialized terms come through right.
One voice at a time
Overlapping speech is the hardest case for any transcription engine.
Idea

The single biggest accuracy lever is clean input. A decent microphone close to the speaker in a quiet room, with the right quality mode, will outperform any after-the-fact correction. Fix the audio first, then add hints and context.

Diagnose by symptom

Background noise, echo, and a distant microphone are the most common causes of garbled text. Kalima offers two capture quality modes:

  • Enhanced mode turns on echo cancellation, noise suppression, and auto-gain. It is the default and helps a lot with a laptop mic or in loud or echoey spaces.
  • Raw mode records unprocessed audio and is best with a good microphone in a quiet room.

If your transcript is garbled, check that the quality mode matches the room you are actually in. Enhanced is the safer choice for almost every room. You can compare both on a sample with the Test modes option in the audio settings. Also move closer to the microphone and keep a consistent distance. See Audio Sources & Quality for the full breakdown.

Kalima detects language automatically, but detection is guided, not forced, it can stumble on accents, short phrases, or languages that sound similar. The fix is to add expected-language hints.

In the setup wizard's Languages step (or your session settings), search for and add each language you expect to be spoken. This steers recognition toward those languages without locking others out, so mid-conversation switches still work. Adding at least one hint noticeably improves results when you already know what will be spoken.

The picker offers about 60 languages, and you can add more than one. Learn more in Languages & Detection.

Generic recognition often misses people's names, company and product names, medical or legal terms, and other jargon. Give Kalima that vocabulary ahead of time with session context.

Open the session context dialog (or the Context step in setup) and add a short description of the session plus the specific terms and names you expect. This boosts recognition of exactly those words. The quick-description field holds up to 2,000 characters, but keep it concise, very long backgrounds have diminishing returns. You can also save reusable context templates for recurring meeting types. Full guidance is in Improve Accuracy with Context.

Kalima separates speakers automatically (Speaker 1, Speaker 2, and so on), but overlapping or very short speech is the hardest case for any engine, so it may occasionally misattribute a line. To reduce this:

  • Encourage one person to speak at a time, cross-talk is the main cause of mixed-up attribution.
  • Make sure each speaker is clearly audible to the microphone.

Once recording stops you can rename speakers to real names, and the numbering stays consistent across every take in the session. See Speakers & Diarization.

If you sit far from the microphone or use a quiet lapel mic, the engine may not hear you clearly. Adjust the recording gain before you record: open the gain settings (the gear next to the audio source) and boost or cut the level. The range runs from -12 dB to +12 dB in 1 dB steps, default 0, and your setting is remembered for next time. Nudge it up for quiet mics and down for hot, distorted input.

This is expected and not an accuracy problem. Live transcription shows a fast preliminary result first, then refines it into the confirmed version a moment later. Confirmed text can trail the live audio by up to about 2 seconds while the engine finishes each phrase. This brief delay actually improves the final wording by avoiding choppy, over-segmented results.

Kalima shows how confident the engine is in each word using subtle opacity, fainter words indicate lower confidence and are worth a second look. Treat it as a proofreading hint, not a guarantee. Faint words clustered together usually point to a noisy stretch of audio; use the fixes above (quality mode, mic distance, gain) for the next recording, and clean up the current transcript with AutoCorrect+ (below).

Kalima captures audio from the moment you press record and buffers it, so the brief window before the connection completes is minimized. If you still notice clipped opening words, pause for about half a second after pressing record before you start speaking.

Clean up an existing transcript

The fixes above improve future recordings. For a transcript you already have, use AutoCorrect+ (Mark & Fix) to repair the exact errors that matter.

1
In the transcript, select or right-click the misheard word or phrase and mark it with a category (wrong word, name/term, grammar, punctuation, filler to remove, or a note).
2
Open the AutoCorrect+ (Mark & Fix) panel to see all your open marks.
3
For each mark, type the correction yourself or ask AI to suggest one, then apply it.
4
For a name or term that recurs, tick "also fix others", AI checks each other occurrence and only replaces the ones that are genuinely the same mistake.

Corrections are applied at the word level, so playback timing and word-by-word highlighting stay aligned. Full details are in AutoCorrect+ (Mark & Fix).

Note

AutoCorrect+ uses AI and counts against your weekly autocorrections quota. AI fixes are minimal by design, they will not add words that were never spoken. If a transcript stretch was simply too noisy, re-recording with better audio beats correcting word by word.

When accuracy looks worse near a disconnect

If a section of the transcript is marked as recovered or arrived late, it was filled in by Kalima's gap recovery after a network interruption rather than transcribed live. That recovered text uses the same languages, translation, and context as your live session and is merged into the right place. If a gap shows as failed, use the retry option. See Connection & Recovery Issues for more.

Still seeing poor accuracy?

If you have a good microphone, the right quality mode, expected-language hints, and session context in place and accuracy is still poor, the problem is usually the source audio itself, heavy background noise, far-field microphones, or multiple people talking over each other. Re-record under cleaner conditions where you can, and contact support with your microphone model, browser or desktop app, and a short example if you would like a closer look.