MeetingScribe
In researchReal-time transcription of in-person meetings · macOS
MeetingScribe is a macOS app that transcribes in-person meetings in real time. It rests on one premise: audio never leaves the device. Meeting content is not the kind of data that should be uploaded to an external server, so all recognition runs on the device. The app is not yet a finished tool. It is at the stage of measuring how far on-device Korean speech recognition is actually usable.
Screenshots
Screenshots coming soon
Keeping the LLM outside the app
The app handles only audio, transcription, display and files. Organizing, correcting and judging are done by Claude, outside the app, by exchanging files.
This boundary has three consequences.
- 1 Audio is never sent outside the device.
- 2 Changing how notes are organized does not require releasing the app again. Only the instructions given to the side that reads the files need to change.
- 3 If the app terminates abnormally, its output remains as files.
Recognition quality
Apple's on-device Korean recognition is too inaccurate to use for meeting minutes. The same Korean audio was run through two recognizers for comparison.
일단 지금 거 아직 안 드렸어가지고 우유가 안 고 그죠?
Gloss, approximate: "First, since I haven't given you the current one yet, the milk isn't [unintelligible], right?" The Korean output is garbled and has no coherent meaning.
일단 지금 보여드리는 요거를 제가 아직 안보내드려서 아직 공유가 안된 상태이고요. 그쵸?
Gloss: "First, I haven't sent you the thing I'm showing you now yet, so it hasn't been shared yet. Right?"
Every setting that can be adjusted within the app was tried. The results of the full measurement are listed below.
| Attempted | Result |
|---|---|
| Glossary injection contextualStrings | No effect. Output with 0 entries and with 811 entries was identical byte for byte |
| DictationTranscriber | One quarter of the character count. Intended only for short dictation |
| SpeechTranscriber, 5 presets | Same as the current configuration |
| ContentHint.farField | No effect for ko-KR |
| alternativeTranscriptions | The candidates are nearly identical to the original, so they provide no usable signal |
| AGC target level | Already optimal (×4.6; optimal range ×4 to ×8) |
| Low-frequency cutoff | Every value from 60 to 250 Hz is within the noise range |
The remaining options are replacing the recognizer and improving the audio input. Replacement has been measured. Running Whisper on 3-second chunks gives a worst-case latency of 4.4 seconds, a 1.0 GB quantized model and 1.47 GB of resident memory. Audio input depends on conditions at recording time, so it can be verified only in real meetings.
Limits of live-microphone measurement
Quality experiments use recorded audio only. Every live run hears different sound, so conditions cannot be compared with a live microphone. An early A/B test run without this in mind stopped at the conclusion that the problem was input level, not the options. For that reason each session keeps the original audio from before amplification and re-runs that file under different conditions.
Release on hold
The app is not in a state where it can be downloaded and used as is. As described above, Korean recognition quality is still insufficient for meeting minutes, and offering it for download in that state would not be appropriate. The distribution method will be decided again after the recognizer has been replaced.