# Local Whisper transcription test, version 1

ToolGlance, measured 2026-08-28. 24 excerpts / 12 speakers / 213.665 seconds /
569 reference words. Two local models, two speech conditions, one silence probe
per model: 98 completed first calls. No hosted-service test or overall winner.

## Frozen inputs

LibriSpeech clean/test via OpenSLR on Hugging Face, revision
71cacbfb7e2354c4226d01e70d77d5fca3d04ba1, clean/test/0000.parquet.
Parquet SHA2567113aa4c3cf963fb54697145719a7725f984c8836d1c494a554cbb9f1a017df0.
From 40 speakers/2620 clips, select first12 numeric speaker IDs and the first2
lexicographic utterance IDs for each. This is a convenience subset, not random
or representative. Original FLAC bytes unchanged. All selected results retained.

Add NumPy2.3.5 normal white noise to each clip, seed20260828 + selection index,
at10dB whole-clip RMS SNR. Store float32 WAV, uniformly peak-scale to0.99 only
if needed. This is synthetic white noise, not a real acoustic environment.
Add5seconds of digital silence, evaluated separately with VAD off.

Settings/inputs locked21:14:45UTC before inference; calls ended21:16:42UTC.
Original protocol SHA25629f7943d8b67070a8cdef254b8d301022002f561fb357b8b5a4552bb83206816.
Original raw SHA2561eff351998464824d5ae9656240a80d7dfdd630a5062cd9953e98efab9101aba.
This public methods description and portable wrapper were prepared afterward;
they do not change those frozen measurements.

## Runtime

Systran/faster-whisper-tiny.en revision0d3d19a32d3338f10357c0889762bd8d64bbdeba;
Systran/faster-whisper-base.en revision3d3d5dee26484f91867d81cb899cfcf72b96be6c.
faster-whisper1.2.1, CTranslate2 4.7.1; CPU int8,two threads,one worker,English,
beam5,temperature0,task transcribe,VAD false,condition_on_previous_text false.
No prompt,vocabulary hints or reference text supplied to model. Other defaults
are those of the pinned library. Every first output saved without correction.

Windows/Python3.12/Ryzen7 8845HS. Tiny before base; no warm-up or repeated time
trials, background load uncontrolled. Per-clip time includes decoding and lazy
segment consumption, excludes model loading/download. Not a portable speed claim.

## Scoring

Jiwer4.0.0 Compose: ToLowerCase,RemovePunctuation,RemoveMultipleSpaces,Strip,
ReduceToListOfListOfWords. No title,number,spelling,contraction or synonym mapping.
WER=(substitutions+deletions+insertions)/reference words, summed across clips per
condition/model. Not an unweighted average of clip rates. WER can exceed100%.
Silence is a separate emitted-word observation, not normal speech WER.

Primary edit totals: tiny clean27,noise60;base clean27,noise40;each denominator569.
Both emit You on silence. No population hallucination rate inferred.
15th/FIFTEENTH and not-so/NOT SO demonstrate orthographic normalization effects.
No human audio re-adjudication, confidence interval or statistical superiority.
No meetings, overlapping voices, diarization, speaker labels or non-English test.
Public benchmark overlap with model training is unknown. Not a novel corpus.

## Sources

https://www.openslr.org/12/
https://huggingface.co/datasets/openslr/librispeech_asr
https://github.com/SYSTRAN/faster-whisper/tree/v1.2.1
https://huggingface.co/Systran/faster-whisper-tiny.en
https://huggingface.co/Systran/faster-whisper-base.en
https://jitsi.github.io/jiwer/
https://creativecommons.org/licenses/by/4.0/
