The Work
Keep the input stable
The capture path uses one persistent SoX process to stream raw PCM audio into Python. Python handles level detection and splitting the stream into recordings. The design avoids repeatedly opening the capture device and avoids relying on a SoX silence effect that the project notes identify as problematic with virtual audio devices.
The recorder checks the environment and supports a continuous-monitor mode as well as an interactive shell. Configuration controls capture settings, output locations, and the transcription model. I kept those choices outside the main processing logic so the capture setup can change without rewriting the whole pipeline.
Compare alternatives for the same recording
The original audio is preserved. The pipeline processes copies with configured SoX profiles, then runs faster-whisper on each version. The documented configuration includes baseline, denasal, and narrowband approaches. Different preprocessing choices can help or hurt a recording, so selection happens after transcription rather than assuming one filter always wins.
The implementation records preprocessing and transcription failures alongside successful attempts. When debugging is enabled, it saves the processed audio, transcripts, and JSON metrics. That makes it possible to compare what the model heard across profiles rather than judge only the final text. Multiple attempts also add processing cost, so this is a quality-oriented design tradeoff.
Treat confidence as a signal, not proof
The scoring function combines average log probability, no-speech probability, compression ratio, and text plausibility checks. It penalizes repeated phrases, known unwanted output patterns, and implausible word rates. The highest-scoring attempt is selected, with a minimum-score check for heavily penalized results.
Domain corrections address common warehouse vocabulary and number formatting before the result is appended to the daily log. These rules can make the text more useful, but they are still heuristics. A confident model can be wrong, and a vocabulary substitution can also be wrong. The retained source and debug artifacts support review when the output needs to be checked.
| Stage | Processing decision | What remains available for review |
|---|---|---|
| Capture | Keep one SoX stream open and segment it in Python | Original audio recordings |
| Preprocess | Run configured profiles on copies of the same recording | Processed audio and failures when debugging is enabled |
| Transcribe and score | Apply number and domain corrections, then compare model signals, repetition, and plausible word rates | Alternative transcripts and JSON metrics when debugging is enabled |
| Select and log | Choose the highest-scoring attempt, check its minimum score, and append accepted text | Daily log plus source audio for checking uncertain results |
Keep a path for learning from labeled examples
The recorder includes a guided labeling mode, and a companion training script reads labeled audio-text pairs for LoRA fine-tuning. The training path is designed for a constrained GPU and includes conversion for the inference runtime. That creates a way to adapt the model if enough good examples are available.
I connected capture, model integration, domain handling, and result selection into one workflow. Preserving the audio and alternative attempts makes it possible to investigate a missed term or incorrect number and identify examples that could improve later training.
THE DELIVERABLES
What I delivered
- Continuous audio capture and segmentation pipeline
- Configurable preprocessing and transcription attempts
- Confidence and unwanted-output screening with daily logs
- Debug artifacts, labeling mode, and a LoRA training path
A CLOSER LOOK
Project exhibits
Route the audio before processing it
Voicemeeter Banana is the third-party audio-routing tool used in the capture setup. Its input and routing controls support the recording stage of the pipeline.

THE OUTCOME
Where the work stands
The pipeline produced searchable text from radio recordings with configurable preprocessing and domain handling. It provides a way to inspect alternative attempts when transcription quality is uncertain.
SCOPE AND LIMITATIONS
What the work establishes
Transcription accuracy and the effect of fine-tuning have not been measured against a labeled evaluation set.
Confidence and heuristic checks cannot guarantee a correct transcript.
