·14 min read·By Michael

How Parkbench Ensures Voice Quality, Part 2: Listening to the Mistakes

Catching how the voice sounds wrong, not just what it says.

Visualization of audio quality analysis and defect detection

Where We Left Off

Part 1 focused on what was missing. We were seeing truncated sentences, dropped endings, and cases where the autoregressive model simply stopped when it ran out of context. The solution was a verification loop: generate audio, transcribe it with Whisper, compare it to the original, and regenerate if anything important is missing.


That closed one class of failure.


It also made another one much easier to hear.


Once truncations are handled, smaller defects become more noticeable. Pops, hisses, elongations, and subtle distortions start to stand out. These are cases where the audio is perfectly transcribable and still clearly wrong to a listener.


This leads to a less comfortable but important reality. When working with stochastic generation at scale, glitches are not rare events. They are expected. The goal is not to eliminate them entirely, but to detect them, contain them, and be able to explain them when they occur.


A Failure That Passed Everything

A user reported a noise burst at exactly 7:20 in a meditation we had published that morning. This was not a missing word or a clipped sentence, but a short burst of high-frequency noise in the middle of otherwise clean audio.


When we inspected the waveform, we found roughly two seconds of energy that clearly did not belong, surrounded by normal speech. Nothing else in the file looked unusual, and none of our existing checks flagged it. Whisper transcribed the segment without issue, and every system we had in place would have passed the file. By our existing checks, it was a valid result: the words were all there, and the transcription matched, but we had no way of judging whether the audio itself was actually acceptable to listen to.


The Five Ways Audio Can Go Wrong

Once we started categorising failures by how they sound, a small taxonomy emerged:

  • Spikes and pops
    Single-sample glitches or exaggerated plosives.
  • High-frequency noise bursts
    Sustained hiss that does not belong, typically lasting one to three seconds.
  • Elongation (droning)
    A syllable stretches unnaturally as the model loses confidence.
  • Unexpected silence
    The model stalls and produces low-energy noise instead of speech.
  • Voice drift
    Over longer generations, the voice gradually shifts away from the intended identity.

Each of these produces a different signal pattern and requires a different detection approach. Each also introduces its own false positive risks.


The False-Positive Problem

Every detector involves a trade-off. If it is too aggressive, it begins to modify audio that was already correct.

A few examples we implemented and later rolled back:

Plosives flagged as spikes

An amplitude-based detector successfully caught real pops, but it also flagged natural consonants such as p, t, and k. The result was unnecessary modification of clean speech.

Silence as a fallback

Replacing failed segments with silence drew more attention than the original defect. We moved to omission instead. A short gap is generally less disruptive than an abrupt silence.

Sibilants flagged as noise

Early high-frequency detectors triggered on soft, breathy speech. The adjustment was to require contiguous energy rather than total energy. Noise bursts tend to be sustained, while speech is more variable.

Each detector is now evaluated with a simple question: what legitimate audio could look like this? If the answer is “some,” the detector is either refined further or not used.


Defense in Depth

No single check is sufficient. The system is structured as a set of layers:

Diagram showing defense-in-depth layers for audio quality checking

The aim is not to make any single layer perfect. Instead, failures should have to pass through multiple independent checks before reaching the final output.


Is It Still the Same Voice?

The most subtle failure mode in that stack, and the one that took the longest to measure reliably, is voice drift.


Each personality on Parkbench is anchored by a reference voice clip. The TTS model uses it to establish timbre, pitch range, and speaking style. Over short generations, that identity holds. Over longer ones, especially something like a 20-minute meditation composed of many chunks, it can begin to wander. The words remain correct and the pacing is consistent, but later chunks no longer quite sound like the same speaker as earlier ones.


This is difficult to detect because nothing is obviously broken. A careful listener will notice that something feels off, but it is hard to point to a specific issue.


The problem is that there is no ground truth for what a given chunk should sound like. Instead, we compare each generated chunk back to the personality's reference clip using MFCCs. These provide a compact summary of the voice's spectral characteristics that is stable within a single speaker and distinct across different speakers. By taking the cosine similarity between a chunk and the reference, we reduce this comparison to a single score.


In practice, a clean chunk in the intended voice scores around 0.90 or higher. Noticeable drift tends to fall below 0.85, and more severe failures drop below 0.75.


We initially set the failure threshold at 0.80, which only caught the more extreme cases. Tightening that threshold to 0.85 significantly improved detection of the more subtle cases that users notice but cannot easily describe. As with the other detectors, there is a trade-off. Speaking style, pacing, and content-specific reference clips can shift MFCC values in legitimate ways, and overly aggressive thresholds introduce false positives. The 0.85 threshold was chosen because it increased detection rates substantially while keeping false positives within an acceptable range across a broad validation set.


When a chunk falls below the threshold, it follows the same recovery path as other failures: retry generation, fall back to smaller segments, and omit if necessary. In practice, this results in a brief, almost unnoticeable pause rather than an extended stretch where the voice no longer sounds like itself.


Voice drift sits outside what transcription-based verification can catch. Whisper will faithfully transcribe the words regardless of who appears to be speaking. Detecting this class of failure required its own measurement, its own threshold, and the same discipline about how tightly to tune it.


Loudness Is a Content Decision

The same voice should not sound identical across all contexts.

Conversations typically sit around −16 LUFS, while meditations are closer to −20 LUFS. This difference is particularly noticeable when listening on headphones.

We moved to a content-aware loudness normalization process:

  • Measure perceptual loudness using LUFS
  • Adjust gain to a target level
  • Apply a peak limiter to prevent clipping

The previous approach relied on a simple RMS threshold. It identified obvious issues but left many outputs sounding too quiet. The updated system improves consistency without introducing distortion.


Every Chunk Has a Black Box

The most useful addition was not a new detector, but better observability.

Each generated chunk now records:

  • Input script
  • Reference voice
  • Output audio
  • QC results
  • Performance metrics

Failures also include full model logs.

A report such as “something sounds off at 7:20” can now be handled systematically:

  1. Locate the relevant chunk
  2. Listen to it in isolation
  3. Review QC results
  4. Inspect logs
  5. Regenerate or omit

This reduces investigation time from hours to minutes.


From Incidents to Trends

Per-chunk forensics answers a very specific question: what happened in this generation?


It does not answer a broader question that matters just as much: what has been happening across the system over time?


A single chunk failing its high-frequency noise check is not especially meaningful. The same check failing on 1.5% of chunks last week and 4% this week is a real signal. Something has shifted, and a user is likely to notice before we do unless we are actively looking for it.

To address this, we built a layer on top of the per-chunk forensics that focuses on aggregation, classification, and proactive notification.

Per-generation severity classification

Every completed generation is assigned a severity tier: clean, minor, major, or broken. This classification is derived from post-mix structural integrity, clipping checks, and chunk-level QC failures. Only generations above a certain threshold trigger an email to the engineering team. Most are clean and produce no notification.

Daily audio-quality digest

A scheduled job reviews all generations from the previous 24 hours, aggregates statistics by content type, and produces an HTML summary. This includes counts of total and clean generations, silence-region and spectral stability metrics, calmness scores for meditation content, and a shortlist of generations worth reviewing.

Severity-tiered notifications

A broken classification, which also blocks publishing, triggers a real-time alert. A degraded classification is included in the daily digest. A clean classification produces no notification. The system is designed to interrupt only when there is a clear reason to do so.

Calmness metrics roll-up

Meditation and sleep content receive a per-generation calmness score based on features such as crest factor, dynamic range, spectral centroid, and the loudness difference between counting cues and narration. Tracking these over time helps detect drift toward more energetic delivery, which can occur after model or reference changes.

The primary value of this layer is not just in investigating individual failures. It is in identifying patterns early, and in building a clear picture of what normal looks like so that deviations are immediately visible.


What This Taught Us

A few principles have emerged:

  • Prefer false negatives over false positives
  • Detection should lead to a clear corrective action
  • Omission is often better than substitution
  • Capture detailed information on failure, keep success paths lightweight
  • Quality emerges from multiple small, context-aware decisions
  • Observability without aggregation is just storage
    Per-incident forensics explains individual failures. Aggregated metrics reveal system-level drift. Both are necessary: one for investigation, the other for early warning.
  • Notify only when it is worth being interrupted
    Real-time alerts for broken, daily summaries for degraded, and silence for clean. Matching notification to severity is what keeps the system useful rather than noisy.

Glitches Will Happen

We do not aim to eliminate glitches entirely.

Instead, the system is designed to:

  • Catch most issues before delivery
  • Avoid publishing clearly broken output
  • Record what went wrong
  • Recover automatically when possible
  • Adapt to different content types

The goal is not perfection, but reliability and explainability.


The next set of challenges is already clear:

  • Emotional consistency across long audio
  • Voice drift detection
  • Improved balance between music and voice
  • Pronunciation validation

Each will be approached in the same way, as a set of incremental improvements rather than a single solution.


Parkbench generates personalised AI audio including meditation, sleep, conversations, and narration. Everything runs locally using open-source models. No user data ever leaves our infrastructure.