Making Streaming Voice Work on CPUs
How a smaller model, quantized CPU inference, and a concurrent sentence pipeline turned batch-only TTS into continuous conversational speech.

Generating audio in the background is relatively forgiving. A personality can create a meditation, narrate a journal entry, or produce a longer wellness session without it mattering very much whether ten seconds of audio takes fifteen or twenty seconds to render. The work happens in the background, and the finished audio is delivered when it is ready.
Spoken conversation has a different constraint. We want playback to begin while the response is still being produced: the language model generates text, complete sentences are passed to text-to-speech, and the browser begins playing them without waiting for the full response. For that to work, speech has to be generated quickly enough that playback does not catch up with the synthesizer.
On our production CPU infrastructure, NeuTTS Air could take more than twenty seconds to generate ten seconds of speech. That is inconvenient but manageable for a batch job. In a conversation, it means the playback buffer eventually empties and the voice stops while another piece of audio is being prepared. We had already built a streaming implementation around Air, but its performance meant we could not sensibly run it in production.
Trying more hardware
We initially treated this as an infrastructure problem and tested NeuTTS Air on an AWS g4dn.xlarge with an NVIDIA T4. We tried the normal float32 path, float16 and torch.compile, but none produced the improvement we needed. Float16 introduced numerical instability in the attention layers, torch.compile did not materially improve throughput, and moving inference to the T4 still did not bring Air close enough to real-time synthesis to justify dedicated GPU infrastructure.
That result became easier to understand when we looked at the workload itself. Air is autoregressive, so its audio tokens are generated sequentially, with each token depending on those before it. A GPU can still help with parts of that computation, but the model cannot take advantage of parallelism in the same way as a workload in which a much larger portion of the calculation can happen simultaneously.
Our existing CPU path also had advantages. It used a quantized GGUF model and inference software designed around sequential token generation, so moving the same model to a GPU did not automatically produce a better system.
We tested larger CPU instances as well. Our normal TTS workers use m5.xlarge instances with 4 vCPUs, while a separate test used a c5.4xlarge with 16 vCPUs. A representative ten-second clip fell from roughly 22 seconds of generation time to around 15–18 seconds. That improvement was measurable but still well short of what streaming required, and the larger instance would have added roughly $280 per month to infrastructure that costs us around $2,000 per month overall.
We therefore left the streaming implementation in place but stopped running it as a production feature. The endpoint, playback code and deployment machinery still existed; the synthesis workload simply did not fit the amount of infrastructure we were prepared to allocate to it.
Moving from Air to Nano
Neuphonic released NeuTTS Nano on January 14, 2026. Nano retained the parts of Air that mattered to us, including voice cloning and GGUF support, while reducing the active parameter count from 360 million to 120 million. Neuphonic also reported roughly a 1.8x improvement in CPU backbone throughput, which made it worth testing against our own workloads.
Before comparing the two models, we established a broader Air baseline on our production hardware using six voices and eight content types, for 48 inference runs in total. We measured Real-Time Factor, or RTF, which is the amount of generation time required for a given duration of finished audio. An RTF of 2 means that one second of audio takes two seconds to produce, while anything below 1 means synthesis is running faster than playback.
| Voice | Mean RTF | Characteristic |
|---|---|---|
calm_storyteller | 1.955 | Clear, measured delivery |
golden_voice | 2.089 | Warm, mid-pace |
dynamic_host_british | 2.524 | Energetic, varied cadence |
soul_narrator | 3.086 | Deep, deliberate |
southern_comfort | 3.256 | Slow, rich tone |
midnight_poet | 3.980 | Dramatic, drawn phrasing |
| Mean | 2.815 | Range: 1.855–6.553 |
Air averaged 2.815 RTF across the sweep. On the same m5.xlarge class used by our 4-vCPU workers, Nano Q4 brought the mean down to roughly 1.5 RTF, making it close to twice as fast as our Air baseline on comparable hardware.
An RTF of 1.5 is still slower than real time if an entire response is treated as a single generation job. Conversation gives us another option, though, because the response is already arriving incrementally from the language model.
Streaming as a pipeline
For streaming speech, the useful question is not whether the entire response can be synthesized faster than its total playback duration. What matters is whether enough audio can be prepared ahead of the listener to keep the playback buffer supplied.
We also allocate more CPU to interactive speech than to batch generation. Priority and background TTS workers run on m5.xlarge instances with 4 vCPUs and 16 GB of memory, while the streaming worker uses an m5.2xlarge with 8 vCPUs and 32 GB. On that worker, Nano can reach around 0.75 RTF with eight threads when we measure the inference workload itself.
Inference is only one part of the path to the browser. Once the model has generated audio tokens, NeuCodec still has to decode them into a waveform, and there is additional work in request handling, sentence processing, networking and buffering. When we measure more of the complete path, steady-state sentence generation is more commonly around 1.27–1.39 RTF.
The two sets of figures describe different boundaries around the system. The 0.75 measurement is useful for understanding the model on the dedicated worker, while the higher figures include more of the work required to produce playable audio. Neither figure by itself determines whether the conversation can sound continuous, because synthesis and playback do not have to happen serially.
Our earlier implementation effectively did that: generate sentence one, play it, then move on to sentence two. The resulting pauses were noticeable. The current pipeline allows two TTS requests to be in flight concurrently, so while one sentence is playing, the next may already be complete and another can be generating.
Concurrency introduces an ordering problem because later requests occasionally finish first. Completed audio is therefore buffered until its position in the response reaches the front of the playback queue. The benefit is that synthesis can continue while earlier audio is playing, so an end-to-end RTF slightly above 1 does not necessarily translate into gaps in the conversation.
Where the latency remains
The largest remaining problem is at the beginning of a spoken response. The first TTS request on a fresh worker is noticeably slower than subsequent ones and, in our measurements, can take around twice the steady-state RTF. We have not isolated that difference to a single part of the runtime, so for now we treat it as an observed cold-start characteristic rather than attributing it to a particular cache or library.
TTS is also not the first component on the critical path. Our self-hosted language model runs through Ollama on a g5.xlarge GPU worker, and on a cold path its time to first token can be around 6.5 seconds. Speech synthesis cannot begin until enough text has arrived to form the first sentence.
A recent production request looked like this:
| Event | Time | Duration |
|---|---|---|
| User sends message | 0s | |
| LLM produces first token | 6.5s | 6.5s |
| First sentence detected | 7.3s | 0.8s later |
| TTS begins sentence 1 | 7.3s | |
| Sentence 1 ready | 19.7s | 12.4s |
| First audio begins | 19.7s | |
| Sentence 2 ready | 25.8s | ~6.1s |
| Sentence 3 ready | 35.6s | ~9.8s |
| Playback finishes | 43.4s | Continuous after start |
A roughly twenty-second wait for the first audio is still longer than we want. The important distinction from the earlier Air implementation is that the delay is now concentrated near the beginning of the request. Once playback starts, later sentences can be prepared quickly enough that the stream continues without repeatedly stopping to wait for synthesis.
There is a related timing problem in the interface. The language model can produce written text much faster than the voice can speak it, which meant an early implementation could display several sentences ahead of the audio. Someone might be reading sentence four while still hearing sentence two.
We now reveal text according to audio readiness rather than directly exposing the timing of the language model. The model continues generating normally, but the interface releases sentences in coordination with the speech pipeline. Text generation and audio playback operate at very different rates, so keeping them aligned turns out to be part of the streaming problem as well.
Speech quality
The performance improvement would have been less useful if Nano introduced an obvious reduction in speech quality. Across meditation, wellness, conversation and journal content, however, we found it broadly comparable to Air. Air retained a little more warmth in some samples, but the difference was small enough that it did not present a practical problem for us.
We also stopped seeing one failure mode that had appeared often enough with Air to require special handling. Certain voices would occasionally produce elongated vowels, turning a short word into several seconds of sustained audio. Our audio pipeline includes automated checks for suspicious generations partly because these failures were unpredictable.
We did not encounter the same behavior during our Nano evaluation using similar text and voice references. That does not establish that Nano cannot produce it, or that the smaller model is responsible for its absence; it is simply what we observed in this test set. We have kept the same quality checks in place rather than treating the model change as a reason to remove them.
The production arrangement
Nano now runs across the Parkbench audio system. Background workers generate longer material such as meditations, wellness sessions and journal narration, while priority workers handle user-initiated jobs. Conversational speech is handled separately by the streaming worker.
Kubernetes currently places those workloads across three nodes. The priority and background deployments use the same neutts-cpu-nodes node group, but anti-affinity prevents them from occupying the same node, so each currently lands on an m5.xlarge. The streaming deployment explicitly selects an m5.2xlarge with 8 vCPUs and 32 GB of memory, which also prevents a long background generation from consuming CPU reserved for an active conversation.
The rest of the audio pipeline continues to operate in the same way. Whisper verification, spike detection, truncation repair and our other checks all run against Nano output because they were built around properties of the finished audio rather than around a specific synthesis model.
Keeping the earlier streaming implementation was useful for a similar reason. When Nano appeared, the endpoint, deployment configuration, playback logic and model-selection machinery were already available. Evaluating another model was largely a matter of determining whether its performance changed the constraints enough for that existing design to work.
Measuring the system rather than the model
One lesson from this work is that there is no particularly useful single RTF for the whole system. We have a comparison between Air and Nano on 4-vCPU workers, a narrower inference measurement on the 8-vCPU streaming worker, a broader measurement that includes decoding and request overhead, and finally the latency experienced before audio actually reaches the browser. Those figures answer related but different questions.
Neuphonic's backbone benchmark told us that Nano was worth evaluating, while our 4-vCPU tests showed how it compared with Air on hardware similar to our existing workers. The 8-vCPU measurements showed what additional CPU capacity could provide, and the end-to-end timings made clear how much work remained outside model inference itself.
The earlier T4 and larger-CPU experiments were useful even though neither produced a configuration we wanted to run. They showed that Air's performance could not be changed enough through infrastructure alone, at least within the resources we were prepared to devote to each active conversation.
The configuration we use now combines a smaller autoregressive model with quantized CPU inference, an 8-vCPU worker for interactive speech, sentence-sized requests, concurrent generation and ordered playback. The interface then has to coordinate that audio with text being generated on a different clock.
None of those parts on its own makes streaming voice work. Nano mainly reduced the amount of computation enough that the rest of the pipeline could keep pace with playback on infrastructure that was reasonable for us to operate. For this workload, that turned out to matter more than running the larger model.
Parkbench is an AI companion platform. This post is based on our production experience running language models, text-to-speech, and audio generation systems.