A Toeplitz autocorrelation matrix cuts ACELP codebook search complexity and memory use while preserving speech perceptual quality.
Maps high-frequency speech power to lower mel bands and resynthesizes MFCC-based audio to improve clarity for hearing-impaired listeners.
Sequential phonetic and prosodic generation with an LM-TTS adaptor cuts speech latency while improving accuracy, style, and emotion.
A two-stage TTS workflow uses control vectors to adjust pitch, energy, and duration without extra training while preserving speech quality.
Performance metrics flag weak text sequences so speech models can request targeted actor recordings and cut training time and data use.
Cursor position and motion guide TTS pace and speech characteristics, improving reading alignment, engagement, and comprehension.
Adjustment dictionaries target repeated pronunciation and prosody errors in neural speech synthesis without repeated user input.
A two-phase, expressivity-scored training approach helps neural TTS generate more natural speech with stronger emotional expression.
Machine learning extracts noise and speech features to synthesize clean voice audio from mobile or non-studio recordings.
Neural spline flows model pitch, energy, and unvoiced regions to make text-to-speech output more natural and expressive.
Binned phoneme pitch vectors and neural prediction layers enable fine pitch tuning for more expressive and context-aware text-to-speech.
A neural TTS model skips less informative acoustic frames and interpolates them later to cut latency and bandwidth without degrading speech quality.
Thresholded and activated attention vectors improve text-to-speech alignment, yielding more natural timing and better non-speech synthesis with limited data.
Multi-stage embedding training enables natural synthetic voices while reducing traceability to real speakers and resisting noise defects.
A two-stage audio pipeline filters hallucinated events and shifts signals with impulse response and noise to build DAS-ready datasets.
By removing speaker and language cues during training and reapplying them at inference, this case improves multilingual speech quality without fine-tuning.
Role and emotion extraction guides reference text and audio selection, improving synthesis authenticity while reducing labeled-data training cost.
A 2D voice search and parameter-mixing approach builds a target-like synthetic voice when only minimal recordings are available.
A shared linguistic-acoustic prosody model adds contextual and speaker cues to generate more natural, engaging synthesized speech.
Simultaneous subband generation cuts neural vocoder forward propagations, enabling faster real-time speech waveform synthesis.
Separating timbre and emotion features cuts training data needs while preserving consistent voice quality and emotion intensity in synthesized speech.
Machine learning extracts target accent features and adapts live speech while preserving the speaker's natural voice for clearer conversations.