Learnable Speech Speed Control via Phoneme Context Encoding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current speech synthesis methods rely on phone duration prediction models that control speech speed by applying a single factor to all phonemes, leading to unnatural speech generation, and lack the ability to control speed without parallel data in end-to-end models.
Innovation Solution
A method that encodes context associated with phonemes, aligns them to target acoustic frames, and recursively generates mel-spectrogram features using a computer system with processors and neural networks, allowing for speech synthesis at varying speeds without parallel data, mimicking human speech speed control.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a single controlling factor is applied to all phonemes for speed control, then the system is simple to operate, but the speech generation becomes unnatural
Solution Approach 1:
The patent segments the speech synthesis process into distinct components: phoneme encoding, context encoding, alignment to acoustic frames, and mel-spectrogram generation. Each component handles specific aspects of speech production independently, allowing for more precise control over different aspects of speech naturalness while maintaining operational simplicity through modular design.
Solution Approach 2:
The patent applies local quality by encoding context-specific information for each phoneme individually rather than applying a uniform controlling factor. Each phoneme's duration and characteristics are adjusted based on its local context (surrounding phonemes, position in utterance, stress patterns), resulting in more natural speech variation while maintaining overall system simplicity.
2Ease of manufacture
If traditional phone duration prediction models are used for speed control, then the method is easy to implement, but the ability to control speed without parallel data is lost
Solution Approach 1:
The patent introduces an intermediary alignment step that maps phonemes to target acoustic frames. This intermediary layer acts as a bridge between the simple phoneme-based input and the complex acoustic output, enabling speed control through the alignment process itself rather than requiring complex parallel data models. The alignment mechanism provides the necessary adaptability while keeping the overall system relatively simple to implement.
Solution Approach 2:
The patent transitions from traditional one-dimensional phoneme duration control to a multi-dimensional approach by introducing acoustic frame alignment as an additional dimension. This allows speed control to be achieved through temporal scaling of the alignment process rather than simple duration multiplication, providing greater flexibility in speed control without requiring parallel data for training.
3Manufacturing precision
If phonemes are aligned to target acoustic frames with context encoding, then speech synthesis quality improves, but computational complexity increases
Solution Approach 1:
The patent performs preliminary encoding of phoneme contexts before the alignment and synthesis stages. By pre-computing context representations and storing them for later use, the system reduces the computational burden during the actual synthesis and alignment phases. This preliminary action maintains high speech synthesis quality while managing overall system complexity through efficient preprocessing.
Data Source
AI summary
A method, computer program, and computer system is provided for synthesizing speech at one or more speeds. A context associated with one or more phonemes corresponding to a speaking voice is encoded, and the one or more phonemes are aligned to one or more target acoustic frames based on the encoded context. One or more mel-spectrogram features are recursively generated from the aligned phonemes and target acoustic frames, and a voice sample corresponding to the speaking voice is synthesized using the generated mel-spectrogram features.


