Learnable Speech Speed Control via Phoneme Context Encoding

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis methods rely on phone duration prediction models that control speech speed by applying a single factor to all phonemes, leading to unnatural speech generation, and lack the ability to control speed without parallel data in end-to-end models.

Innovation Solution

A method that encodes context associated with phonemes, aligns them to target acoustic frames, and recursively generates mel-spectrogram features using a computer system with processors and neural networks, allowing for speech synthesis at varying speeds without parallel data, mimicking human speech speed control.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of operation

If a single controlling factor is applied to all phonemes for speed control, then the system is simple to operate, but the speech generation becomes unnatural

Engineering Contradiction:
Improvespeed control simplicityVSAvoidspeech naturalness
Core Design Contradiction:
Ease of operationVSManufacturing precision

Solution Approach 1:

The patent segments the speech synthesis process into distinct components: phoneme encoding, context encoding, alignment to acoustic frames, and mel-spectrogram generation. Each component handles specific aspects of speech production independently, allowing for more precise control over different aspects of speech naturalness while maintaining operational simplicity through modular design.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent applies local quality by encoding context-specific information for each phoneme individually rather than applying a uniform controlling factor. Each phoneme's duration and characteristics are adjusted based on its local context (surrounding phonemes, position in utterance, stress patterns), resulting in more natural speech variation while maintaining overall system simplicity.

Inventive Principle:
Principle #3Local quality

2Ease of manufacture

If traditional phone duration prediction models are used for speed control, then the method is easy to implement, but the ability to control speed without parallel data is lost

Engineering Contradiction:
Improveimplementation simplicityVSAvoidspeed control flexibility
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent introduces an intermediary alignment step that maps phonemes to target acoustic frames. This intermediary layer acts as a bridge between the simple phoneme-based input and the complex acoustic output, enabling speed control through the alignment process itself rather than requiring complex parallel data models. The alignment mechanism provides the necessary adaptability while keeping the overall system relatively simple to implement.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent transitions from traditional one-dimensional phoneme duration control to a multi-dimensional approach by introducing acoustic frame alignment as an additional dimension. This allows speed control to be achieved through temporal scaling of the alignment process rather than simple duration multiplication, providing greater flexibility in speed control without requiring parallel data for training.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

3Manufacturing precision

If phonemes are aligned to target acoustic frames with context encoding, then speech synthesis quality improves, but computational complexity increases

Engineering Contradiction:
Improvespeech synthesis qualityVSAvoidsystem computational complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary encoding of phoneme contexts before the alignment and synthesis stages. By pre-computing context representations and storing them for later use, the system reduces the computational burden during the actual synthesis and alignment phases. This preliminary action maintains high speech synthesis quality while managing overall system complexity through efficient preprocessing.

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS11682379B2Learnable speed control of speech synthesis
Publication Date: 2023.06.20 TENCENT AMERICA LLC
  • US11682379B2 patent drawing
  • US11682379B2 patent drawing
  • US11682379B2 patent drawing

AI summary

A method, computer program, and computer system is provided for synthesizing speech at one or more speeds. A context associated with one or more phonemes corresponding to a speaking voice is encoded, and the one or more phonemes are aligned to one or more target acoustic frames based on the encoded context. One or more mel-spectrogram features are recursively generated from the aligned phonemes and target acoustic frames, and a voice sample corresponding to the speaking voice is synthesized using the generated mel-spectrogram features.