Speech Synthesis Expression Control Using Contrastive Text Features
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional speech synthesis systems fail to accurately describe the expression state of generated speech, resulting in a poor presentation effect.
Innovation Solution
A method involving a contrastive learning module to obtain a reference description feature, construct a target description feature, and generate target speech content based on an input phoneme sequence to control the expression state.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If conventional speech synthesis systems are used, then speech content can be generated, but the expression state cannot be accurately described resulting in poor presentation effect
Solution Approach 1:
The patent segments the speech synthesis process into distinct feature extraction and expression control components. The reference description feature is separated from the target description feature, allowing independent optimization of each component's contribution to expression state accuracy.
Solution Approach 2:
The patent changes the parameter representation by introducing expression state features as additional dimensions in the speech synthesis process. This transforms the synthesis from basic phoneme generation to controlled expression generation through parameter expansion.
2Ease of operation
If expression state control is added to speech synthesis, then presentation effect is enhanced, but system complexity increases
Solution Approach 1:
The patent creates a universal framework where the reference description feature can be extracted from any speech content and applied to control expression states across different synthesis scenarios. This multi-functional approach handles various expression requirements through a unified mechanism.
Solution Approach 2:
The patent introduces description features as intermediary representations between input text and output speech. These features act as mediators that carry expression state information through the synthesis process without requiring direct complex interactions between all system components.
Data Source
AI summary
A method, an apparatus, a device, and a storage medium for speech synthesis are provided. A reference description feature corresponding to prompted speech content is obtained, the reference description feature includes a text encoding representation determined by processing the prompted speech content with a contrastive learning module, and the text encoding representation describes a first expression state of the prompted speech content. Based on the reference description feature, a target description feature for indicating a target expression state is constructed. Target speech content corresponding to the target expression state is generated based on an input phoneme sequence including the target description feature.


