Emotion Speech Synthesis with Continuous Intensity Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing speech synthesis technologies struggle to generate speeches with varying emotion intensities, as they rely on data-driven methods that fail to capture the continuous nature of human emotion perception, resulting in synthesized speech that lacks emotional intensity continuity.
Innovation Solution
A method is introduced to generate a continuous emotion intensity feature vector set by extracting acoustic statistical features from emotion speech audio, allowing for the synthesis of speeches with adjustable and diverse emotion intensities using machine learning models.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If data-driven methods are used to synthesize emotion speech by collecting speech data and constructing acoustic parameter models for each emotion type, then highly realistic and natural emotion speech can be generated, but the system cannot continuously adjust emotion intensity levels
Solution Approach 1:
The patent extracts acoustic statistical feature parameters from original emotion speech audio to generate a continuous emotion intensity feature vector set, where each specific emotion intensity corresponds to a parameter value in the set. This allows continuous adjustment of emotion intensity by modifying the feature vector parameters, resolving the contradiction between achieving emotion intensity continuity and maintaining system complexity.
2Ease of operation
If simple classification is used to mark emotion intensity at a few discrete levels, then modeling can be performed for each level, but the synthesized speech cannot reflect the continuous nature of actual emotion perception
Solution Approach 1:
The patent transforms the static discrete emotion intensity levels into a dynamic continuous system. By extracting acoustic statistical features and generating a continuous feature vector set, the system can dynamically adjust emotion intensity to any value within the continuous range, making the synthesized speech reflect the continuous nature of actual emotion perception while maintaining ease of operation.
Data Source
Figure 1A
Figure 1B
Figure 2
AI summary
This application provides an emotion speech synthesis method and device. In the method, an emotion intensity feature vector is set for a target synthesis text, an acoustic feature vector corresponding to an emotion intensity is generated based on the emotion intensity feature vector by using an acoustic model, and a speech corresponding to the emotion intensity is synthesized based on the acoustic feature vector. The emotion intensity feature vector is continuously adjustable, and emotion speeches of different intensities can be generated based on values of different emotion intensity feature vectors, so that emotion types of a synthesized speech are more diversified. This application may be applied to a human-computer interaction process in the artificial intelligence (AI) field, to perform intelligent emotion speech synthesis.