Emotion Speech Synthesis with Continuous Intensity Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing speech synthesis technologies struggle to generate speeches with varying emotion intensities, as they rely on data-driven methods that fail to capture the continuous nature of human emotion perception, resulting in synthesized speech that lacks emotional intensity continuity.

Innovation Solution

A method is introduced to generate a continuous emotion intensity feature vector set by extracting acoustic statistical features from emotion speech audio, allowing for the synthesis of speeches with adjustable and diverse emotion intensities using machine learning models.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Manufacturing precision

If data-driven methods are used to synthesize emotion speech by collecting speech data and constructing acoustic parameter models for each emotion type, then highly realistic and natural emotion speech can be generated, but the system cannot continuously adjust emotion intensity levels

Engineering Contradiction:
Improveemotion intensity continuityVSAvoidsystem complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent extracts acoustic statistical feature parameters from original emotion speech audio to generate a continuous emotion intensity feature vector set, where each specific emotion intensity corresponds to a parameter value in the set. This allows continuous adjustment of emotion intensity by modifying the feature vector parameters, resolving the contradiction between achieving emotion intensity continuity and maintaining system complexity.

Inventive Principle:
Principle #35Parameter changes

2Ease of operation

If simple classification is used to mark emotion intensity at a few discrete levels, then modeling can be performed for each level, but the synthesized speech cannot reflect the continuous nature of actual emotion perception

Engineering Contradiction:
Improveemotion intensity adjustment flexibilityVSAvoidemotion intensity accuracy
Core Design Contradiction:
Ease of operationVSMeasurement precision

Solution Approach 1:

The patent transforms the static discrete emotion intensity levels into a dynamic continuous system. By extracting acoustic statistical features and generating a continuous feature vector set, the system can dynamically adjust emotion intensity to any value within the continuous range, making the synthesized speech reflect the continuous nature of actual emotion perception while maintaining ease of operation.

Inventive Principle:
Principle #15Dynamics

Data Source

PatentEP3859731B1Speech synthesis method
Publication Date: 2025.12.10 HUAWEI TECH CO LTD
  • EP3859731B1 patent drawingFigure 1A
  • EP3859731B1 patent drawingFigure 1B
  • EP3859731B1 patent drawingFigure 2

AI summary

This application provides an emotion speech synthesis method and device. In the method, an emotion intensity feature vector is set for a target synthesis text, an acoustic feature vector corresponding to an emotion intensity is generated based on the emotion intensity feature vector by using an acoustic model, and a speech corresponding to the emotion intensity is synthesized based on the acoustic feature vector. The emotion intensity feature vector is continuously adjustable, and emotion speeches of different intensities can be generated based on values of different emotion intensity feature vectors, so that emotion types of a synthesized speech are more diversified. This application may be applied to a human-computer interaction process in the artificial intelligence (AI) field, to perform intelligent emotion speech synthesis.