Emotion Control in Speech Synthesis via Acoustic Feature Mapping

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current speech synthesis methods produce emotionless audio due to the lack of emotional sound libraries, which are difficult and inefficient to create, resulting in weak expressive force in synthesized speech.

Innovation Solution

A method that acquires text and specifies an emotion type, determines corresponding acoustic features, and inputs these into a pre-trained speech synthesis model trained without the specified emotion type, generating target audio with the desired emotional characteristics using existing corpora.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If an emotional sound library is created to enhance the expressive force of synthesized audio, then the emotion and expressiveness of the audio is improved, but the workload and time required for recording staff increases significantly

Engineering Contradiction:
Improveexpressive force of audioVSAvoidrecording efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent uses acoustic features extracted from emotional audio samples as templates or copies to guide the synthesis process. Instead of requiring recording staff to manually create emotional sound libraries, the system copies acoustic characteristics (pitch, energy, duration) from existing emotional audio to generate new emotional speech, dramatically reducing recording workload while maintaining emotional expressiveness

Inventive Principle:
Principle #26Copying

Solution Approach 2:

The patent replaces the mechanical process of manual emotional audio recording with an automated computational system. The speech synthesis model automatically generates emotional audio by controlling acoustic features through algorithms, substituting the mechanical recording process with an automated information processing system that eliminates the need for extensive manual recording work

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Reliability

If a sound library with emotion is created to make synthesized audio more expressive, then the emotion of the audio is improved, but the workload for recording staff increases heavily

Engineering Contradiction:
Improveemotion of audioVSAvoidworkload complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent extracts essential emotional characteristics from audio samples in the form of acoustic features (pitch contour, energy distribution, duration patterns). By taking out only the critical emotional elements rather than requiring complete emotional sound libraries, the system reduces the complexity of data preparation while preserving emotional expressiveness in the synthesized output

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent controls emotional expression by adjusting acoustic parameters (fundamental frequency, energy, duration) generated by the speech synthesis model. By changing these parameters according to acoustic features derived from emotional samples, the system achieves emotional audio output without requiring complex emotional sound libraries, thereby reducing workload complexity

Inventive Principle:
Principle #35Parameter changes

3Reliability

If emotional sound libraries are manually created to enhance audio expressiveness, then the expressiveness of synthesized speech is improved, but the time required for creation increases significantly

Engineering Contradiction:
Improveexpressiveness of speechVSAvoidcreation time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The patent performs preliminary extraction of acoustic features from emotional audio samples before the actual synthesis process. By pre-processing and storing acoustic characteristics (pitch, energy, duration patterns) in a structured format, the system prepares emotional expressions in advance, enabling rapid generation of emotional speech during synthesis without time-consuming manual recording during the actual creation phase

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system copies acoustic feature patterns from pre-processed emotional samples to guide the speech synthesis model. This copying mechanism allows rapid generation of emotional speech by reusing extracted acoustic characteristics, dramatically reducing the time required to create emotional audio compared to manual recording while maintaining high expressiveness

Inventive Principle:
Principle #26Copying

Data Source

PatentUS20230306954A1Speech synthesis method, apparatus, readable medium and electronic device
Publication Date: 2023.09.28 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US20230306954A1 patent drawing
  • US20230306954A1 patent drawing
  • US20230306954A1 patent drawing

AI summary

The present disclosure relates to a speech synthesis method, apparatus, readable medium and electronic device, which relates to the technical field of electronic information processing. The method comprises: acquiring a text to be synthesized and a specified emotion type (101), determining specified acoustic features corresponding to the specified emotion type (102), and inputting the text to be synthesized and the specified acoustic features into a pre-trained speech synthesis model, to acquire a target audio with the specified emotion type corresponding to the text to be synthesized which is output by the speech synthesis model (102). The acoustic features of the target audio match with the specified acoustic features, and the speech synthesis model is trained from a corpus without the specified emotion type.