Synthesized Speech Emotion Correction via Feedback Loop

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current methods for generating synthesized speech using text and emotion information lack verification of whether the generated speech accurately reflects the intended emotion, leading to inconsistencies and difficulties in ensuring the desired emotion is expressed.

Innovation Solution

A method that uses a deep learning model to generate synthesized speech by comparing the emotion information vector of the input text with the emotion vector extracted from the synthesized speech, determining if correction is needed based on a preconfigured threshold, and adjusting the emotion vector to ensure the generated speech aligns with the intended emotion.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If emotion information is applied as input for generating synthesized speech, then the synthesized speech should reflect the intended emotion, but it is not possible to verify whether the emotion information is duly expressed

Engineering Contradiction:
Improveemotion expression accuracyVSAvoidverification system complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent implements a feedback mechanism by extracting emotion information from the generated synthesized speech and comparing it with the input emotion information. A loss value is calculated to quantify the difference, and if this loss exceeds a threshold, the emotion information is corrected and the speech is regenerated. This closed-loop feedback system ensures verification of emotion expression without requiring complex manual verification processes.

Inventive Principle:
Principle #23Feedback

2Measurement precision

If emotion information correction is implemented by comparing loss values and regenerating speech, then the emotion accuracy is improved, but the processing time and computational resources increase

Engineering Contradiction:
Improveemotion information accuracyVSAvoidspeech generation time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent introduces a loss value parameter that quantifies the difference between input and extracted emotion information. By comparing this parameter against a threshold, the system determines whether correction is necessary. This parameter-based approach enables efficient decision-making about regeneration needs, balancing accuracy requirements with processing time constraints.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system does not always regenerate speech when emotion information is provided. Instead, it selectively regenerates only when the loss value exceeds the threshold, performing partial action rather than exhaustive regeneration in all cases. This optimizes processing time by avoiding unnecessary regenerations while still ensuring accuracy when needed.

Inventive Principle:
Principle #16Partial or excessive action

3Manufacturing precision

If the loss value threshold is set low for strict emotion verification, then the emotion accuracy is improved, but the number of regenerations and processing time increase

Engineering Contradiction:
Improveemotion information precisionVSAvoidspeech generation efficiency
Core Design Contradiction:
Manufacturing precisionVSProductivity

Solution Approach 1:

The patent allows the threshold parameter to be configured based on specific application requirements. By adjusting this parameter, users can balance between precision and efficiency: a lower threshold provides stricter verification for high-precision applications, while a higher threshold improves efficiency for applications where approximate emotion expression is acceptable. This parameter adaptability resolves the contradiction between precision and productivity.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11636845B2Method for synthesized speech generation using emotion information correction and apparatus
Publication Date: 2023.04.25 LG ELECTRONICS INC
  • US11636845B2 patent drawing
  • US11636845B2 patent drawing
  • US11636845B2 patent drawing

AI summary

A method includes generating first synthesized speech by using text and a first emotion vector configured for the text, extracting a second emotion vector included in the first synthesized speech, determining whether correction of the second emotion information vector is needed by comparing a loss value calculated by using the first emotion information vector and the second emotion information vector with a preconfigured threshold, re-performing speech synthesis by using a third emotion information vector generated by correcting the second emotion information vector, and outputting the generated synthesized speech, thereby configuring emotion information of speech in a more effective manner. A speech synthesis apparatus may be associated with an artificial intelligence module, drone (unmanned aerial vehicle, UAV), robot, augmented reality (AR) devices, virtual reality (VR) devices, devices related to 5G services, and the like.