Synthesized Speech Emotion Correction via Feedback Loop
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for generating synthesized speech using text and emotion information lack verification of whether the generated speech accurately reflects the intended emotion, leading to inconsistencies and difficulties in ensuring the desired emotion is expressed.
Innovation Solution
A method that uses a deep learning model to generate synthesized speech by comparing the emotion information vector of the input text with the emotion vector extracted from the synthesized speech, determining if correction is needed based on a preconfigured threshold, and adjusting the emotion vector to ensure the generated speech aligns with the intended emotion.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If emotion information is applied as input for generating synthesized speech, then the synthesized speech should reflect the intended emotion, but it is not possible to verify whether the emotion information is duly expressed
Solution Approach 1:
The patent implements a feedback mechanism by extracting emotion information from the generated synthesized speech and comparing it with the input emotion information. A loss value is calculated to quantify the difference, and if this loss exceeds a threshold, the emotion information is corrected and the speech is regenerated. This closed-loop feedback system ensures verification of emotion expression without requiring complex manual verification processes.
2Measurement precision
If emotion information correction is implemented by comparing loss values and regenerating speech, then the emotion accuracy is improved, but the processing time and computational resources increase
Solution Approach 1:
The patent introduces a loss value parameter that quantifies the difference between input and extracted emotion information. By comparing this parameter against a threshold, the system determines whether correction is necessary. This parameter-based approach enables efficient decision-making about regeneration needs, balancing accuracy requirements with processing time constraints.
Solution Approach 2:
The system does not always regenerate speech when emotion information is provided. Instead, it selectively regenerates only when the loss value exceeds the threshold, performing partial action rather than exhaustive regeneration in all cases. This optimizes processing time by avoiding unnecessary regenerations while still ensuring accuracy when needed.
3Manufacturing precision
If the loss value threshold is set low for strict emotion verification, then the emotion accuracy is improved, but the number of regenerations and processing time increase
Solution Approach 1:
The patent allows the threshold parameter to be configured based on specific application requirements. By adjusting this parameter, users can balance between precision and efficiency: a lower threshold provides stricter verification for high-precision applications, while a higher threshold improves efficiency for applications where approximate emotion expression is acceptable. This parameter adaptability resolves the contradiction between precision and productivity.
Data Source
AI summary
A method includes generating first synthesized speech by using text and a first emotion vector configured for the text, extracting a second emotion vector included in the first synthesized speech, determining whether correction of the second emotion information vector is needed by comparing a loss value calculated by using the first emotion information vector and the second emotion information vector with a preconfigured threshold, re-performing speech synthesis by using a third emotion information vector generated by correcting the second emotion information vector, and outputting the generated synthesized speech, thereby configuring emotion information of speech in a more effective manner. A speech synthesis apparatus may be associated with an artificial intelligence module, drone (unmanned aerial vehicle, UAV), robot, augmented reality (AR) devices, virtual reality (VR) devices, devices related to 5G services, and the like.


