Emotion-Aware Speech Synthesis Using Context and Semantic Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional text-to-text (TTS) processing systems fail to effectively convey the emotional intent and interactive meaning of text, resulting in speech outputs that lack emotional depth and realism, failing to engage users in interactive conversations.

Innovation Solution

A speech synthesis method that generates emotion information based on situation explanation information and context analysis, combining metadata with text to produce speech that reflects emotional attributes such as 'neutral', 'love', 'happy', 'anger', 'sad', 'worry', and 'sorry', using deep neural network learning and emotion expression models to create realistic emotional outputs.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If conventional TTS processing is used to deliver semantic contents, then text-to-speech conversion is achieved, but emotional intent and interactive meaning are lost

Engineering Contradiction:
Improveemotional intentVSAvoidprocessing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system segments the text processing into multiple stages: semantic analysis to extract meaning, context analysis to understand situation, and emotion information generation to derive emotional attributes. Each stage processes specific aspects independently before combining results, preventing information loss while maintaining manageable complexity through modular processing.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary semantic and context analysis on the input text before speech synthesis. By pre-processing the text to extract emotion-relevant features and situation information, the system prepares emotion information in advance that will be integrated during synthesis, ensuring emotional intent is preserved without adding complexity to the core TTS engine.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If emotion information is generated through semantic and context analysis, then speech with emotional attributes is produced, but processing time and computational resources increase

Engineering Contradiction:
Improveemotional accuracyVSAvoidprocessing time
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system applies partial action by selectively analyzing only the portions of text that contain emotional or contextual significance. Rather than processing every word uniformly, the system identifies key semantic elements and context indicators that contribute to emotion determination, performing analysis only where necessary to achieve accurate emotional representation without unnecessary computational overhead.

Inventive Principle:
Principle #16Partial or excessive action

Solution Approach 2:

The system changes parameters by adjusting the depth and intensity of semantic and context analysis based on the specific input characteristics. For different types of text or emotional contexts, the system can modify analysis parameters to optimize between accuracy and processing speed, achieving reliable emotion detection while adapting computational resource usage to the task requirements.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11074904B2Speech synthesis method and apparatus based on emotion information
Publication Date: 2021.07.27 LG ELECTRONICS INC
  • US11074904B2 patent drawing
  • US11074904B2 patent drawing
  • US11074904B2 patent drawing

AI summary

A speech synthesis method and apparatus based on emotion information are disclosed. A speech synthesis method based on emotion information extracts speech synthesis target text from received data and determines whether the received data includes situation explanation information. First metadata corresponding to first emotion information is generated on the basis of the situation explanation information. When the extracted data does not include situation explanation information, second metadata corresponding to second emotion information generated on the basis of semantic analysis and context analysis is generated. One of the first metadata and the second metadata is added to the speech synthesis target text to synthesize speech corresponding to the extracted data. A speech synthesis apparatus of this disclosure may be associated with an artificial intelligence module, drone (unmanned aerial vehicle, UAV), robot, augmented reality (AR) devices, virtual reality (VR) devices, devices related to 5G services, and the like.