Emotion Classification Text-to-Speech Semantic Analysis

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional Text-To-Speech (TTS) systems fail to effectively convey the intent and emotion of a user's message, lacking the ability to analyze semantic content and context to produce emotionally nuanced speech outputs.

Innovation Solution

An emotion classification information-based TTS method that determines whether emotion metadata is set in a message, using semantic and context analysis to generate metadata for the speech synthesis engine, allowing it to produce speech that reflects the intended emotion, even if no explicit emotion classification is provided.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If conventional TTS processing is used to transmit semantic contents, then the basic text-to-speech function is achieved, but the ability to convey user intent and emotion is lost

Engineering Contradiction:
Improveemotion informationVSAvoidprocessing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system segments the TTS processing into distinct modules: a semantic analysis module that extracts emotion information from text, a context analysis module that considers conversational context, and a speech synthesis module that generates emotionally nuanced speech. This segmentation allows each module to specialize in one aspect of emotion detection and speech generation, reducing overall system complexity while preserving emotion information.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The system performs preliminary semantic and context analysis on the input text before speech synthesis to pre-determine the appropriate emotion classification. By preparing emotion metadata in advance through analysis of the message content and context, the system ensures emotion information is captured before the speech generation process begins, preventing loss of emotional nuance.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If emotion classification information is manually set in messages, then speech with intended emotion can be produced, but user operation complexity increases

Engineering Contradiction:
Improveemotion accuracyVSAvoiduser operation
Core Design Contradiction:
ReliabilityVSEase of operation

Solution Approach 1:

The system automatically performs semantic and context analysis on user messages to self-determine the appropriate emotion classification without requiring manual user input. The semantic analysis module extracts emotional cues from the message content, while the context analysis module considers the conversational context, allowing the system to autonomously generate accurate emotion metadata that reflects the user's intended emotion.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system uses feedback from semantic analysis results and context analysis results to automatically adjust and determine the most appropriate emotion classification. By analyzing the message content and comparing it against known emotional patterns and contextual cues, the system refines its emotion determination to accurately reflect user intent without manual intervention.

Inventive Principle:
Principle #23Feedback

3Measurement precision

If semantic analysis and context analysis are performed to generate emotion metadata, then speech synthesis accuracy is improved, but processing time increases

Engineering Contradiction:
Improveemotion detection precisionVSAvoidprocessing time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The system performs partial semantic analysis and context analysis focused specifically on extracting emotion-relevant features from the message, rather than comprehensive analysis of all text characteristics. By concentrating computational resources on emotion-specific linguistic cues and contextual indicators, the system achieves sufficient emotion detection precision without the time cost of exhaustive analysis.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11514886B2Emotion classification information-based text-to-speech (TTS) method and apparatus
Publication Date: 2022.11.29 LG ELECTRONICS INC
  • US11514886B2 patent drawing
  • US11514886B2 patent drawing
  • US11514886B2 patent drawing

AI summary

Disclosed are an emotion classification information-based text-to-speech (TTS) method and device. The emotion classification information-based TTS method according to an embodiment of the present invention may, when emotion classification information is set in a received message, transmit metadata corresponding to the set emotion classification information to a speech synthesis engine and, when no emotion classification information is set in the received message, generate new emotion classification information through semantic analysis and context analysis of sentences in the received message and transmit the metadata to the speech synthesis engine. The speech synthesis engine may perform speech synthesis by carrying emotion classification information based on the transmitted metadata.