Emotion-Targeted Speech Synthesis via Candidate Selection

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current computer speech synthesis systems lack the ability to generate voice with emotional content, resulting in unnatural sounding speech that degrades the user experience.

Innovation Solution

The system generates a plurality of emotionally diverse candidate speech segments for a given message using crowd-sourcing and machine learning algorithms, which are then selected based on predetermined emotion types, allowing for the inclusion of emotional content in speech output.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Ease of manufacture

If straight text-to-speech conversion is used, then the speech generation is simple and efficient, but the emotional content is absent making the voice sound unnatural

Engineering Contradiction:
Improvespeech generation simplicityVSAvoidnaturalness of voice
Core Design Contradiction:
Ease of manufactureVSReliability

Solution Approach 1:

The speech generation process is segmented into multiple candidate speech segments, each representing different emotional interpretations of the same text. This segmentation allows the system to explore multiple emotional possibilities without complicating the overall generation pipeline, resolving the contradiction between simplicity and naturalness.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Multiple candidate speech segments with different emotional contents are generated in advance as preliminary options. This preliminary action enables the system to pre-compute diverse emotional expressions, which are then selected based on the target emotion type, maintaining efficiency while improving naturalness.

Inventive Principle:
Principle #10Preliminary action

2Reliability

If multiple candidate speech segments are generated to include emotional content, then the naturalness and emotional expression are improved, but the system complexity increases

Engineering Contradiction:
Improveemotional content accuracyVSAvoidsystem structure
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

Multiple candidate speech segments are generated in advance and stored, eliminating the need for complex real-time emotional synthesis. This preliminary preparation reduces online computational complexity while maintaining high emotional content accuracy during actual speech generation.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

Instead of generating entirely new speech segments with emotional content, the system creates multiple candidate versions (copies) of the same text with different emotional characteristics. This copying approach maintains system simplicity while providing diverse emotional expressions for selection.

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If crowd-sourcing is used to generate candidate speech segments, then the diversity of emotional content is enhanced, but the time and resources required for generation increase

Engineering Contradiction:
Improveemotional content diversityVSAvoidcandidate generation time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

Candidate speech segments with diverse emotional content are generated in advance through crowd-sourcing and stored for future use. This preliminary generation eliminates the need for real-time crowd-sourcing, reducing time loss during actual speech generation while maintaining emotional diversity.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates multiple candidate copies of speech segments with different emotional interpretations, collected through crowd-sourcing. These pre-generated copies provide diverse emotional content without requiring ongoing crowd-sourcing efforts, balancing versatility with time efficiency.

Inventive Principle:
Principle #26Copying

Data Source

PatentUS10803850B2Voice generation with predetermined emotion type
Publication Date: 2020.10.13 MICROSOFT TECHNOLOGY LICENSING LLC
  • US10803850B2 patent drawing
  • US10803850B2 patent drawing
  • US10803850B2 patent drawing

AI summary

Techniques for generating voice with predetermined emotion type. In an aspect, semantic content and emotion type are separately specified for a speech segment to be generated. A candidate generation module generates a plurality of emotionally diverse candidate speech segments, wherein each candidate has the specified semantic content. A candidate selection module identifies an optimal candidate from amongst the plurality of candidate speech segments, wherein the optimal candidate most closely corresponds to the predetermined emotion type. In further aspects, crowd-sourcing techniques may be applied to generate the plurality of speech output candidates associated with a given semantic content, and machine-learning techniques may be applied to derive parameters for a real-time algorithm for the candidate selection module.