Text-to-Speech Synthesis with Emotional Context Prediction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current text-to-speech synthesis systems lack a humane touch, producing mechanical voices that fail to convey emotions and context, particularly in messages with emotional undertones, leading to an unnatural reading experience.

Innovation Solution

A method and system that predict the emotional state of a sender's voice based on intermediate emotional states and receiver context, using a processor to synthesize customized natural voice by analyzing present and previous messages, responses, and receiver context, incorporating emotional analysis and Bayesian neural networks to assign weightages and generate a final emotional vector for voice synthesis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Productivity

If automated text-to-speech synthesis is used, then text can be converted to speech efficiently, but the voice output sounds mechanical and lacks emotional expression

Engineering Contradiction:
Improvetext-to-speech conversion efficiencyVSAvoidnaturalness of voice output
Core Design Contradiction:
ProductivityVSEase of manufacture

Solution Approach 1:

The system dynamically changes voice parameters including pitch, tone, speed, and volume based on the predicted emotional state of the sender. This allows the same text to be rendered with different emotional expressions by adjusting these acoustic parameters, thereby resolving the contradiction between efficient automated synthesis and natural emotional expression.

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The system performs preliminary emotional analysis of the text message and predicts the sender's emotional state before generating the voice output. This preliminary emotional context analysis enables the synthesis engine to pre-configure appropriate voice parameters, ensuring natural emotional expression is built into the synthesis process from the start rather than added afterward.

Inventive Principle:
Principle #10Preliminary action

2Ease of manufacture

If emotional context analysis is incorporated into TTS synthesis, then emotional expression in voice output is improved, but system complexity increases

Engineering Contradiction:
Improveemotional expression capabilityVSAvoidsystem architecture complexity
Core Design Contradiction:
Ease of manufactureVSDevice complexity

Solution Approach 1:

The system introduces an intermediary emotional analysis module that acts as a bridge between the text input and the voice synthesis engine. This module predicts the sender's emotional state and transforms textual information into emotional context parameters, which then guide the voice synthesis process. This intermediary approach adds emotional expression capability while keeping the overall system architecture modular and manageable.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The system incorporates feedback mechanisms where the predicted emotional state and receiver context continuously influence the voice parameter adjustments during synthesis. This feedback loop ensures that the emotional expression remains consistent and appropriate throughout the voice output, achieving natural emotional conveyance without requiring overly complex manual control systems.

Inventive Principle:
Principle #23Feedback

3Ease of manufacture

If voice parameters are dynamically adjusted based on emotional state, then naturalness of voice is improved, but computational requirements increase

Engineering Contradiction:
Improvenaturalness of voiceVSAvoidcomputational energy consumption
Core Design Contradiction:
Ease of manufactureVSUse of energy by moving object

Solution Approach 1:

The system applies partial emotional analysis by focusing on key emotional indicators and dominant emotions rather than analyzing every aspect of the text in equal detail. This selective approach allows the system to achieve sufficient emotional expression in the voice output while reducing unnecessary computational overhead and energy consumption.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS10424288B2System and method for rendering textual messages using customized natural voice
Publication Date: 2019.09.24 WIPRO LTD
  • US10424288B2 patent drawing
  • US10424288B2 patent drawing
  • US10424288B2 patent drawing

AI summary

This disclosure relates generally to the text-to-speech synthesis and more particularly to a system and method for rendering textual messages using customized natural voice. In one embodiment, a system for rendering textual messages using customized natural voice, is disclosed, comprising a processor and a memory communicatively coupled to the processor. The memory stores processor instructions, which, on execution, causes the processor to receive present textual messages and at least one of previous textual messages, response to the previous textual messages or receiver's context. The processor further predicts final emotional state of sender's customized natural voice based on an intermediate emotional state and the receiver's context. The processor further synthesizes the sender's customized natural voice based on the predicted final emotional state of the sender's customized natural voice, voice samples and voice parameters associated with the at least one sender.