Personalized Text-to-Speech Synthesis with Incidental Voice Capture

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Text-to-speech synthesis systems fail to personalize the voice of the sender, resulting in indistinguishable voices for multiple users and loss of emotional expressiveness in instant messaging and other communication environments.

Innovation Solution

A method and system that receive incidental audio input from a sender during communication, generate a voice dataset, and synthesize text input to produce personalized speech that mimics the sender's voice, incorporating emotional and visual expressions through voice morphing and image synthesis.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Loss of information

If TTS synthesis is used to convert text to speech in communication systems, then text can be converted to audible speech output, but the synthesized speech loses the sender's identity and all users sound the same

Engineering Contradiction:
Improvesender's identityVSAvoidvoice personalization capability
Core Design Contradiction:
Loss of informationVSEase of operation

Solution Approach 1:

The system performs preliminary voice training by capturing audio samples from the sender during communication sessions before text-to-speech conversion is needed. This advance preparation creates a personalized voice model that can be applied later when text needs to be synthesized, thereby preserving the sender's identity without interfering with the actual text conversion process.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates a copy or model of the sender's voice characteristics by analyzing audio samples and generating a voice profile. This copied voice model is then used to synthesize speech from text, allowing the system to reproduce the sender's unique vocal patterns, pitch, and timbre without requiring the sender to be physically present.

Inventive Principle:
Principle #26Copying

2Loss of information

If TTS synthesis converts text to speech, then text content can be delivered audibly, but emotional expressiveness and vocal nuances are lost

Engineering Contradiction:
Improveemotional expressivenessVSAvoidemotional accuracy in synthesis
Core Design Contradiction:
Loss of informationVSManufacturing precision

Solution Approach 1:

The system incorporates feedback mechanisms by analyzing the sender's audio samples for emotional cues, tone variations, and expressive patterns. This feedback information is then used to adjust and refine the synthesized speech output, ensuring that emotional expressiveness is preserved and accurately conveyed in the converted speech.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system dynamically changes multiple speech parameters including pitch, speed, volume, and tone during synthesis to match the sender's emotional state and vocal characteristics. By adjusting these parameters based on the analyzed audio samples, the system can convey emotional nuances and maintain the expressive quality of the original communication.

Inventive Principle:
Principle #35Parameter changes

3Ease of manufacture

If voice recording modules require users to repeat words for training, then personalized speech can be synthesized, but the process becomes time-consuming and intrusive

Engineering Contradiction:
Improvevoice training processVSAvoidtime for voice input
Core Design Contradiction:
Ease of manufactureVSLoss of time

Solution Approach 1:

The system enables self-service voice training by automatically capturing audio samples from the sender's normal communication without requiring active participation or repetition of words. The sender's voice is recorded incidentally during regular use, and the system automatically processes these samples to create the voice model, eliminating the need for dedicated training sessions.

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The system performs voice training in advance by continuously or periodically capturing audio samples during communication sessions. This preliminary action ensures that the voice model is ready before text-to-speech conversion is needed, eliminating the need for time-consuming on-demand voice recording and processing.

Inventive Principle:
Principle #10Preliminary action

4Ease of operation

If incidental audio input is used for voice training, then voice personalization can be achieved without dedicated input, but the system complexity increases

Engineering Contradiction:
Improveincidental voice captureVSAvoidaudio processing system
Core Design Contradiction:
Ease of operationVSDevice complexity

Solution Approach 1:

The system achieves multi-functionality by using the same audio input hardware and processing pipelines for both normal communication and voice training purposes. The audio recording module serves dual functions: capturing communication content and collecting voice samples for personalization, thereby avoiding the need for separate dedicated voice input hardware or processing systems.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The system merges the voice training function with the existing communication framework by integrating audio capture and processing into the regular communication flow. Instead of separate voice recording and training modules, the system combines these functions with the existing audio communication infrastructure, reducing overall system complexity while enabling incidental voice capture.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS9368102B2Method and system for text-to-speech synthesis with personalized voice
Publication Date: 2016.06.14 CERENCE OPERATING CO
  • US9368102B2 patent drawing
  • US9368102B2 patent drawing
  • US9368102B2 patent drawing

AI summary

A method and system are provided for text-to-speech synthesis with personalized voice. The method includes receiving an incidental audio input (403) of speech in the form of an audio communication from an input speaker (401) and generating a voice dataset (404) for the input speaker (401). The method includes receiving a text input (411) at the same device as the audio input (403) and synthesizing (312) the text from the text input (411) to synthesized speech including using the voice dataset (404) to personalize the synthesized speech to sound like the input speaker (401). In addition, the method includes analyzing (316) the text for expression and adding the expression (315) to the synthesized speech. The audio communication may be part of a video communication (453) and the audio input (403) may have an associated visual input (455) of an image of the input speaker. The synthesis from text may include providing a synthesized image personalized to look like the image of the input speaker with expressions added from the visual input (455).