Personalized Text-to-Speech Synthesis with Incidental Voice Capture
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Text-to-speech synthesis systems fail to personalize the voice of the sender, resulting in indistinguishable voices for multiple users and loss of emotional expressiveness in instant messaging and other communication environments.
Innovation Solution
A method and system that receive incidental audio input from a sender during communication, generate a voice dataset, and synthesize text input to produce personalized speech that mimics the sender's voice, incorporating emotional and visual expressions through voice morphing and image synthesis.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If TTS synthesis is used to convert text to speech in communication systems, then text can be converted to audible speech output, but the synthesized speech loses the sender's identity and all users sound the same
Solution Approach 1:
The system performs preliminary voice training by capturing audio samples from the sender during communication sessions before text-to-speech conversion is needed. This advance preparation creates a personalized voice model that can be applied later when text needs to be synthesized, thereby preserving the sender's identity without interfering with the actual text conversion process.
Solution Approach 2:
The system creates a copy or model of the sender's voice characteristics by analyzing audio samples and generating a voice profile. This copied voice model is then used to synthesize speech from text, allowing the system to reproduce the sender's unique vocal patterns, pitch, and timbre without requiring the sender to be physically present.
2Loss of information
If TTS synthesis converts text to speech, then text content can be delivered audibly, but emotional expressiveness and vocal nuances are lost
Solution Approach 1:
The system incorporates feedback mechanisms by analyzing the sender's audio samples for emotional cues, tone variations, and expressive patterns. This feedback information is then used to adjust and refine the synthesized speech output, ensuring that emotional expressiveness is preserved and accurately conveyed in the converted speech.
Solution Approach 2:
The system dynamically changes multiple speech parameters including pitch, speed, volume, and tone during synthesis to match the sender's emotional state and vocal characteristics. By adjusting these parameters based on the analyzed audio samples, the system can convey emotional nuances and maintain the expressive quality of the original communication.
3Ease of manufacture
If voice recording modules require users to repeat words for training, then personalized speech can be synthesized, but the process becomes time-consuming and intrusive
Solution Approach 1:
The system enables self-service voice training by automatically capturing audio samples from the sender's normal communication without requiring active participation or repetition of words. The sender's voice is recorded incidentally during regular use, and the system automatically processes these samples to create the voice model, eliminating the need for dedicated training sessions.
Solution Approach 2:
The system performs voice training in advance by continuously or periodically capturing audio samples during communication sessions. This preliminary action ensures that the voice model is ready before text-to-speech conversion is needed, eliminating the need for time-consuming on-demand voice recording and processing.
4Ease of operation
If incidental audio input is used for voice training, then voice personalization can be achieved without dedicated input, but the system complexity increases
Solution Approach 1:
The system achieves multi-functionality by using the same audio input hardware and processing pipelines for both normal communication and voice training purposes. The audio recording module serves dual functions: capturing communication content and collecting voice samples for personalization, thereby avoiding the need for separate dedicated voice input hardware or processing systems.
Solution Approach 2:
The system merges the voice training function with the existing communication framework by integrating audio capture and processing into the regular communication flow. Instead of separate voice recording and training modules, the system combines these functions with the existing audio communication infrastructure, reducing overall system complexity while enabling incidental voice capture.
Data Source
AI summary
A method and system are provided for text-to-speech synthesis with personalized voice. The method includes receiving an incidental audio input (403) of speech in the form of an audio communication from an input speaker (401) and generating a voice dataset (404) for the input speaker (401). The method includes receiving a text input (411) at the same device as the audio input (403) and synthesizing (312) the text from the text input (411) to synthesized speech including using the voice dataset (404) to personalize the synthesized speech to sound like the input speaker (401). In addition, the method includes analyzing (316) the text for expression and adding the expression (315) to the synthesized speech. The audio communication may be part of a video communication (453) and the audio input (403) may have an associated visual input (455) of an image of the input speaker. The synthesis from text may include providing a synthesized image personalized to look like the image of the input speaker with expressions added from the visual input (455).


