Personalized Text-to-Speech Voice Synthesis for Message Differentiation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current text-to-voice synthesis systems use a single pre-selected voice for all incoming messages, leading to monotonous tone, difficulty in message interpretation, confusion between different messages, and failure to represent the sender's personality, affecting user response.
Innovation Solution
A communication device and method that synthesize an output voice using voice characteristic information associated with the sender, such as speaking speed, pitch, and volume, to audibly present text messages, allowing differentiation between messages and representing the sender's personality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If a single pre-selected voice is used for all incoming messages, then the system is simple to operate, but the tone becomes monotonous and message interpretation becomes difficult
Solution Approach 1:
The system changes voice parameters (pitch, speaking rate, volume) based on the sender's characteristics and message content. The processor retrieves voice characteristic information associated with each sender and dynamically adjusts synthesis parameters to generate distinct voices for different senders, thereby preventing monotony while maintaining ease of operation.
2Device complexity
If a single voice presents all messages, then the device complexity is reduced, but the user cannot distinguish between different messages
Solution Approach 1:
The system uses voice characteristic information stored in memory associated with each sender to dynamically change voice parameters during message presentation. The processor retrieves the appropriate voice characteristics and adjusts pitch, speaking rate, and other parameters to create distinct voices for different senders, enabling message differentiation without increasing device complexity.
3Quantity of substance
If a single pre-selected voice is used, then the system requires minimal data storage, but the sender's personality is not represented vocally
Solution Approach 1:
The system stores compact voice characteristic information (pitch, speaking rate, volume levels) associated with each sender in memory. When a message is received, the processor retrieves these parameters and applies them to the text-to-speech synthesis to generate a voice that reflects the sender's personality characteristics, thereby representing sender identity without requiring extensive data storage.
Data Source
AI summary
A communication device and method are provided for audibly outputting a received text message to a user, the text message being received from a sender. A text message to present audibly is received. An output voice to present the text message is retrieved, wherein the output voice is synthesized using predefined voice characteristic information to represent the sender's voice. The output voice is used to audibly present the text message to the user.


