Voice Model Generation for Hands-Free Sender Identification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current messaging systems are limited in hands-free or eyes-free settings, leading to disruptions and safety concerns as users need to interact multiple times to identify senders in group chats, especially in environments where attention cannot be diverted from current activities.
Innovation Solution
A voice model is generated using audio samples from users to create a synthetic audio output that resembles their voice, allowing for immediate identification of senders through a personalized and familiar interaction, reducing the need for multiple communications to identify the sender.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If text messaging is used in hands-free or eyes-free settings, then communication can occur without visual attention, but user safety is compromised and interaction delays increase due to multiple back-and-forth communications needed to identify senders
Solution Approach 1:
The system performs preliminary action by generating and storing voice models of users in advance. When a message is received, the voice model enables immediate identification of the sender through voice recognition, eliminating the need for multiple interactive back-and-forth communications and allowing users to maintain attention on safety-critical tasks while still receiving and processing messages
Solution Approach 2:
The voice model acts as an intermediary between the text message and the user. Instead of requiring direct visual interaction with the messaging interface, the voice model synthesizes speech that conveys sender identity and message content, allowing users to receive information without diverting attention from current activities
2Loss of information
If text messaging requires multiple interactions to identify senders in group chats, then sender identification can be achieved, but time delay and cognitive delay increase
Solution Approach 1:
Voice models are generated and stored in advance for each user. When a message arrives in a group chat, the system immediately matches the sender's voice model against the incoming message, providing instant sender identification without requiring multiple interactive steps or cognitive processing delays
Solution Approach 2:
The system provides immediate feedback by synthesizing voice output that clearly identifies the sender. This feedback mechanism eliminates the need for users to manually query or interact multiple times to determine who sent a message, reducing both time delay and cognitive delay
3Ease of operation
If generic text-to-speech is used for message delivery, then audio output can be generated, but personalization and sender identification are lost
Solution Approach 1:
Instead of using a single generic text-to-speech system, the patent implements local quality by creating personalized voice models for each individual user. Each voice model captures the unique characteristics of that user's voice, allowing the system to generate audio messages that not only convey information but also clearly identify the sender through their distinctive voice characteristics
Solution Approach 2:
The system creates accurate voice copies of each user by training voice models on their audio samples. These voice copies are then used to synthesize messages that sound like the actual sender, providing personalization and clear sender identification while maintaining the convenience of automated audio message delivery
Data Source
AI summary
Disclosed herein a system, a method and a device for generating a voice model for a user. A device can include an encoder and a decoder to generate a voice model for converting text to an audio output that resembles a voice of the person sending respective text. The encoder can includes a neural network and can receive a plurality of audio samples from a user. The encoder can generate a sequence of values and provide the sequence of values to the decoder. The decoder can establish, using the sequence of values and one or more speaker embeddings of the user, a voice model corresponding to the plurality of audio samples of the user.


