Speaker Verification Using Co-location Data
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Speaker verification systems face challenges in accurately identifying enrolled users in multi-user environments, particularly in noisy conditions and when multiple potential impostors are present, leading to higher false acceptance rates due to the lack of sufficient information for decision-making.
Innovation Solution
The system enhances speaker verification by utilizing co-location information and sharing speaker models between devices to generate and normalize scores, thereby improving the accuracy of identifying the correct user through the use of multiple speaker models and scores from co-located devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If speaker verification systems use traditional single-device speaker models, then the system is simple to operate, but the measurement precision of speaker identification deteriorates in multi-user environments
Solution Approach 1:
The patent merges speaker models from multiple co-located devices to create a combined speaker model. This combination allows the system to leverage information from multiple devices simultaneously, improving speaker identification accuracy in multi-user environments by comparing the audio signal against aggregated speaker characteristics from all co-located devices rather than relying on a single device's speaker model.
Solution Approach 2:
The patent enables speaker verification systems to serve multiple functions: individual device speaker verification and multi-device collaborative verification. The system can operate in either mode depending on the situation, making it universally applicable to both simple single-user scenarios and complex multi-user environments, thereby improving measurement precision without permanently increasing system complexity.
2Speed
If speaker verification systems continuously listen for predefined phrases, then the system can quickly respond to user input, but the false acceptance rate increases when multiple potential impostors are present
Solution Approach 1:
The patent introduces co-location information as an intermediary factor in the speaker verification decision process. Before making a verification decision, the system checks whether the detected speaker matches the expected user based on co-location data from multiple devices. This intermediary check reduces false acceptance rates by providing an additional layer of verification that distinguishes between legitimate users and impostors in the vicinity.
Solution Approach 2:
The patent performs preliminary verification using co-location information before final speaker verification. By pre-establishing which devices are co-located and what users should be associated with each device, the system can quickly eliminate impostors who are physically present but not authorized, thereby maintaining fast response times while improving reliability.
3Reliability
If speaker verification systems use multiple speaker models from co-located devices, then the false acceptance rate decreases, but the loss of information increases due to the need to share and normalize data between devices
Solution Approach 1:
The patent extracts only the necessary co-location information and speaker model data from individual devices, rather than sharing complete device states. By extracting and transmitting only the relevant speaker characteristics and co-location metadata, the system reduces data sharing overhead while still enabling improved verification accuracy through multiple speaker models.
Data Source
AI summary
A computer-implemented method that includes receiving audio data corresponding to an utterance of a voice command captured by a user device. The user device has a plurality of different users. The method includes determining a particular user among the plurality of different users of the user device as a speaker of the utterance based on a comparison between the audio data and corresponding speaker verification data stored on memory hardware for each user of the plurality of different users of the user device. The method further includes, based on determining the particular user among the plurality of different users of the user device as the speaker of the utterance, providing, for output from the user device, a message comprising a speaker identifier associated with the particular user.


