Speaker Verification Using Co-Location Models

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Speaker verification systems face challenges in distinguishing between enrolled and non-enrolled users, particularly in multi-user environments where impostors are frequently present, leading to high false acceptance rates due to the lack of sufficient information for decision-making.

Innovation Solution

The implementation of a method that utilizes co-location information and shares speaker models between devices to improve verification decisions, generating and normalizing scores from multiple speaker models to enhance the accuracy of identifying the correct user.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If speaker verification systems use traditional single-device speaker models, then the system is simple to operate, but the false acceptance rate increases in multi-user environments

Engineering Contradiction:
Improvefalse acceptance rateVSAvoidsystem complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges speaker models from multiple co-located devices to create a collective verification system. Instead of relying on a single device's speaker model, the system combines models from multiple devices in the same physical location, thereby improving reliability by reducing false acceptance rates while managing complexity through coordinated multi-device operation

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent introduces co-location detection as an intermediary mechanism that identifies when multiple devices are present in the same physical space. This intermediary layer enables the system to dynamically activate multiple speaker models only when needed, rather than always using all models, thus improving reliability without proportionally increasing complexity

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If speaker verification systems consider multiple speaker models from co-located devices, then verification accuracy improves, but information processing requirements increase

Engineering Contradiction:
Improveverification accuracyVSAvoidinformation processing load
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

The patent performs preliminary co-location detection and speaker model retrieval before the actual verification process. By identifying co-located devices and preparing their speaker models in advance, the system reduces the processing load during the critical verification moment, as the models are already available and the system knows exactly which models to compare against the utterance

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The patent retrieves speaker models from multiple co-located devices, which may be more than strictly necessary (excessive action), but this ensures that verification accuracy is maximized. The system processes more speaker models than a single-device system would, but only for co-located devices, thereby improving verification accuracy while keeping the information processing load manageable through selective multiplication rather than universal expansion

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS11676608B2Speaker verification using co-location information
Publication Date: 2023.06.13 GOOGLE LLC
  • US11676608B2 patent drawing
  • US11676608B2 patent drawing
  • US11676608B2 patent drawing

AI summary

A method includes generating an audio signal encoding an utterance captured by a microphone of a user device and transmitting the audio signal encoding the utterance to a server. The server is configured to determine a speaker of the utterance from one of a plurality of different users of the user device based on a comparison between the audio signal encoding the utterance and corresponding speaker verification data, and process the audio signal encoding the utterance using a speech recognition module to identify a particular action. The method also includes executing the particular action identified by the server to cause a particular application to launch on the user device based on user permissions associated with the speaker determined by the server to access the particular data.