Speaker Recognition via Environmental Sensor Fusion

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Key phrase-based speaker recognition in computing systems is prone to false rejections and identifications due to the brevity of voice data, necessitating the augmentation of identification processes with environmental contextual information.

Innovation Solution

The method involves monitoring a computing environment with sensors to detect key phrases and utilize additional acoustic, image, location, and behavioral data collected before and after the utterance to determine the probability that the key phrase was spoken by an identified user, thereby increasing the accuracy of speaker recognition.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Speed

If key phrase-based speaker recognition is used, then user identification speed is improved, but reliability deteriorates due to false rejections and identifications

Engineering Contradiction:
Improveuser identification speedVSAvoidspeaker recognition reliability
Core Design Contradiction:
SpeedVSReliability

Solution Approach 1:

The patent combines key phrase detection with multiple environmental sensor data sources (acoustic, image, location, behavioral) to create a multi-modal recognition system. This merging of data sources allows the system to maintain fast key phrase-based identification while improving reliability through contextual verification from additional sensors, thereby resolving the contradiction between speed and reliability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The system collects environmental sensor data before and after the key phrase utterance to establish contextual information in advance. This preliminary action of gathering contextual data allows the system to quickly process the key phrase while having verification information ready, thus maintaining identification speed while enhancing reliability through pre-collected contextual evidence.

Inventive Principle:
Principle #10Preliminary action

2Loss of time

If only key phrase data is used for identification, then processing time is reduced, but measurement precision deteriorates

Engineering Contradiction:
Improveprocessing timeVSAvoidspeaker identification accuracy
Core Design Contradiction:
Loss of timeVSMeasurement precision

Solution Approach 1:

The system uses the key phrase as a partial trigger for identification, processing this brief data segment quickly to initiate the recognition process. Environmental sensor data is then selectively applied to verify the identification when needed, rather than continuously processing all data streams. This partial action approach reduces processing time for the critical key phrase while maintaining accuracy through selective verification.

Inventive Principle:
Principle #16Partial or excessive action

3Loss of information

If environmental sensor data is collected at different times, then contextual information is improved, but device complexity increases

Engineering Contradiction:
Improvecontextual information qualityVSAvoidsensor data integration complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The system employs a multi-functional data processing architecture where a single processing pipeline handles multiple sensor types (acoustic, image, location, behavioral) using unified methods. This universal approach allows the system to collect and process environmental sensor data at different times without requiring separate specialized processing paths for each sensor type, thereby improving contextual information while managing device complexity through consolidation.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11270695B2Augmentation of key phrase user recognition
Publication Date: 2022.03.08 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11270695B2 patent drawing
  • US11270695B2 patent drawing
  • US11270695B2 patent drawing

AI summary

Examples for augmenting user recognition via speech are provided. One example method comprises, on a computing device, monitoring a use environment via one or more sensors including an acoustic sensor, detecting utterance of a key phrase via data from the acoustic sensor, and based upon the selected data from the acoustic sensor and also on other environmental sensor data collected at different times than the selected data from the acoustic sensor, determining a probability that the key phrase was spoken by an identified user. The method further includes, if the probability meets or exceeds a threshold probability, then performing an action on the computing device.