Voice Interface Device Speaker Identification in Noise
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice interface devices struggle to handle multiple users effectively, particularly in noisy environments, as they lack efficient methods for identifying speakers, coordinating responses among devices, and mitigating noise interference.
Innovation Solution
The implementation of a system that uses voice models to identify speakers, negotiates leadership among multiple devices for response, and detects noise levels to suggest alternative wake-up methods, ensuring personalized and accurate responses while maintaining effective communication in noisy conditions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If voice interface devices use hotword detection to wake up, then they can be activated by voice input, but they fail to accurately identify speakers and provide personalized responses in noisy environments
Solution Approach 1:
The system performs preliminary voice model training for each user before actual operation. During this training phase, the device collects and stores voice characteristics of each user, creating personalized voice models that enable accurate speaker identification even in noisy environments. This preliminary preparation allows the system to distinguish between different speakers by comparing their voice patterns against stored models rather than relying solely on hotword detection.
Solution Approach 2:
The system creates voice model copies or representations of each user's voice characteristics. These voice models serve as digital templates that capture unique speaking patterns, pitch, tone, and other vocal features. When a voice input is received, the system compares it against these stored voice model copies to identify the speaker, enabling personalized responses without being affected by background noise.
2Adaptability or versatility
If multiple voice interface devices are deployed in a location, then user coverage is improved, but user confusion increases due to multiple devices responding simultaneously
Solution Approach 1:
The system introduces a coordination mechanism that acts as an intermediary between multiple voice interface devices. This coordination system manages which device should respond to a given voice input, preventing multiple devices from responding simultaneously. The intermediary layer ensures that only one device outputs a response at a time, eliminating user confusion while maintaining the benefits of having multiple devices available for different locations and users.
Data Source
AI summary
A method at an electronic device with one or more microphones and a speaker includes receiving a first voice input; comparing the first voice input to one or more voice models; based on the comparing, determining whether the first voice input corresponds to any of a plurality of occupants, and according to the determination, authenticating an occupant and presenting a response, or restricting functionality of the electronic device.


