Audio Trigger Phrase Detection in Multi-Microphone Voice Systems
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Devices struggle to accurately recognize a specific user's voice command in a multi-user environment and often require disabling audio playback before initiating a voice recognition session.
Innovation Solution
An audio capturing device with multiple microphone pairs and a processor that simultaneously monitors audio input channels to detect an audio trigger phrase and initiate a voice recognition session on the correct channel, using beamforming and noise suppression to isolate the user's voice.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple users are present in a room, then the device can serve multiple users, but it becomes difficult to accurately recognize a specific user's voice command while others are talking
Solution Approach 1:
The audio input is divided into multiple independent channels, each monitored separately for trigger phrases. This segmentation allows the system to identify which specific channel detected the trigger and focus processing on that channel, improving accuracy in multi-user environments by isolating the relevant user's voice from others.
2Measurement precision
If the device disables audio playback before receiving user queries, then voice recognition can be accurately processed, but the user experience is interrupted and less seamless
Solution Approach 1:
The system performs preliminary detection of trigger phrases on audio input channels before initiating full voice recognition processing. By detecting the trigger advance and preparing to process only after the trigger is confirmed, the system can maintain audio playback continuity while still achieving accurate voice recognition when needed.
3Adaptability or versatility
If the device monitors multiple audio input channels simultaneously, then it can detect trigger phrases from any user, but the complexity of audio processing increases
Solution Approach 1:
Each audio input channel is equipped with dedicated trigger phrase detection capability, allowing independent monitoring. This local quality approach enables simultaneous multi-channel detection without requiring complex centralized processing of all channels together, as each channel can independently identify triggers and the system can selectively process the relevant channel.
Data Source
AI summary
A method, a system, and a computer program product for detecting an audio trigger phrase at a particular audio input channel and initiating a voice recognition session. The method includes capturing audio content by a plurality of microphone pairs of an audio capturing device, wherein each microphone pair of the plurality of microphone pairs is associated with an audio input channel of a plurality of audio input channels of the audio capturing device. The method further includes simultaneously monitoring, by a processor of the audio capturing device, audio content on each of the audio input channels. The method further includes: independently detecting, by the processor, an audio trigger phrase on at least one audio input channel of the plurality of audio input channels; and in response to detecting the audio trigger phrase, commencing a voice recognition session using the at least one audio input channel as an audio source.


