Voice Command Verification Using Sound Source Positioning and Signal Comparison
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
AI speakers often malfunction due to noise from home devices such as TVs or radios, incorrectly identifying voice outputs from these sources as user commands, leading to unintended activation.
Innovation Solution
An electronic device equipped with multiple microphones and a processor that identifies the position of a sound source and determines whether it is within a specific zone relative to the device and an access point, comparing communication signals to differentiate between user voice and ambient noise, thereby determining the authenticity of voice commands.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Ease of operation
If the AI speaker activates speech recognition service based on voice wake-up method, then the convenience of voice-controlled operation is improved, but the reliability deteriorates due to incorrect activation by noise from home devices
Solution Approach 1:
The system segments the voice recognition process into multiple independent verification stages: initial voice wake-up detection, sound source positioning through multiple microphones, communication signal pattern matching, and user authentication. Each stage acts as a separate filter to reduce false activations while maintaining convenient voice control.
Solution Approach 2:
The system introduces communication signal pattern matching as an intermediary verification layer between voice detection and command execution. By comparing communication signal patterns against preset patterns, the system mediates between the user's voice input and the final activation decision, preventing noise-induced false positives.
2Reliability
If the electronic device uses multiple microphones and communication signal comparison to verify user voice, then the reliability of command execution is improved, but the device complexity increases
Solution Approach 1:
The communication module serves multiple functions: it transmits data packets for voice recognition, monitors communication signal patterns for user verification, and provides positioning information. By making the communication module multi-functional, the system achieves reliable verification without adding separate dedicated hardware components.
Solution Approach 2:
The system uses its own communication signals and existing microphone array to perform self-verification. The electronic device monitors its own communication signal patterns and uses its built-in multiple microphones for sound source positioning, eliminating the need for external verification devices and reducing overall system complexity.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
Effectively distinguishes between user voice commands and noise from other devices, preventing incorrect activation of functions and ensuring that commands are only executed when clearly given by the user, enhancing the reliability of voice-controlled systems.
Implementation Method 1
identify a position of a sound source based on a voice received through a plurality of microphones
Implementation Method 2
an access point that transmits and receives a communication signal with the electronic device
Data Source
AI summary
An electronic device includes a communication module, a plurality of microphones, and a processor. The processor is configured to identify a position of a sound source based on a voice received through the plurality of microphones, identify whether the position of the sound source is included in a first zone between the electronic device and an access point that transmits and receives a communication signal with the electronic device, identify whether the voice has been uttered by a user based on a comparison between the communication signal and a preset communication signal, when the position of the sound source is included in the first zone, and determine whether to execute a command included in the voice based on the identification of whether the voice has been uttered by the user.


