Audio Signal Direction Detection for Voice Recognition
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition systems face challenges in accurately determining a target sound source among multiple external sound sources, particularly in noise environments, leading to deteriorated performance when controlling devices like TVs and Bluetooth audio devices.
Innovation Solution
An electronic device equipped with multiple microphones and a processor that determines the direction of sound sources using angle and phase information, separates audio signals, and transmits only the signal from the target sound source to a voice recognition server, enabling effective voice recognition by identifying the target sound source based on the duration of sound source directions and wake-up words.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If microphone-array technique is used to strengthen sound source and eliminate noise, then voice recognition performance is improved, but device complexity increases
Solution Approach 1:
The patent segments the audio signal processing into distinct functional modules: a sound source direction determination unit that identifies directions of multiple sound sources, and a voice signal extraction unit that separates target voice from noise based on directional information. This modular segmentation improves voice recognition in noise environments while maintaining manageable system complexity through functional decomposition.
2Adaptability or versatility
If additional electronic device with microphone-array is used for remote control, then control capability is extended, but voice recognition performance deteriorates due to audio output from controlled device
Solution Approach 1:
The patent segments the audio field into multiple directional components by determining sound source directions and separating audio signals according to these directions. This allows the system to isolate and eliminate audio output from controlled devices (TV, Bluetooth audio) while preserving the user's voice commands, thereby maintaining remote control capability without voice recognition deterioration.
Solution Approach 2:
The patent introduces an intermediary processing stage between audio signal reception and voice recognition. The sound source direction determination unit and voice signal extraction unit act as intermediaries that analyze and separate audio signals before transmission to the voice recognition engine, enabling the system to handle remote control scenarios with audio interference.
3Ease of operation
If voice recognition is performed in noise environment with low SNR, then system usability is maintained, but recognition accuracy deteriorates
Solution Approach 1:
The patent segments the mixed audio signal into multiple directional audio signals based on determined sound source directions. By separating the target voice signal from background noise through directional segmentation, the system maintains usability in noise environments while significantly improving recognition accuracy through enhanced signal isolation.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach allows for stable and accurate voice recognition in various noise environments, reducing computational complexity and improving recognition accuracy by isolating the target sound source and reducing noise interference.
Implementation Method 1
determining the direction in which each of the multiple sound sources is located with reference to the electronic device, on the basis of the multiple audio signals received through the multiple microphones
Data Source
Figure 1
Figure 2
Figure 3
AI summary
An electronic device is disclosed. The electronic device comprises: multiple microphones for receiving audio signals generated by multiple sound sources; a communication unit for communicating with a voice recognition server; and a processor for determining the direction in which each of the multiple sound sources is located with reference to the electronic device, on the basis of the multiple audio signals received through the multiple microphones, determining at least one target sound source among the multiple sound sources on the basis of the duration of the determined direction of each of the sound sources, and controlling the communication unit such that the communication unit transmits, to the voice recognition server, an audio signal of a target sound source from which a predetermined voice is generated among the at least one target sound source.