Multi-Channel Microphone Array Speech Recognition in Noise
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In environments with continuous noise, existing speech recognition technologies struggle to effectively isolate and recognize voice commands from multiple audio sources, leading to reduced recognition performance and efficiency.
Innovation Solution
An electronic device with a multi-channel microphone array analyzes audio signals from various sources, determining the direction and duration of signal input to identify and process only the target source for speech recognition, thereby isolating voice commands from noise.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If speech recognition is attempted for all audio signals input through multiple microphones, then comprehensive speech recognition coverage is achieved, but recognition accuracy deteriorates in noisy environments with continuous noise
Solution Approach 1:
The patent segments audio signals by separating them into different spatial channels based on direction of arrival. The microphone array divides the audio environment into multiple directional beams, allowing the system to process speech and noise separately according to their spatial locations. This enables selective attention to speech-containing directions while filtering out noise from other directions.
Solution Approach 2:
The patent applies different processing qualities to different spatial regions. Directions identified as containing speech receive enhanced processing and prioritization, while directions identified as noise sources receive suppression or filtering. This local differentiation of processing quality improves overall recognition accuracy by focusing computational resources on relevant audio sources.
2Object-affected harmful factors
If beamforming is applied to reinforce audio signals from desired directions, then noise reduction is improved, but device complexity increases due to multiple microphone processing requirements
Solution Approach 1:
The patent implements dynamic beamforming where the directional reinforcement patterns are continuously adjusted based on real-time analysis of audio signal characteristics. The system dynamically identifies speech and noise sources, then adapts the beamforming weights accordingly. This dynamic adaptation allows effective noise reduction while maintaining speech quality without requiring overly complex fixed processing structures.
3Reliability
If all audio signals are processed for speech recognition, then no speech is missed, but processing time and computational resources increase unnecessarily
Solution Approach 1:
The patent performs preliminary classification of audio signals before full speech recognition processing. By first analyzing directional information and signal characteristics to identify potential speech sources, the system prepares a filtered set of candidate signals for detailed recognition. This preliminary sorting action prevents unnecessary processing of clearly noisy or non-speech signals, reducing overall processing time while maintaining reliable speech detection.
Data Source
AI summary
An electronic device for speech recognition includes a multi-channel microphone array required for remote speech recognition. The electronic device improves efficiency and performance of speech recognition of the electronic device in a space where noise other than speech to be recognized exists. A control method includes receiving a plurality of audio signals output from a plurality of sources through a plurality of microphones and analyzing the audio signals and obtaining information on directions in which the audio signals are input and information on input times of the audio signals. A target source for speech recognition among the plurality of sources is determined on the basis of the obtained information on the directions in which the plurality of audio signals are input, and the obtained information on the input times of the plurality of audio signals, and an audio signal obtained from the determined target source is processed.


