Microphone Array Speech Recognition Source Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Electronic apparatuses face challenges in identifying user utterances amidst multiple sound sources in an environment, as beamforming microphones collect audio components from both the user and other sound sources, making it difficult to isolate the user's command.
Innovation Solution
The apparatus employs multiple microphone sets positioned differently to focus on specific angles and uses a processor to identify corresponding sound-source components, select the most reliable microphone sets, and perform speech recognition based on the identified user command.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If beamforming-type microphone is used to focus on user utterance direction, then speech recognition accuracy is improved, but audio components from other sound sources in the same direction are also collected making it difficult to identify user utterance
Solution Approach 1:
The audio signal is segmented into multiple sound source components through source separation processing. The processor separates the mixed audio signal into distinct components corresponding to different sound sources, allowing identification and selection of the user utterance component even when multiple sources are present in the beamforming direction.
Solution Approach 2:
The system changes the parameter of audio signal processing by applying source separation algorithms that transform the mixed audio signal into separated sound source components. This parameter transformation enables differentiation between user utterance and other sound sources based on their distinct acoustic characteristics.
2Reliability
If multiple microphone sets are used to collect audio from different positions, then ability to identify user utterance is improved, but device complexity increases
Solution Approach 1:
The system adds spatial dimension by deploying multiple microphone sets at different positions and angles. This dimensional expansion creates multiple observation perspectives of the same sound field, enabling the processor to identify and isolate user utterance components through comparative analysis of audio signals from different spatial locations.
Solution Approach 2:
Each microphone set serves multiple functions: collecting audio from its specific direction, providing spatial reference for source separation, and contributing to overall system reliability. The universal design allows any microphone set to potentially capture the user utterance depending on the acoustic environment and user position.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
This approach effectively isolates user utterances from multiple sound sources, enhancing the accuracy of speech recognition and enabling reliable operation even in environments with various sound sources.
Implementation Method 1
A beamforming-type microphone strengthens a channel of a sound received in a direction where the trigger word is detected, and weakens channels of sounds received in other directions
Data Source
AI summary
An electronic apparatus is provided. The electronic apparatus includes an interface configured to receive a first audio signal from a first microphone set and receive a second audio signal from a second microphone set provided at a position different from that of the first microphone set; a processor configured to: obtain a plurality of first sound-source components based on the first audio signal and a plurality of second sound-source components based on the second audio signal; identify a first sound-source component, from among the plurality of first sound-source components, and a second sound-source component, from among the plurality of second sound-source components, that correspond to each other; identify a user command based on the first sound-source component and the second sound-source component; and control an operation corresponding to the user command.


