Voice Recognition via Directional Metadata Filtering
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice recognition technologies face challenges in adapting to changes in the location of external devices and effectively filtering voice signals from both user and external devices, leading to noise interference and inefficient recognition.
Innovation Solution
An electronic device equipped with multiple microphones and a processor that receives metadata signals from external devices to identify and filter voice signals based on direction information, using techniques like DOA and BSS to distinguish between user and external device voice signals, allowing for real-time adaptation and accurate voice recognition.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If a database regarding direction information of external devices is built in advance, then voice signal filtering can be performed by comparing direction information, but it is difficult to adaptably respond to the change of the location of the external device
Solution Approach 1:
The patent applies dynamics by transitioning from a static pre-built database to a dynamic real-time direction information acquisition system. The electronic device continuously receives metadata signals from external devices to obtain current direction information, allowing the system to adapt to location changes while maintaining filtering accuracy.
2Measurement precision
If the external device inserts identification information into the voice signal and outputs it, then the electronic device can identify and filter the voice signal of the external device, but both the voice signal output by the external device and the voice signal of the user in the same frequency band may be filtered
Solution Approach 1:
The patent applies segmentation by separating the filtering criterion from the frequency band and focusing on direction information instead. By using direction information from metadata signals, the system can identify external device voice signals based on their spatial origin without affecting user voice signals, even when they occupy the same frequency band.
3Productivity
If voice recognition is performed without filtering external device voice signals, then the system responds quickly, but noise interference from external devices increases
Solution Approach 1:
The patent applies the taking out principle by extracting and removing external device voice signals from the mixed audio input based on direction information. This allows the voice recognition system to process only relevant user voice signals, reducing noise interference while maintaining quick response times.
Applied Scientific Principles
This section explains which scientific principles are used to turn an abstract innovation direction into a practical engineering solution.
Function Achieved in This Case
The solution enables accurate voice recognition by filtering out external device noise and adapting to changes in device location, ensuring smooth operation regardless of environmental changes.
Implementation Method 1
receive a plurality of voice signals and a metadata signal in a non-audible frequency band regarding at least one of the plurality of voice signals
Implementation Method 2
obtain direction information and frequency band information regarding each of the plurality of voice signals and the metadata signal
Data Source
AI summary
An electronic device for performing a voice recognition and a controlling method are provided. The method includes receiving a plurality of voice signals and a metadata signal in a non-audible frequency band regarding at least one of the plurality of voice signals, through the plurality of microphones, obtaining direction information and frequency band information regarding each of the plurality of voice signals and the metadata signal, identifying the plurality of voice signals and the metadata signal, respectively, based on the direction information and the frequency band information, identifying a voice signal of which direction information is same as direction information of the metadata signal and a voice signal of which direction information is different from direction information of the metadata signal, respectively, among the plurality of voice signals, and performing a voice recognition based on the voice signal of which direction information is different from the direction information of the metadata signal.


