Microphone-Array-Agnostic Voice Recognition With Masked Beamforming
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing multi-channel input voice recognition technologies are restricted by specific microphone array forms, requiring consistent microphone array shapes and numbers, and lack stability across different forms, leading to performance degradation in complex environments.
Innovation Solution
A device and method that utilize a time-frequency transformer, speaker and noise mask estimator, beamformer estimator, and learning machine to transform and filter voice signals independently of microphone array form, enabling stable recognition through meta-learning and fine-tuning with a small amount of data.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If existing multi-channel input voice recognition technology is used, then recognition performance can be improved in specific microphone array forms, but the system cannot adapt to different microphone array forms and requires retraining with new data
Solution Approach 1:
The patent applies universality by designing a voice recognition model that can process multiple microphone array forms (linear, circular, tetrahedral, etc.) with different numbers of microphones using a unified architecture. The model uses channel attention mechanisms and spatial encoding that are form-agnostic, allowing the same model to adapt to various microphone configurations without requiring separate models or extensive retraining for each form.
Solution Approach 2:
The patent employs parameter changes by dynamically adjusting attention weights and spatial encoding parameters based on the input microphone array configuration. The channel attention module learns optimal weighting for different channels, and the spatial encoding parameters are adapted to match the specific geometry of the microphone array being used, enabling the model to optimize its behavior for each configuration while maintaining a single universal model.
2Measurement precision
If training is performed with a specific microphone array form, then recognition accuracy is improved for that form, but the system fails to perform sufficiently when a different microphone array form is used
Solution Approach 1:
The patent applies preliminary action by pre-training the model on a diverse dataset encompassing multiple microphone array forms and configurations. This pre-training establishes a robust foundation that enables the model to generalize to unseen microphone array forms. The channel attention mechanisms and spatial encoding are pre-configured to handle various geometries, allowing the model to maintain reliable performance when deployed with different microphone arrays without requiring form-specific retraining.
3Adaptability or versatility
If multiple devices with different microphone array forms are used, then device versatility is improved, but each device requires additional training tasks and new data collection
Solution Approach 1:
The patent eliminates the need for device-specific training by creating a universal voice recognition model that can be deployed across multiple devices with different microphone array forms. The model's architecture, featuring channel attention mechanisms and form-agnostic spatial encoding, allows it to adapt to any microphone configuration through fine-tuning on small datasets or even direct deployment, thereby saving significant time and resources that would otherwise be required for separate training processes.
4Ease of manufacture
If fixed microphone array forms are required, then model training is simplified, but the system lacks flexibility for real-world applications with varying microphone configurations
Solution Approach 1:
The patent maintains training simplicity while achieving flexibility by using parameter-based adaptation rather than architecture changes. The model employs learnable parameters in the channel attention module and spatial encoding that automatically adjust to match different microphone array forms. This approach keeps the overall model architecture simple and easy to train, while the parameters dynamically adapt to accommodate various microphone configurations in real-world applications.
Data Source
AI summary
A device for recognizing a multi-channel input voice includes a time-frequency transformer that receives channel audio signals extracted from voice data recorded through microphones having unspecified microphone array forms and transforms the channel audio signals into time-frequency domain signals, a speaker and noise mask estimator that receives the time-frequency domain signals and estimates a time-frequency domain mask for voices and noise for speakers, a beamformer estimator that estimates a time-frequency domain signal for voice signals of the speakers from which the noise has been removed from the time-frequency domain signals by using the time-frequency domain mask, a time-frequency inverse transformer that inversely transforms the time-frequency domain signal from which the noise has been removed into a time domain signal, and a learning machine that trains the speaker and noise mask estimator based on a loss function obtained by comparing the inversely-transformed time domain signal and a pre-defined answer signal.


