Microphone-Array-Agnostic Voice Recognition With Masked Beamforming

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Existing multi-channel input voice recognition technologies are restricted by specific microphone array forms, requiring consistent microphone array shapes and numbers, and lack stability across different forms, leading to performance degradation in complex environments.

Innovation Solution

A device and method that utilize a time-frequency transformer, speaker and noise mask estimator, beamformer estimator, and learning machine to transform and filter voice signals independently of microphone array form, enabling stable recognition through meta-learning and fine-tuning with a small amount of data.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If existing multi-channel input voice recognition technology is used, then recognition performance can be improved in specific microphone array forms, but the system cannot adapt to different microphone array forms and requires retraining with new data

Engineering Contradiction:
Improvevoice recognition performanceVSAvoidmicrophone array form adaptability
Core Design Contradiction:
Measurement precisionVSAdaptability or versatility

Solution Approach 1:

The patent applies universality by designing a voice recognition model that can process multiple microphone array forms (linear, circular, tetrahedral, etc.) with different numbers of microphones using a unified architecture. The model uses channel attention mechanisms and spatial encoding that are form-agnostic, allowing the same model to adapt to various microphone configurations without requiring separate models or extensive retraining for each form.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Solution Approach 2:

The patent employs parameter changes by dynamically adjusting attention weights and spatial encoding parameters based on the input microphone array configuration. The channel attention module learns optimal weighting for different channels, and the spatial encoding parameters are adapted to match the specific geometry of the microphone array being used, enabling the model to optimize its behavior for each configuration while maintaining a single universal model.

Inventive Principle:
Principle #35Parameter changes

2Measurement precision

If training is performed with a specific microphone array form, then recognition accuracy is improved for that form, but the system fails to perform sufficiently when a different microphone array form is used

Engineering Contradiction:
Improverecognition accuracyVSAvoidperformance stability across different forms
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The patent applies preliminary action by pre-training the model on a diverse dataset encompassing multiple microphone array forms and configurations. This pre-training establishes a robust foundation that enables the model to generalize to unseen microphone array forms. The channel attention mechanisms and spatial encoding are pre-configured to handle various geometries, allowing the model to maintain reliable performance when deployed with different microphone arrays without requiring form-specific retraining.

Inventive Principle:
Principle #10Preliminary action

3Adaptability or versatility

If multiple devices with different microphone array forms are used, then device versatility is improved, but each device requires additional training tasks and new data collection

Engineering Contradiction:
Improvedevice compatibilityVSAvoidtraining time and data collection time
Core Design Contradiction:
Adaptability or versatilityVSLoss of time

Solution Approach 1:

The patent eliminates the need for device-specific training by creating a universal voice recognition model that can be deployed across multiple devices with different microphone array forms. The model's architecture, featuring channel attention mechanisms and form-agnostic spatial encoding, allows it to adapt to any microphone configuration through fine-tuning on small datasets or even direct deployment, thereby saving significant time and resources that would otherwise be required for separate training processes.

Inventive Principle:
Principle #6Universality (Multi-functionality)

4Ease of manufacture

If fixed microphone array forms are required, then model training is simplified, but the system lacks flexibility for real-world applications with varying microphone configurations

Engineering Contradiction:
Improvemodel training simplicityVSAvoidmicrophone array form flexibility
Core Design Contradiction:
Ease of manufactureVSAdaptability or versatility

Solution Approach 1:

The patent maintains training simplicity while achieving flexibility by using parameter-based adaptation rather than architecture changes. The model employs learnable parameters in the channel attention module and spatial encoding that automatically adjust to match different microphone array forms. This approach keeps the overall model architecture simple and easy to train, while the parameters dynamically adapt to accommodate various microphone configurations in real-world applications.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS20250292765A1Device for recognizing multi-channel input voice independent on microphone array form and learning method thereof
Publication Date: 2025.09.18 ELECTRONICS & TELECOMM RES INST
  • US20250292765A1 patent drawing
  • US20250292765A1 patent drawing
  • US20250292765A1 patent drawing

AI summary

A device for recognizing a multi-channel input voice includes a time-frequency transformer that receives channel audio signals extracted from voice data recorded through microphones having unspecified microphone array forms and transforms the channel audio signals into time-frequency domain signals, a speaker and noise mask estimator that receives the time-frequency domain signals and estimates a time-frequency domain mask for voices and noise for speakers, a beamformer estimator that estimates a time-frequency domain signal for voice signals of the speakers from which the noise has been removed from the time-frequency domain signals by using the time-frequency domain mask, a time-frequency inverse transformer that inversely transforms the time-frequency domain signal from which the noise has been removed into a time domain signal, and a learning machine that trains the speaker and noise mask estimator based on a loss function obtained by comparing the inversely-transformed time domain signal and a pre-defined answer signal.