Speech Processing Model Selection for Dynamic Acoustic Environments

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Microphone arrays introduce time-varying spectral modifications to speech signals due to changes in speaker position and adaptive beamforming, affecting speech processing systems' robustness and accuracy.

Innovation Solution

A computer-implemented method that receives inputs on speaker location and microphone array orientation, generates time-varying spectrally-augmented signals by modeling acoustic variations, and trains speech processing systems to account for these changes, enabling more robust speech recognition in dynamic environments.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If microphone arrays are used to receive speech signals, then spatial selectivity and noise rejection are improved, but time-varying spectral modifications are introduced due to speaker position changes and adaptive beamforming

Engineering Contradiction:
Improvenoise rejectionVSAvoidspectral accuracy
Core Design Contradiction:
ReliabilityVSManufacturing precision

Solution Approach 1:

The system pre-trains multiple speech processing models with different acoustic characteristics (including various beamforming effects and speaker positions) before actual speech recognition. This preliminary preparation allows the system to quickly select an appropriate pre-trained model during runtime without needing to adapt in real-time, thus maintaining spectral accuracy while benefiting from microphone array noise rejection

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system dynamically selects among multiple pre-trained speech processing models based on runtime acoustic conditions (such as speaker position and beamforming configuration). This dynamic adaptation allows the system to maintain spectral accuracy by matching the current acoustic environment with the most appropriate pre-trained model, while still utilizing the noise rejection capabilities of the microphone array

Inventive Principle:
Principle #15Dynamics

2Measurement precision

If adaptive beamforming is used to steer towards speakers, then speech signal quality is improved, but time variations in the speech spectrum are introduced

Engineering Contradiction:
Improvespeech signal qualityVSAvoidspectral stability
Core Design Contradiction:
Measurement precisionVSStability of the object's composition

Solution Approach 1:

Multiple speech processing models are pre-trained with different beamforming configurations and spectral characteristics. This preliminary action captures various beamforming effects in advance, allowing the system to select an appropriate model during runtime without introducing instability from real-time spectral variations

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system creates multiple copies of speech processing models, each trained with different beamforming configurations. These model copies represent different spectral conditions, allowing the system to select the most appropriate copy for the current acoustic environment, thereby maintaining spectral stability while benefiting from adaptive beamforming

Inventive Principle:
Principle #26Copying

3Adaptability or versatility

If speaker position varies in the beampattern, then spatial flexibility is improved, but speech signals are affected by time-varying filters

Engineering Contradiction:
Improvespatial flexibilityVSAvoidspeech signal consistency
Core Design Contradiction:
Adaptability or versatilityVSManufacturing precision

Solution Approach 1:

The system pre-trains speech processing models with speech signals captured from different spatial positions and beamforming configurations. This preliminary action creates a library of models that account for spatial variations, allowing the system to maintain speech signal consistency by selecting an appropriate pre-trained model for the current speaker position

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system changes the parameter of model selection based on spatial conditions. By adjusting which pre-trained model is used (rather than adapting the model itself in real-time), the system maintains speech signal consistency while accommodating spatial flexibility in speaker positioning

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS11783826B2System and method for data augmentation and speech processing in dynamic acoustic environments
Publication Date: 2023.10.10 MICROSOFT TECHNOLOGY LICENSING LLC
  • US11783826B2 patent drawing
  • US11783826B2 patent drawing
  • US11783826B2 patent drawing

AI summary

A method, computer program product, and computing system for receiving one or more inputs indicative of at least one of: a relative location of a speaker and a microphone array, and a relative orientation of the speaker and the microphone array. One or more reference signals may be received. A speech processing system may be trained using the one or more inputs and the one or more reference signals.