Multi-Microphone Speech Augmentation via Noise Component Modeling

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In multi-microphone system environments, mismatches between speech signals processed by each microphone system lead to significant performance degradations in automated clinical documentation systems.

Innovation Solution

A computer-implemented method that obtains speech signals from multiple devices, selects a noise component model based on these signals, and augments the speech signals in real-time to reduce noise and reverberation, thereby improving signal consistency across different microphone systems.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If data augmentation is used to improve robustness to noise and reverberation, then speech processing accuracy is improved, but mismatch between speech signals from different microphones causes performance degradation

Engineering Contradiction:
Improvespeech processing accuracyVSAvoidcompatibility across microphone systems
Core Design Contradiction:
ReliabilityVSAdaptability or versatility

Solution Approach 1:

The patent transforms speech signals by applying acoustic transfer functions and noise component models to modify acoustic parameters (reverberation characteristics, noise profiles) while preserving speech content. This allows training data to adapt to different microphone acoustics without requiring separate training for each device

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent introduces acoustic transfer functions and noise component models as intermediary representations that capture the relationship between different microphone systems. These intermediaries enable consistent speech processing across devices by serving as a bridge between microphone-specific characteristics

Inventive Principle:
Principle #24Intermediary (Mediator)

2Measurement precision

If speech signals are processed separately by each microphone system, then device-specific optimization is achieved, but significant performance degradation occurs due to signal mismatch

Engineering Contradiction:
Improvedevice-specific processing accuracyVSAvoidcross-device performance consistency
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system applies acoustic transfer functions to transform speech signals between different acoustic domains, enabling a model trained on one microphone to process signals from another microphone with high accuracy by changing the acoustic parameters to match the target device characteristics

Inventive Principle:
Principle #35Parameter changes

Solution Approach 2:

The patent creates a universal speech processing framework where a single trained model can process speech from multiple different microphone systems. The acoustic transfer functions and noise component models enable one model to serve multiple devices by adapting to their specific acoustic characteristics

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12243514B2Data augmentation system and method for multi-microphone systems
Publication Date: 2025.03.04 MICROSOFT TECHNOLOGY LICENSING LLC
  • US12243514B2 patent drawing
  • US12243514B2 patent drawing
  • US12243514B2 patent drawing

AI summary

A method, computer program product, and computing system for obtaining one or more speech signals from a first device, thus defining one or more first device speech signals. One or more speech signals may be obtained from a second device, thus defining one or more second device speech signals. A noise component model may be selected from a plurality of noise component models based upon, at least in part, the one or more first device speech signals and the one or more second device speech signals. The one or more second device speech signals may be augmented, at run-time, based upon, at least in part, the noise component model.