Speech Model Adaptation for Noise Suppression in Audio
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing noise suppression methods in audio communications, such as speech enhancement and online source separation, are inadequate in effectively separating speech from background noise, especially in non-stationary environments, and do not adapt well to individual callers.
Innovation Solution
The method involves applying and updating speech models based on calling frequency and duration, using either caller-specific or generic models, and storing these models for future communications, allowing for improved noise suppression by adapting to individual speakers and environments.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If speech enhancement methods are used to suppress background noise, then stationary noise suppression is improved, but non-stationary noise and speech separation effectiveness deteriorate
Solution Approach 1:
The patent implements dynamic adaptation of speech models through continuous updating during audio communications. The system transitions from static speech enhancement to dynamic model adaptation, where speech models are updated in real-time based on incoming audio data, enabling effective handling of non-stationary noise conditions while maintaining reliability in noise suppression.
Solution Approach 2:
The patent changes the fundamental parameter of speech model adaptability by introducing caller-specific model updating mechanisms. Instead of using fixed speech enhancement parameters, the system adapts model parameters dynamically based on individual caller characteristics and communication context, improving both stationary and non-stationary noise suppression performance.
2Adaptability or versatility
If online source separation with advanced spectral models is used, then non-stationary background handling is improved, but dependency on accurate source model representation worsens
Solution Approach 1:
The patent implements feedback mechanisms where speech models are continuously updated based on the actual audio data received during communications. This feedback loop allows the system to learn from real-world data, improving source model representation accuracy over time and reducing dependency on pre-defined spectral models while maintaining reliable separation effectiveness.
Solution Approach 2:
The patent performs preliminary adaptation of speech models before actual audio processing by updating models during idle periods or using historical communication data. This preliminary action ensures that accurate source models are ready before critical separation tasks, improving reliability without compromising the ability to handle non-stationary backgrounds in real-time.
3Measurement precision
If caller-specific speech models are applied and updated, then speech quality and intelligibility are improved, but system complexity and data storage requirements worsen
Solution Approach 1:
The patent implements a universal framework that handles both generic and caller-specific speech models through a single adaptive system. The same model updating mechanism serves multiple purposes: improving speech quality for individual callers, adapting to different noise environments, and maintaining a library of reusable models. This multi-functionality reduces overall system complexity despite the added capability of caller-specific adaptation.
4Adaptability or versatility
If speech models are updated during audio communication, then adaptability to individual callers is improved, but processing time and computational load worsen
Solution Approach 1:
The patent implements periodic model updating rather than continuous updating during audio communication. The system updates speech models at specific intervals or under certain conditions (e.g., during pauses in speech, after complete utterances, or at scheduled times), reducing computational load and processing time while maintaining effective adaptation to individual callers throughout the communication.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
A method and an apparatus for separating speech data from background data in an audio communication are suggested. The method comprises: applying a speech model to the audio communication for separating the speech data from the background data of the audio communication; and updating the speech model as a function of the speech data and the background data during the audio communication.