ML-Based Audio Gain Control for Speech-Noise Separation

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern automatic gain control systems in conferencing technologies often fail to differentiate between speech and noise effectively, leading to undesirable gain changes that can dampen speech in noisy environments or augment noise in quiet ones, resulting in poor audio quality and inefficient resource usage.

Innovation Solution

A machine learning-based system that estimates the desired signal level by distinguishing between speech and noise using a trained model, removing background noise to enhance the speech signal and adjust the gain accordingly, ensuring a stable and comfortable audio output.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional automatic gain control adjusts gain based on overall signal level, then output volume is normalized, but speech and noise cannot be differentiated leading to dampened speech or augmented noise

Engineering Contradiction:
Improvesignal level estimation accuracyVSAvoidaudio quality
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

The patent segments the audio signal into speech components and noise components using a machine learning model. This segmentation allows the system to estimate the desired signal level (speech) separately from the background noise, enabling precise gain adjustment that enhances speech while suppressing noise, thereby resolving the contradiction between measurement precision and audio quality

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent introduces a machine learning model as an intermediary between the raw audio signal and the gain control mechanism. This intermediary estimates the desired signal level by learning the relationship between input audio features and actual speech levels, enabling accurate speech-level estimation without directly measuring the speech signal, thus improving both measurement precision and audio quality

Inventive Principle:
Principle #24Intermediary (Mediator)

2Reliability

If gain is increased to amplify weak speech, then speech becomes audible, but background noise is also amplified

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidbackground noise
Core Design Contradiction:
ReliabilityVSObject-affected harmful factors

Solution Approach 1:

The patent extracts the desired speech signal level estimation from the mixed audio signal using a machine learning model. By taking out the speech component information from the overall signal, the system can adjust gain based solely on speech levels without amplifying background noise, thereby improving speech intelligibility while eliminating the harmful effect of noise amplification

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies different gain adjustments to different components of the audio signal based on their nature. Speech components receive gain increases to improve intelligibility, while noise components are suppressed. This local quality approach ensures that amplification benefits speech without proportionally amplifying background noise, resolving the contradiction between reliability and harmful factors

Inventive Principle:
Principle #3Local quality

3Object-affected harmful factors

If gain is decreased to suppress noise, then noise is reduced, but speech signal is also weakened

Engineering Contradiction:
Improvebackground noiseVSAvoidspeech signal strength
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The patent extracts speech level information from the audio signal using machine learning, enabling independent control of speech and noise. By taking out the speech component estimation, the system can suppress noise through gain reduction without weakening the speech signal, as the speech level is maintained through the extracted estimation, resolving the contradiction between noise suppression and speech signal strength

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent applies selective gain reduction that affects noise components more than speech components. By understanding the local characteristics of different signal components through machine learning estimation, the system can reduce gain in noise-dominated frequency/temporal regions while maintaining gain in speech-dominated regions, thereby reducing background noise without weakening the speech signal

Inventive Principle:
Principle #3Local quality

4Measurement precision

If machine learning model estimates desired signal level, then gain control accuracy improves, but computational complexity increases

Engineering Contradiction:
Improvedesired signal level estimationVSAvoidprocessing requirements
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent performs preliminary training of the machine learning model offline using labeled audio data. This preliminary action transfers complex learning computations to the training phase, allowing the deployed system to use a pre-trained model that requires minimal real-time computational resources while still achieving high measurement precision in desired signal level estimation, thus resolving the contradiction between precision and complexity

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20240420719A1Automatic gain control based on machine learning level estimation of the desired signal
Publication Date: 2024.12.19 GOOGLE LLC
  • US20240420719A1 patent drawing
  • US20240420719A1 patent drawing
  • US20240420719A1 patent drawing

AI summary

A system includes a memory and a processing device communicably coupled to the memory. The processing device identifies audio data associated with a plurality of input device. The processing devices determines a speech energy level for each input device by providing the audio data as input to a trained model. For each input device, a statistical value associated with the speech energy level is determined. A strongest input device is identified based on the statistical value. In response to determining that the statistical value associated with the speech energy level of the strongest input device satisfies a threshold condition, the processing device updates the gain value of an input device to an estimated target gain value based on the statistical value of the speech energy level of the respective input device.