Machine Learning Speech Level Estimation for Automatic Gain Control

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Modern automatic gain control systems in conferencing technologies often fail to distinguish between speech and noise effectively, leading to undesirable gain changes that can dampen speech in noisy environments or amplify noise in quiet ones, resulting in poor audio quality and inefficient resource usage.

Innovation Solution

A machine learning-based system that estimates the desired signal level by removing background noise from audio input signals using a trained model, allowing for dynamic gain adjustments based on the energy levels of speech signals across multiple frequency ranges to maintain a consistent output volume.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If traditional automatic gain control is used to adjust gain levels in audio conferencing, then output volume normalization is achieved, but speech and noise cannot be effectively distinguished leading to dampening of speech or amplification of noise

Engineering Contradiction:
Improvespeech and noise distinction accuracyVSAvoidaudio quality maintenance
Core Design Contradiction:
Measurement precisionVSEase of operation

Solution Approach 1:

A machine learning model is introduced as an intermediary component between the audio input and the gain control mechanism. This model processes the raw audio signal to estimate speech and noise levels separately, providing accurate measurements to the gain control system. The intermediary model enables precise speech-noise distinction without requiring complex manual tuning of the gain control parameters.

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent replaces traditional mechanical or algorithmic gain control methods with a machine learning-based approach. Instead of using conventional signal processing techniques to distinguish speech from noise, the system employs a trained neural network model that automatically learns and applies speech-noise separation, achieving superior measurement precision.

Inventive Principle:
Principle #28Mechanics substitution (Replace mechanical system)

2Device complexity

If gain control adjusts all frequency ranges uniformly, then simple processing is maintained, but speech quality deteriorates in noisy environments due to inability to selectively enhance speech frequencies

Engineering Contradiction:
Improveprocessing complexityVSAvoidspeech intelligibility
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The audio signal is segmented into multiple frequency ranges, and the machine learning model independently estimates speech and noise levels for each frequency band. This segmentation allows the system to apply different gain adjustments to different frequency ranges, enhancing speech intelligibility by selectively boosting speech-dominated frequencies while suppressing noise-dominated frequencies.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The gain control mechanism applies local quality adjustments by treating each frequency range differently based on its speech-noise characteristics. Instead of uniform gain adjustment across all frequencies, the system optimizes gain for each frequency band individually, ensuring that speech quality is maintained or enhanced in speech-dominated bands while noise is suppressed in noise-dominated bands.

Inventive Principle:
Principle #3Local quality

3Measurement precision

If machine learning model processes each frequency range separately to improve speech detection accuracy, then measurement precision increases, but computational complexity and processing time increase

Engineering Contradiction:
Improvespeech level estimation accuracyVSAvoidmodel processing complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

A single machine learning model is designed to perform multiple functions: it simultaneously processes multiple frequency ranges, estimates both speech and noise levels, and outputs predictions for all frequency bands in one pass. This multi-functional model approach achieves high measurement precision across all frequency ranges without requiring separate processing chains for each band, thereby managing computational complexity efficiently.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS12073845B2Automatic gain control based on machine learning level estimation of the desired signal
Publication Date: 2024.08.27 GOOGLE LLC
  • US12073845B2 patent drawing
  • US12073845B2 patent drawing
  • US12073845B2 patent drawing

AI summary

Method includes receiving, at a server device, from a plurality of input devices, audio data. The audio data of each input device corresponds to a time-related portion of the audio data. The method determines a speech energy level for each input device by providing the time-related audio portion as input to a trained model. For each input device, a statistical value associated with the speech energy level is determined. A strongest input device is identified based on the statistical value. The statistical value associated with the speech energy level of each input device other than the strongest input device is compared to the statistical value of the strongest input device. Depending on the comparison, the method determines whether to update the gain value of an input device to an estimated target gain value based on the statistical value of the speech energy level of the respective input device.