Unified ML Audio Suppression Model for Teleconference Noise Filtering

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current noise suppression systems in teleconference applications require multiple models to filter out background noises and voices, leading to high computational resource usage and delayed audio suppression, making real-time filtering ineffective.

Innovation Solution

A unified machine learning (ML) model is trained to switch between noise suppression modes, allowing it to suppress either background noise or all noise except a user's voice, using a single model that consumes fewer computing resources and provides faster output.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If multiple separate ML models are used for background noise suppression and speech suppression, then comprehensive noise filtering capability is improved, but computational resource usage increases and processing speed decreases

Engineering Contradiction:
Improvenoise filtering capabilityVSAvoidprocessing speed
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent combines multiple separate ML models (background noise suppression model and speech suppression model) into a single unified ML model. This unified model receives audio input and generates suppressed audio output by integrating the functionality of multiple models, thereby reducing computational overhead and improving processing speed while maintaining comprehensive noise filtering capability.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The unified ML model is designed to perform multiple suppression functions simultaneously - it can suppress background noise, speech-like sounds, and other distractions within a single model framework. This multi-functional approach eliminates the need for separate models and enables real-time processing with reduced computational resources.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Reliability

If multiple separate ML models are used for background noise suppression and speech suppression, then comprehensive noise filtering capability is improved, but device complexity increases

Engineering Contradiction:
Improvenoise filtering capabilityVSAvoidmodel architecture complexity
Core Design Contradiction:
ReliabilityVSDevice complexity

Solution Approach 1:

The patent merges multiple separate ML models into a single unified model architecture. Instead of implementing separate background noise suppression and speech suppression models that would increase system complexity, the unified model integrates all suppression functionalities into one coherent structure, simplifying the overall system architecture while maintaining comprehensive filtering capabilities.

Inventive Principle:
Principle #5Merging (Combining)

3Reliability

If multiple ML models run concurrently for real-time audio suppression, then comprehensive noise filtering is achieved, but processing delay increases making real-time suppression ineffective

Engineering Contradiction:
Improvenoise filtering capabilityVSAvoidprocessing delay
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The unified ML model processes audio data in a single pass rather than requiring multiple sequential model inferences. By integrating background noise suppression and speech suppression functionalities into one model, the system eliminates the time delay associated with running multiple models concurrently, enabling effective real-time audio suppression.

Inventive Principle:
Principle #5Merging (Combining)

Data Source

PatentUS20250111857A1Unified audio suppression model
Publication Date: 2025.04.03 AMAZON TECH INC
  • US20250111857A1 patent drawing
  • US20250111857A1 patent drawing
  • US20250111857A1 patent drawing

AI summary

Examples herein provide an approach to enhance an audio mixture of a teleconference application by switching between noise suppression modes using a single model. Specifically, a machine learning (ML) model may be configured to, in response to receiving an audio mixture representation as input, suppress either a background noise of the audio mixture or suppress all noise of the audio mixture except a user's voice. In some examples, the ML model may be trained on speech and background noise training data during a training phase. In addition, the ML model may be trained on a user's voice during an enrollment phase. In addition, during an inference phase, the ML model may enhance the audio mixture by suppressing a portion of the audio mixture.