Joint Audio De-noise and De-reverberation Model for Videoconferencing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Videoconferencing systems face challenges in delivering high-quality audio due to noise and reverberation distortions in captured audio signals, which existing technologies fail to address effectively by requiring separate models for noise and reverberation removal, increasing computational complexity and memory usage.

Innovation Solution

A de-noise and de-reverberation model is trained using guided training with auxiliary teacher models to simultaneously remove noise and reverberation from audio signals, reducing computational complexity and maintaining high audio quality by employing a less complex model structure.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Measurement precision

If separate models are used for noise removal and reverberation removal, then audio quality can be improved, but computational complexity and memory usage increase

Engineering Contradiction:
Improveaudio qualityVSAvoidcomputational complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The patent combines two separate models (noise removal model and reverberation removal model) into a single joint model that performs both functions simultaneously. This merging reduces computational complexity and memory usage while maintaining the audio quality improvements that would otherwise require separate processing stages.

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The joint model is designed to perform multiple functions - both noise removal and reverberation removal - within a single unified architecture. This multi-functionality allows the system to achieve the benefits of separate specialized models without the overhead of maintaining and executing multiple separate processing pipelines.

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Measurement precision

If separate models are used for noise removal and reverberation removal, then audio quality can be improved, but memory usage increases

Engineering Contradiction:
Improveaudio qualityVSAvoidmemory usage
Core Design Contradiction:
Measurement precisionVSQuantity of substance

Solution Approach 1:

By merging the noise removal and reverberation removal functionalities into a single joint model, the patent reduces the total memory footprint. Instead of allocating memory for two separate model structures, weights, and computational buffers, the system uses one consolidated model that shares resources and reduces overall memory consumption.

Inventive Principle:
Principle #5Merging (Combining)

3Measurement precision

If a complex model structure is used for simultaneous noise and reverberation removal, then audio quality improves, but computational complexity increases

Engineering Contradiction:
Improveaudio qualityVSAvoidmodel structure complexity
Core Design Contradiction:
Measurement precisionVSDevice complexity

Solution Approach 1:

The joint model employs a segmented architecture where different components handle specific aspects of noise and reverberation removal. The model processes audio signals through distinct functional segments that can be independently optimized, reducing overall structural complexity while maintaining comprehensive processing capabilities.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent transforms the problem from a two-dimensional approach (separate models for noise and reverberation) to a unified multi-dimensional processing framework. By integrating multiple processing dimensions into a single model architecture, the system achieves complex audio enhancement without proportionally increasing model structure complexity.

Inventive Principle:
Principle #17Another dimension (Dimensionality change)

Data Source

PatentUS12175959B1Joint audio de-noise and de-reverberation for videoconferencing
Publication Date: 2024.12.24 ZOOM VIDEO COMM INC
  • US12175959B1 patent drawing
  • US12175959B1 patent drawing
  • US12175959B1 patent drawing

AI summary

One disclosed example method includes a device receiving an audio signal recorded in a physical environment and applying a de-noise and de-reverberation model onto the audio signal to generate a cleaned audio signal. The de-noise and de-reverberation model is configured to remove noise and reverberation from the audio signal and is trained via a training process. The training process includes training the de-noise and de-reverberation model based on a trained de-noise teacher model and a trained de-reverberation teacher model. The training includes adjusting a portion of parameters of the de-noise and de-reverberation model based on values generated by the de-noise teacher model and the de-reverberation teacher model and then adjusting the parameters of the de-noise and de-reverberation model independently of the de-noise teacher model and the de-reverberation teacher model.