Conference State-Based Audio Processing for Noise Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio conferencing systems are inefficient in managing noise and nuisance from multiple endpoints, leading to degradation of audio quality during conference calls due to their inability to dynamically adjust signal processing based on the conference state, resulting in issues like 'suck' or 'brown out' where important audio is compromised by minor noise and echo suppression.

Innovation Solution

The system analyzes voice activity in both near-end and far-end audio streams to determine the conference state, adjusting parameters such as microphone sensitivity, voice activity detection, and playback levels to optimize audio processing, thereby minimizing noise and improving audio capture and playback during conference calls.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Object-affected harmful factors

If echo suppression is applied to remove noise from conference audio, then noise reduction is improved, but important audio content is lost along with the noise

Engineering Contradiction:
Improvenoise reductionVSAvoidaudio content loss
Core Design Contradiction:
Object-affected harmful factorsVSLoss of information

Solution Approach 1:

The system dynamically adjusts audio processing parameters based on detected conference state. When presentation state is detected, echo suppression is reduced or disabled to preserve audio content. When conversation state is detected, full echo suppression is applied for noise reduction. This dynamic adaptation resolves the contradiction by making the processing intensity dependent on real-time conference conditions.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes processing parameters (echo suppression level, noise reduction intensity) based on conference state detection. The conference state machine transitions between states (presentation, conversation, transition) and adjusts parameters accordingly, allowing the system to optimize between noise reduction and content preservation based on current conference dynamics.

Inventive Principle:
Principle #35Parameter changes

2Reliability

If noise reduction processing is applied to conference audio, then audio quality is improved, but latency increases due to processing time

Engineering Contradiction:
Improveaudio qualityVSAvoidprocessing latency
Core Design Contradiction:
ReliabilityVSLoss of time

Solution Approach 1:

The system dynamically adjusts the intensity of noise reduction processing based on conference state. During presentation states, processing is minimized to reduce latency. During conversation states with higher tolerance for delay, more intensive processing is applied to improve audio quality. This resolves the contradiction by making processing intensity adaptive to conference dynamics.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system applies noise reduction processing periodically rather than continuously, adjusting the processing rate based on conference state changes. During stable presentation states, processing occurs at lower rates to minimize latency. During dynamic conversation states, processing intensity increases to maintain audio quality, accepting higher latency when necessary.

Inventive Principle:
Principle #19Periodic action

3Productivity

If full duplex operation is implemented to allow simultaneous capture and playback, then conferencing efficiency is improved, but acoustic feedback and noise interference increase

Engineering Contradiction:
Improveconferencing efficiencyVSAvoidacoustic feedback and noise
Core Design Contradiction:
ProductivityVSObject-generated harmful factors

Solution Approach 1:

The system dynamically adjusts playback levels and capture sensitivity based on detected conference state and voice activity. During full duplex operation, when presentation is detected, playback is attenuated to prevent acoustic feedback. During conversation modes, playback levels are optimized for clarity while monitoring for feedback conditions. This dynamic control allows full duplex operation while mitigating acoustic feedback and noise interference.

Inventive Principle:
Principle #15Dynamics

4Measurement precision

If voice activity detection is used to identify speaking endpoints, then conference state determination is improved, but false detection of noise as voice increases

Engineering Contradiction:
Improvevoice activity detection accuracyVSAvoidfalse detection rate
Core Design Contradiction:
Measurement precisionVSReliability

Solution Approach 1:

The system uses feedback from multiple detection stages to improve voice activity detection accuracy. The conference state machine receives feedback from voice activity detectors and adjusts detection thresholds based on current state. During presentation states, the system raises thresholds to prevent false detection of background noise. During conversation states, thresholds are lowered to capture all participants. This feedback-based adaptation reduces false detection while maintaining detection precision.

Inventive Principle:
Principle #23Feedback

Solution Approach 2:

The system changes detection parameters (thresholds, sensitivity levels) based on conference state. The state machine transitions between presentation, conversation, and transition states, adjusting voice activity detection parameters for each state. This parameter adaptation allows the system to distinguish between actual voice activity and background noise by contextualizing detections within the overall conference dynamics.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentUS10771631B2State-based endpoint conference interaction
Publication Date: 2020.09.08 DOLBY LABORATORIES LICENSING CORP
  • US10771631B2 patent drawing
  • US10771631B2 patent drawing
  • US10771631B2 patent drawing

AI summary

Systems and methods are described for modifying one of far-end signal playback and capture of local audio on an audio device. Frames of both a far-end audio stream and a near-end audio stream may be analyzed using a measure of voice activity, the analyzing producing voice data associated with each frame. Based on the voice data, a conference state may be determined, and one of playback of the far-end audio stream and capture of local audio on an audio device may be modified based on the determined conference state. By associating the likely intent with a predefined state, the device may further cull or remove unwanted or unlikely content from the device input and output. This may have a substantial advantage in allowing for full duplex operation in the case of more meaningful and continuing voice activity, particularly in the case where there are many connected endpoints.