Audio Conference Endpoint State-Based Processing

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Conventional audio conferencing systems are limited in dynamically adjusting audio playback and capture based on the conference state, leading to issues like 'suck' or 'brown out' due to spurious noise and echo suppression, which degrade the audio quality and productivity during meetings.

Innovation Solution

The system analyzes voice activity in both near-end and far-end audio streams to determine the conference state, adjusting microphone sensitivity, voice activity detection, and playback levels to optimize audio processing, allowing for full duplex operation and minimizing noise and echo.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If conventional audio conferencing systems use echo suppression and noise filtering, then audio quality is improved, but spurious noise and 'suck' or 'brown out' effects occur degrading audio quality

Engineering Contradiction:
Improveaudio qualityVSAvoidspurious noise and echo suppression artifacts
Core Design Contradiction:
ReliabilityVSObject-generated harmful factors

Solution Approach 1:

The system dynamically adjusts audio processing parameters including echo suppression strength, noise filtering intensity, and playback levels based on the detected conference state. The conference state machine transitions between states such as 'far-end talking', 'near-end talking', and 'presentation' to optimize processing in real-time, preventing the fixed processing from causing spurious noise artifacts.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system changes processing parameters based on conference state detection. When a presentation state is detected, the system reduces echo suppression intensity and adjusts playback levels to prevent 'suck' or 'brown out' effects. The voice activity detection thresholds and processing gains are adjusted according to the current state to maintain audio quality without introducing artifacts.

Inventive Principle:
Principle #35Parameter changes

2Productivity

If the system operates in full duplex mode with simultaneous microphone capture and speaker playback, then communication efficiency is improved, but acoustic coupling and echo occur

Engineering Contradiction:
Improvecommunication efficiencyVSAvoidacoustic coupling and echo
Core Design Contradiction:
ProductivityVSObject-generated harmful factors

Solution Approach 1:

The system enables full duplex operation by dynamically adjusting echo cancellation parameters and playback levels based on the conference state. When near-end voice activity is detected, the system increases echo suppression intensity. The adaptive noise gate and playback level adjustments prevent acoustic coupling while maintaining simultaneous capture and playback capability.

Inventive Principle:
Principle #15Dynamics

3Object-affected harmful factors

If the system uses aggressive noise suppression and echo cancellation, then noise and echo are reduced, but audio quality and naturalness degrade

Engineering Contradiction:
Improvenoise and echo levelsVSAvoidaudio quality and naturalness
Core Design Contradiction:
Object-affected harmful factorsVSReliability

Solution Approach 1:

The system adjusts noise suppression and echo cancellation parameters based on the detected conference state. In presentation states, aggressive suppression is reduced to maintain naturalness. The system dynamically modifies processing intensity, gain levels, and filtering strength to prevent over-processing artifacts while maintaining noise and echo reduction effectiveness.

Inventive Principle:
Principle #35Parameter changes

Data Source

PatentEP3280123B1State-based endpoint conference interaction
Publication Date: 2022.05.11 DOLBY LABORATORIES LICENSING CORP
  • EP3280123B1 patent drawingFigure 1
  • EP3280123B1 patent drawingFigure 2
  • EP3280123B1 patent drawingFigure 3

AI summary

Systems and methods are described for modifying one of far-end signal playback and capture of local audio on an audio device. Frames of both a far-end audio stream and a near-end audio stream may be analyzed using a measure of voice activity, the analyzing producing voice data associated with each frame. Based on the voice data, a conference state may be determined, and one of playback of the far-end audio stream and capture of local audio on an audio device may be modified based on the determined conference state. By associating the likely intent with a predefined state, the device may further cull or remove unwanted or unlikely content from the device input and output. This may have a substantial advantage in allowing for full duplex operation in the case of more meaningful and continuing voice activity, particularly in the case where there are many connected endpoints.