Audio Conference Endpoint State-Based Processing
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio conferencing systems are limited in dynamically adjusting audio playback and capture based on the conference state, leading to issues like 'suck' or 'brown out' due to spurious noise and echo suppression, which degrade the audio quality and productivity during meetings.
Innovation Solution
The system analyzes voice activity in both near-end and far-end audio streams to determine the conference state, adjusting microphone sensitivity, voice activity detection, and playback levels to optimize audio processing, allowing for full duplex operation and minimizing noise and echo.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If conventional audio conferencing systems use echo suppression and noise filtering, then audio quality is improved, but spurious noise and 'suck' or 'brown out' effects occur degrading audio quality
Solution Approach 1:
The system dynamically adjusts audio processing parameters including echo suppression strength, noise filtering intensity, and playback levels based on the detected conference state. The conference state machine transitions between states such as 'far-end talking', 'near-end talking', and 'presentation' to optimize processing in real-time, preventing the fixed processing from causing spurious noise artifacts.
Solution Approach 2:
The system changes processing parameters based on conference state detection. When a presentation state is detected, the system reduces echo suppression intensity and adjusts playback levels to prevent 'suck' or 'brown out' effects. The voice activity detection thresholds and processing gains are adjusted according to the current state to maintain audio quality without introducing artifacts.
2Productivity
If the system operates in full duplex mode with simultaneous microphone capture and speaker playback, then communication efficiency is improved, but acoustic coupling and echo occur
Solution Approach 1:
The system enables full duplex operation by dynamically adjusting echo cancellation parameters and playback levels based on the conference state. When near-end voice activity is detected, the system increases echo suppression intensity. The adaptive noise gate and playback level adjustments prevent acoustic coupling while maintaining simultaneous capture and playback capability.
3Object-affected harmful factors
If the system uses aggressive noise suppression and echo cancellation, then noise and echo are reduced, but audio quality and naturalness degrade
Solution Approach 1:
The system adjusts noise suppression and echo cancellation parameters based on the detected conference state. In presentation states, aggressive suppression is reduced to maintain naturalness. The system dynamically modifies processing intensity, gain levels, and filtering strength to prevent over-processing artifacts while maintaining noise and echo reduction effectiveness.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Systems and methods are described for modifying one of far-end signal playback and capture of local audio on an audio device. Frames of both a far-end audio stream and a near-end audio stream may be analyzed using a measure of voice activity, the analyzing producing voice data associated with each frame. Based on the voice data, a conference state may be determined, and one of playback of the far-end audio stream and capture of local audio on an audio device may be modified based on the determined conference state. By associating the likely intent with a predefined state, the device may further cull or remove unwanted or unlikely content from the device input and output. This may have a substantial advantage in allowing for full duplex operation in the case of more meaningful and continuing voice activity, particularly in the case where there are many connected endpoints.