Server-Based Echo Suppression for Web Audio Conferences
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Web browser-based audio communication systems face challenges in effectively suppressing acoustic echoes due to limitations in manipulating audio data, making it difficult to implement optimal echo cancellation or suppression without specialized hardware or software.
Innovation Solution
A server-based echo suppression system processes audio streams from multiple users, identifying the active speaker and designating their stream as such, while silencing or replacing other streams with silence or comfort noise to prevent feedback loops, thereby achieving effective echo suppression without requiring software or hardware installation on user devices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If echo cancellation is performed using specialized hardware or software, then echo suppression effectiveness is improved, but device complexity and installation requirements increase
Solution Approach 1:
The patent introduces a media server as an intermediary component between users and the echo cancellation system. The server performs echo cancellation on audio streams before distributing them to participants, eliminating the need for specialized hardware or software on user devices. This mediator approach maintains high echo suppression effectiveness while reducing client-side complexity to zero.
Solution Approach 2:
The system enables self-service echo cancellation by having the media server automatically process audio streams without requiring user configuration or specialized client software. The server independently analyzes and cancels echoes in real-time, making the system easy to deploy and use while maintaining professional-grade echo suppression.
2Reliability
If echo cancellation is performed in real-time software, then echo suppression is achieved, but computational cost and processing time increase
Solution Approach 1:
By offloading computational intensive echo cancellation tasks to a centralized media server with dedicated processing resources, the patent eliminates the need for high computational power on client devices. The server handles all real-time audio processing, allowing clients to use minimal computational resources while maintaining effective echo suppression.
3Ease of operation
If browser-based applications are used, then ease of access and distribution is improved, but ability to manipulate audio data for echo suppression deteriorates
Solution Approach 1:
The media server acts as an intermediary that compensates for the limited audio manipulation capabilities of browser-based applications. Since browsers cannot perform complex audio processing, the server performs echo cancellation on incoming streams before delivering them to browser clients, maintaining both web-based convenience and effective echo suppression.
Solution Approach 2:
Instead of attempting to perform echo cancellation in the browser (client-side), the patent inverts the approach by implementing it on the server-side. This reversal allows browser-based applications to maintain their simplicity and wide availability while still achieving professional-grade echo suppression through server processing.
4Reliability
If half-duplex approach is used, then echo feedback loop is prevented, but communication flexibility and user experience deteriorate
Solution Approach 1:
The media server intermediary enables full-duplex communication while preventing echo feedback loops through active echo cancellation processing. The server can simultaneously receive and process multiple audio streams, identify and cancel echoes in real-time, and distribute cleaned streams to all participants, maintaining both communication flexibility and echo prevention.
Data Source
AI summary
A system and method for performing echo suppression on a server in browser-based online audio conferences without downloading or installing software on a participant's computing device is disclosed. Streams of audio communication data from the participants in an audio conference are received at the server. An echo suppression application determines the first party that speaks by analyzing the streams to locate speech data, and assigns that party as the “owner” of the audio channel. The speech data is sent to the other participants in the conference. The application then determines whether newly received audio from the owner of the channel is new speech; if so, then the party remains the owner of the channel, and the new speech data is also sent to the other parties in the conference. The channel is surrendered if no new speech is received from the owner in a defined period, and the next party that speaks becomes the new owner of the channel. The other audio data from the participants is replaced by silence.


