Video-Guided Audio Suppression for Conference Noise
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current conferencing systems lack an intelligent method to detect and suppress unwanted audio during electronic conferences, leading to poor conference quality due to background noise, which conventional methods like muting or audio signature-based suppression fail to address effectively.
Innovation Solution
The system processes video data to identify users' environments and actions, combined with contextual information, to determine unwanted audio portions and take remedial actions such as suppression, using machine learning algorithms and voice recognition to differentiate between wanted and unwanted audio.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If manual muting or audio signature-based suppression is used, then unwanted audio can be controlled to some extent, but the system cannot intelligently distinguish between wanted and unwanted audio, leading to poor conference quality in complex environments
Solution Approach 1:
The audio stream is segmented into multiple components including speech, music, noise, and applause using source separation technology. This allows the system to process and suppress only unwanted segments (noise, unwanted music) while preserving wanted segments (speech, desired music), thereby improving conference quality without requiring complete muting.
Solution Approach 2:
The system introduces an intermediary audio processing pipeline that includes speech enhancement, source separation, and selective suppression modules. This intermediary layer sits between the raw microphone input and the conference output, intelligently filtering unwanted audio while preserving wanted audio, thus improving reliability without directly modifying the core conferencing system.
2Object-affected harmful factors
If audio signature-based suppression is used, then certain types of audio can be suppressed, but the system suppresses all audio matching the signature including wanted audio in contexts like music lessons or product releases
Solution Approach 1:
The system applies different processing qualities to different audio sources and time periods. Music detected during product releases or music lessons is preserved with high quality, while music in other contexts is suppressed. Speech is always preserved with high quality, while noise and unwanted sounds are suppressed. This local differentiation resolves the contradiction between suppressing harmful audio and preserving wanted audio in specific contexts.
Solution Approach 2:
The suppression parameters are dynamically adjusted based on real-time context analysis. The system continuously monitors audio signatures, speech activity, and meeting context to dynamically change what is suppressed and what is preserved. This dynamic approach allows the system to adapt to different meeting types and contexts, preserving wanted audio while suppressing unwanted audio.
3Object-affected harmful factors
If complete muting is applied to control unwanted audio, then background noise is eliminated, but legitimate user audio is also blocked requiring constant manual intervention
Solution Approach 1:
The system performs self-service audio management by automatically detecting and suppressing unwanted audio (noise, unwanted music) while preserving wanted audio (speech, desired music). This eliminates the need for manual muting operations by users, as the system autonomously manages audio quality based on real-time analysis of audio signatures, speech activity, and contextual information.
Solution Approach 2:
The system implements continuous feedback loops where audio is analyzed in real-time, suppression decisions are made, and the results are monitored. Speech activity detection and audio signature analysis provide feedback that automatically adjusts suppression levels, ensuring that unwanted audio is suppressed while wanted audio remains clear, without requiring manual user intervention.
Data Source
AI summary
A method includes receiving a video data associated with a user in an electronic conference. The method further includes receiving an audio data associated with the user in the electronic conference. It is appreciated that the video data is processed to determine one or more actions taken by the user, and wherein the processing identifies a physical surrounding of the user. The method further includes identifying a portion of the audio data to be suppressed based on the one or more actions taken by the user during the electronic conference and further based on the identification of the physical surrounding of the user.


