Audio Encoding Subset Selection for Conference Feedback Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In audio and video conferencing, the complexity of networking architectures and the increase in participants lead to difficulties in ensuring clear audio transmission due to microphone picking up both local voices and speaker signals, along with the computational intensity of encoding audio signals from multiple participants, making it challenging to maintain efficient communication.
Innovation Solution
The use of encoder pools, where loudest participants are identified and their audio signals are separately encoded and broadcast, while other participants are assigned to encoder pools based on their codecs, reducing the number of times audio signals are encoded and minimizing feedback.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If audio signals from all participants are encoded separately for each participant, then audio quality is maintained, but computing resources and encoding time increase significantly
Solution Approach 1:
The patent segments participants into different groups: active speakers who receive individual encoded audio streams, and inactive participants who receive mixed audio streams. This segmentation allows the system to apply different encoding strategies to different participant groups, reducing overall encoding complexity while maintaining audio quality for active participants.
Solution Approach 2:
Instead of fully encoding audio for all participants, the patent applies partial encoding only to active participants who need individual audio streams. The mixed audio stream serves as a sufficient approximation for inactive participants, reducing the excessive computational action required for full individual encoding of all participants.
2Quantity of substance
If the microphone picks up both local voices and speaker signals, then all audio inputs are captured, but feedback and audio quality deteriorate
Solution Approach 1:
The patent extracts and removes the speaker signal component from the microphone input before encoding. By separating the harmful feedback component (speaker signal) from the useful signal (local participant voices), the system eliminates feedback while preserving the audio input coverage from local participants.
Solution Approach 2:
The patent implements feedback cancellation by detecting the speaker output signal and subtracting it from the microphone input signal. This feedback mechanism actively removes the harmful feedback component while maintaining the desired audio input from local participants.
3Quantity of substance
If multiple participants speak simultaneously, then all voices are captured, but it becomes difficult to hear individual participants
Solution Approach 1:
The patent dynamically adjusts the audio streaming strategy based on real-time detection of active participants. When multiple participants speak simultaneously, the system identifies the most active participants and provides them with individual encoded streams, while other participants receive mixed audio. This dynamic adaptation maintains audio distinguishability during simultaneous speech events.
Solution Approach 2:
The patent applies different audio quality levels to different participants based on their activity status. Active participants who need to be heard clearly receive individually encoded high-quality audio streams, while inactive participants receive lower-quality mixed audio streams. This local quality differentiation maintains distinguishability for active speakers while reducing overall processing load.
4Productivity
If encoder pools are used to reduce encoding cycles, then computing resources are saved, but participants with different codecs require separate pool management
Solution Approach 1:
The patent creates encoder pools that are universal in their ability to handle multiple codec types. Each encoder pool is designed to work with different codecs, allowing the same pool infrastructure to serve participants with diverse codec requirements. This universality reduces the need for separate pool management for each codec type while maintaining encoding efficiency.
Data Source
AI summary
Various example implementations are directed to methods and apparatuses for facilitating conferenced communications. In one of various examples involving audio signals received from a plurality of participants of a digital audio conference, a logic circuit is to process the audio signals via respective audio input circuits respectively associated with each of the endpoint devices, and, in response to a subset of the different audio signals deemed or qualified as having a loudest audio input, encodes audio from only the subset for broadcasting to participants of the digital audio conference.


