Audio Encoding Subset Selection for Conference Feedback Reduction

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In audio and video conferencing, the complexity of networking architectures and the increase in participants lead to difficulties in ensuring clear audio transmission due to microphone picking up both local voices and speaker signals, along with the computational intensity of encoding audio signals from multiple participants, making it challenging to maintain efficient communication.

Innovation Solution

The use of encoder pools, where loudest participants are identified and their audio signals are separately encoded and broadcast, while other participants are assigned to encoder pools based on their codecs, reducing the number of times audio signals are encoded and minimizing feedback.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If audio signals from all participants are encoded separately for each participant, then audio quality is maintained, but computing resources and encoding time increase significantly

Engineering Contradiction:
Improveaudio qualityVSAvoidencoding efficiency
Core Design Contradiction:
ReliabilityVSProductivity

Solution Approach 1:

The patent segments participants into different groups: active speakers who receive individual encoded audio streams, and inactive participants who receive mixed audio streams. This segmentation allows the system to apply different encoding strategies to different participant groups, reducing overall encoding complexity while maintaining audio quality for active participants.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

Instead of fully encoding audio for all participants, the patent applies partial encoding only to active participants who need individual audio streams. The mixed audio stream serves as a sufficient approximation for inactive participants, reducing the excessive computational action required for full individual encoding of all participants.

Inventive Principle:
Principle #16Partial or excessive action

2Quantity of substance

If the microphone picks up both local voices and speaker signals, then all audio inputs are captured, but feedback and audio quality deteriorate

Engineering Contradiction:
Improveaudio input coverageVSAvoidfeedback
Core Design Contradiction:
Quantity of substanceVSObject-affected harmful factors

Solution Approach 1:

The patent extracts and removes the speaker signal component from the microphone input before encoding. By separating the harmful feedback component (speaker signal) from the useful signal (local participant voices), the system eliminates feedback while preserving the audio input coverage from local participants.

Inventive Principle:
Principle #2Taking out (Extraction)

Solution Approach 2:

The patent implements feedback cancellation by detecting the speaker output signal and subtracting it from the microphone input signal. This feedback mechanism actively removes the harmful feedback component while maintaining the desired audio input from local participants.

Inventive Principle:
Principle #23Feedback

3Quantity of substance

If multiple participants speak simultaneously, then all voices are captured, but it becomes difficult to hear individual participants

Engineering Contradiction:
Improveaudio input coverageVSAvoidparticipant audio distinguishability
Core Design Contradiction:
Quantity of substanceVSMeasurement precision

Solution Approach 1:

The patent dynamically adjusts the audio streaming strategy based on real-time detection of active participants. When multiple participants speak simultaneously, the system identifies the most active participants and provides them with individual encoded streams, while other participants receive mixed audio. This dynamic adaptation maintains audio distinguishability during simultaneous speech events.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The patent applies different audio quality levels to different participants based on their activity status. Active participants who need to be heard clearly receive individually encoded high-quality audio streams, while inactive participants receive lower-quality mixed audio streams. This local quality differentiation maintains distinguishability for active speakers while reducing overall processing load.

Inventive Principle:
Principle #3Local quality

4Productivity

If encoder pools are used to reduce encoding cycles, then computing resources are saved, but participants with different codecs require separate pool management

Engineering Contradiction:
Improveencoding cycle efficiencyVSAvoidencoder pool management
Core Design Contradiction:
ProductivityVSDevice complexity

Solution Approach 1:

The patent creates encoder pools that are universal in their ability to handle multiple codec types. Each encoder pool is designed to work with different codecs, allowing the same pool infrastructure to serve participants with diverse codec requirements. This universality reduces the need for separate pool management for each codec type while maintaining encoding efficiency.

Inventive Principle:
Principle #6Universality (Multi-functionality)

Data Source

PatentUS11800017B1Encoding a subset of audio input for broadcasting conferenced communications
Publication Date: 2023.10.24 8X8 INC
  • US11800017B1 patent drawing
  • US11800017B1 patent drawing
  • US11800017B1 patent drawing

AI summary

Various example implementations are directed to methods and apparatuses for facilitating conferenced communications. In one of various examples involving audio signals received from a plurality of participants of a digital audio conference, a logic circuit is to process the audio signals via respective audio input circuits respectively associated with each of the endpoint devices, and, in response to a subset of the different audio signals deemed or qualified as having a loudest audio input, encodes audio from only the subset for broadcasting to participants of the digital audio conference.