Selective Audio Mixing for Conference Endpoints
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In large conferencing scenarios, existing technologies require significant resources for decoding and mixing of audio and video, leading to inefficient processes.
Innovation Solution
A method for selectively combining audio from a subset of endpoints based on audio level information, where only audio streams exceeding a predetermined threshold or having the highest levels are decoded and combined, reducing the computational intensity and system resource usage.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If all audio streams from all endpoints are decoded and combined, then complete audio coverage is achieved, but system resource usage and processing complexity increase significantly
Solution Approach 1:
The patent extracts only the necessary audio streams for mixing by using audio level information to identify and select only those streams that contain actual speech content, rather than processing all received audio streams. This extraction approach reduces processing complexity while maintaining audio coverage completeness.
Solution Approach 2:
The patent changes the parameter used for audio stream selection from binary presence/absence to continuous audio level measurement. By measuring audio levels and comparing against a threshold, the system can dynamically determine which streams to process, resolving the contradiction between complete coverage and processing complexity.
2Productivity
If audio level information is used to select only certain endpoints, then system resource usage decreases, but audio quality may be compromised
Solution Approach 1:
The patent implements feedback by using audio level information from all endpoints to make selection decisions, then mixing only the selected streams. This feedback mechanism ensures that only streams with actual content are processed, improving processing efficiency while maintaining audio quality through intelligent selection rather than arbitrary filtering.
Solution Approach 2:
The patent uses audio level as a dynamic parameter to control stream selection, allowing the system to adapt to changing conference conditions. This parameter-based approach maintains audio quality by selecting streams based on actual content presence rather than static criteria.
3Loss of information
If all audio streams are processed in large conferences, then no audio information is lost, but significant computational resources are required
Solution Approach 1:
The patent extracts only the essential audio streams by filtering out streams with audio levels below the threshold. This extraction eliminates unnecessary processing of silent or near-silent streams, reducing computational resource usage while preventing loss of meaningful audio information.
Solution Approach 2:
The patent introduces audio level threshold as a parameter to control the trade-off between information completeness and resource usage. By adjusting this parameter, the system can optimize processing efficiency while maintaining audio information completeness for streams that actually contain speech.
Data Source
AI summary
Selective audio combination for a conference. The conference may be initiated between a plurality of participants at respective participant locations. The conference may be performed using a plurality of conferencing endpoints at each of the participant locations. Audio may be received from each of the plurality of conferencing endpoints. Audio level information may also be received from each of the plurality of conferencing endpoints. The audio may be combined from a plural subset of the plurality of conferencing endpoints to produce conference audio. The plural subset is less than all of the plurality of conferencing endpoints. The audio may be combined based on the audio level information. The conference audio may be provided to the plurality of conferencing endpoints.


