Conference Audio Encoding Reduces CPU Load via Participant Merging

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current voice-over-IP conference systems do not effectively reduce computational complexity by recognizing and eliminating redundant audio encodings among participants, leading to inefficient CPU usage and increased background noise as the number of participants grows.

Innovation Solution

A system and method that processes audio from voice-over-IP conference participants, recognizes participants using the same audio encoding format, and encodes the audio only once for those with similar attributes, eliminating redundant operations and reducing CPU usage by transmitting encoded audio efficiently.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Reliability

If audio is encoded separately for each participant in a voice-over-IP conference, then each participant receives personalized audio processing, but CPU usage increases and computational complexity increases

Engineering Contradiction:
Improveaudio processing qualityVSAvoidCPU usage
Core Design Contradiction:
ReliabilityVSUse of energy by moving object

Solution Approach 1:

The patent merges identical audio processing operations by recognizing when multiple participants are receiving the same audio content. Instead of encoding audio separately for each participant, the system identifies participants who are not actively speaking and groups them to share the same encoded audio stream, thereby reducing redundant CPU operations while maintaining audio processing quality for active speakers

Inventive Principle:
Principle #5Merging (Combining)

Solution Approach 2:

The patent implements a universal audio encoding approach where a single encoded audio stream can be distributed to multiple participants simultaneously. The system determines participant states and creates a unified processing path for participants with identical audio requirements, allowing the same encoding operation to serve multiple participants and reduce overall computational load

Inventive Principle:
Principle #6Universality (Multi-functionality)

2Manufacturing precision

If audio is encoded for each participant individually, then audio processing is thorough, but redundant operations increase computational complexity

Engineering Contradiction:
Improveaudio encoding accuracyVSAvoidcomputational complexity
Core Design Contradiction:
Manufacturing precisionVSDevice complexity

Solution Approach 1:

The patent combines redundant audio encoding operations by identifying participants who are in the same state (not actively speaking) and delivering them the same encoded audio. This merging approach maintains encoding accuracy for participants who need it while eliminating duplicate encoding operations for those who don't, thereby reducing computational complexity without sacrificing audio processing precision

Inventive Principle:
Principle #5Merging (Combining)

3Loss of information

If all participant audio is processed and combined, then complete audio coverage is achieved, but background noise increases

Engineering Contradiction:
Improveaudio coverage completenessVSAvoidbackground noise
Core Design Contradiction:
Loss of informationVSObject-generated harmful factors

Solution Approach 1:

The patent extracts and separates actively speaking participants from the general participant group. By identifying who is currently speaking and who is not, the system extracts only the necessary audio sources for mixing while excluding silent participants whose audio would contribute only background noise. This ensures complete audio coverage of active speakers while minimizing harmful background noise from inactive participants

Inventive Principle:
Principle #2Taking out (Extraction)

Data Source

PatentEP3031048B1Encoding of participants in a conference setting
Publication Date: 2020.02.19 INTERACTIVE INTELLIGENCE INC
  • EP3031048B1 patent drawingFigure 1
  • EP3031048B1 patent drawingFigure 2

AI summary

A system and method are presented for the encoding of participants in a conference setting. In an embodiment, audio from conference participants in a voice-over-IP setting may be received and processed by the system. In an embodiment, audio may be received in a compressed form and de-compressed for processing. For each participant, return audio is generated, compressed (if applicable) and transmitted to the participant. The system may recognize when participants are using the same audio encoding format and are thus receiving audio that may be similar or identical. The audio may only be encoded once instead of for each participant. Thus, redundant encodings are recognized and eliminated resulting in less CPU usage.