Audio Conference Device Encoder Switchover Synchronization
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Audio conferencing systems face challenges in managing audio signals from multiple participants, leading to increased computing power requirements and decreased speech intelligibility due to background noise and echo effects, especially in large conferences, where dynamically changing active and inactive participants cause encoding and decoding errors during switchover between encoders.
Innovation Solution
A method for classifying audio data flows into homogeneous groups based on classification information, allowing for uniform signal processing, attenuation, or amplification of audio signals, and a switchover method that synchronizes encoding parameters between encoders to maintain audio quality during participant status changes.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If all (N-1) speech signals of N participants are superimposed to form mixed signals for each participant, then complete conference audio coverage is achieved, but computing power requirements increase and background noise becomes perceptible
Solution Approach 1:
The patent segments the N participant signals into multiple groups, where only a subset of signals (M < N-1) are superimposed to form the mixed signal for each participant. This segmentation reduces the computational complexity from O(N²) to O(M×N) while maintaining acceptable audio coverage by selectively including only the most relevant participant signals in each mixed signal.
2Reliability
If all (N-1) speech signals of N participants are superimposed to form mixed signals, then complete conference audio coverage is achieved, but speech intelligibility decreases due to background noise superimposition
Solution Approach 1:
The patent applies local quality by differentiating between useful speech signals and background noise sources. Instead of uniformly superimposing all participant signals, the system selectively includes only M participants whose signals contain useful speech content, excluding those contributing primarily to background noise. This selective approach maintains audio coverage completeness while improving speech intelligibility by reducing noise interference.
3Power
If the set of active and inactive participants is dynamically adapted based on audio signals, then computing outlay is reduced, but audio quality suffers due to abrupt appearance/disappearance of background noises and speech clipping
Solution Approach 1:
The patent applies preliminary action by pre-defining a fixed set of M participants whose signals will be superimposed to form the mixed signal, rather than dynamically adapting the participant set during the conference. This preliminary selection is based on initial audio signal analysis, and the same M participants are consistently included throughout the conference duration. This approach prevents abrupt quality changes and speech clipping that would occur with dynamic adaptation, while still reducing computing outlay compared to including all N-1 participants.
4Adaptability or versatility
If channels are dynamically created and mixed signals are dynamically connected to changing destination participants, then adaptability is improved, but encoding and decoding errors occur during switchover between encoders
Solution Approach 1:
The patent applies universality by using a fixed set of M participants for all mixed signal generation throughout the conference, rather than dynamically reconfiguring which participants are included. This universal approach ensures that the same participant signals are consistently processed by the same encoders, eliminating the encoding and decoding errors that would occur during dynamic switchover between different encoder configurations. The system maintains adaptability through volume control adjustments while preserving encoding stability.
Data Source
AI summary
A method and an audio conference device for carrying out an audio conference are disclosed, whereby classification information associated with a respective audio date flow is recorded for supplied audio data flows. According to a result of an evaluation of the classification information, the audio data flows are associated with at least three groups which are homogeneous with regard to the results. The individual audio data flows are processed uniformly in each group in terms of the signals thereof, and said audio data flows processed in this way are superimposed in order to form audio conference data flows to be transmitted to the communication terminals.


