Conference Audio Mixing via Speech Activity Classification
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing conference call systems lack efficient methods to mix and prioritize voice signals from multiple participants, leading to poor audio quality and difficulty in distinguishing actively speaking individuals, especially in scenarios with multiple simultaneous speakers.
Innovation Solution
A method and system that classify speech activity into different categories, apply gains from a mixing table based on this classification, and rank signals to prioritize those with higher likelihoods of being actively spoken, thereby improving audio mixing and output quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Reliability
If all participant signals are mixed equally in a conference call, then all participants are included in the output, but actively speaking participants cannot be distinguished and audio quality deteriorates
Solution Approach 1:
The patent applies local quality by assigning different gain values to different participant signals based on their speech activity. Instead of uniform mixing, each signal receives localized quality adjustment through gain control, where actively speaking participants receive higher gains and non-speaking participants receive lower or zero gains, thereby distinguishing active speakers while maintaining overall conference inclusion
Solution Approach 2:
The patent changes the mixing parameter dynamically by adjusting gain values based on speech activity detection. The system monitors speech activity parameters in real-time and modifies the mixing gains accordingly, transitioning between different gain configurations to reflect the current speaking state of participants, thus improving audio quality adaptively
2Reliability
If speech activity classification is implemented to prioritize actively speaking participants, then audio quality and speaker distinction improve, but system complexity increases
Solution Approach 1:
The patent applies preliminary action by pre-defining multiple mixing tables with different gain configurations corresponding to different speech activity scenarios. Instead of calculating optimal gains in real-time, the system classifies the current speech activity state and selects the pre-prepared mixing table that best matches the current scenario, thereby reducing real-time computational complexity while maintaining speaker distinction capability
Solution Approach 2:
The patent implements dynamics by creating a dynamic mixing system that adapts to changing speech activity conditions. The system continuously monitors speech activity and dynamically switches between different pre-defined mixing tables, allowing the mixing configuration to evolve with the conference state without requiring complex real-time optimization algorithms
Data Source
AI summary
A network entity, method and computer program product are provided for effectuating a conference session. The method may include receiving a plurality of signals representative of voice communication of the participants. In this regard, the signals may be received from a plurality of terminals of a respective plurality of participants at one of the locations, each of at least some of the terminals otherwise being configured for voice communication independent of at least some of the other terminals. The method of this aspect also includes classifying speech activity of the conference session according to a speech pause, or one or more actively-speaking participants, during the conference session. The signals of the respective participants may then be mixed into a at least one mixed signal for output to one or more other participants at one or more other locations, the signals being mixed based upon classification of the speech activity.


