Speech-Selective Audio Mixing for Conference Noise Reduction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conferencing sessions are disrupted by noise interference and speech collisions, where participants often forget to mute their microphones and simultaneous talking occurs, leading to embarrassing moments and disruptions.
Innovation Solution
A conferencing system designates endpoints as primary and secondary talkers, using speech detectors to characterize audio as speech or noise, and adjusts gain settings through faders to reduce noise and mitigate speech collisions, implemented in a conferencing bridge with modules for speech selective mixing and collision handling.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If manual muting is used to prevent noise interference, then noise disruption is reduced, but operational complexity and user burden increase
Solution Approach 1:
The system automatically detects speech and noise conditions and adjusts audio mixing without requiring user intervention. The speech-selective mixer monitors audio inputs and autonomously determines which signals to amplify or attenuate based on speech detection algorithms, eliminating the need for manual muting operations while maintaining audio quality.
Solution Approach 2:
The patent replaces manual mechanical muting operations with an automated electronic speech detection and audio mixing system. Speech detectors and digital signal processing algorithms substitute for manual button presses, using electronic detection and computational methods to identify speech patterns and automatically adjust audio levels.
2Reliability
If automatic speech detection is implemented, then speech collision mitigation improves, but device complexity increases
Solution Approach 1:
The speech-selective mixer performs multiple functions using a single integrated system: it detects speech, identifies primary and secondary talkers, adjusts audio mixing in real-time, and mitigates speech collisions. This multi-functional approach consolidates what could be separate complex systems into one unified device that handles all aspects of audio management.
Solution Approach 2:
The system dynamically changes audio parameters such as gain levels and mixing ratios based on real-time speech detection. By adjusting these parameters automatically according to detected speech conditions, the system achieves reliable speech collision handling without requiring complex hardware modifications, relying instead on software-based parameter optimization.
3Measurement precision
If speech-selective mixing is used, then audio clarity improves, but processing requirements increase
Solution Approach 1:
The system applies speech-selective mixing selectively rather than uniformly to all audio inputs. It identifies primary talkers and applies full processing only to their signals, while using simplified processing for secondary talkers and background noise. This partial application of complex processing reduces overall energy consumption while maintaining audio clarity for the most important signals.
Data Source
AI summary
A conference apparatus reduces or eliminates noise in audio for endpoints in a conference. Endpoints in the conference are designated as a primary talker and as secondary talkers. Audio for the endpoints is processed with speech detectors to characterize the audio as speech or not and to determine energy levels of the audio. As the audio is written to buffers and then read from the buffers, decisions for the gain settings of faders for read audio of the endpoints being combined in the speech selective mix. In addition, the conference apparatus can mitigate the effects of a possible speech collision that may occur during the conference between endpoints.


