Headset Audio Processing for Teleconference Double Talk
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
During teleconferencing, participants in the same room experience 'double talk' and reduced audio quality due to imperfect noise cancellation, which causes them to hear each other's voices both directly and through the headset, making communication more difficult.
Innovation Solution
A teleconference system that includes a controller to detect similarity between speech audio signals received by a user's device and teleconference equipment, allowing it to adjust audibility and activate a hear-through mode in the headset, reducing noise cancellation and amplifying only sounds from relevant directions, thereby minimizing direct hearing of nearby participants' voices.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Object-affected harmful factors
If noise cancellation is activated in the headset, then ambient noise is attenuated, but participants in the same room can still hear each other's voices both directly and through the headset (double talk)
Solution Approach 1:
The system applies different audio processing to different sound sources based on their spatial origin. Speech from nearby participants (same room) is attenuated or blocked, while speech from distant participants (teleconference) is maintained. This local differentiation of audio quality based on spatial origin resolves the contradiction by selectively applying noise cancellation only where needed.
Solution Approach 2:
The teleconference system acts as an intermediary that receives audio from both the headset microphone and the teleconference equipment, processes the signals to identify and separate nearby participant speech from distant participant speech, and selectively blocks only the nearby speech while allowing distant speech to pass through. This intermediary processing resolves the double talk issue.
2Ease of operation
If pass-through mode is activated to hear ambient sound, then users can hear sounds around them, but noise cancellation is reduced and audio quality deteriorates
Solution Approach 1:
The system dynamically adjusts the audio processing mode based on the detected scenario. When participants are detected to be in the same room, the system switches to a selective blocking mode that maintains noise cancellation for general ambient noise while specifically blocking nearby participant speech. This dynamic adaptation allows the system to maintain audio quality while enabling ambient sound hearing when needed.
Solution Approach 2:
The system changes the audio processing parameters dynamically. Instead of a fixed pass-through or noise cancellation mode, the system adjusts the degree of noise cancellation and speech blocking based on the detected presence and location of participants. This parameter change allows the system to optimize audio quality while enabling ambient sound hearing according to the specific usage scenario.
3Object-affected harmful factors
If traditional noise cancellation is used, then general ambient noise is reduced, but speech from nearby participants cannot be distinguished from teleconference speech
Solution Approach 1:
The system segments the audio signal into different components based on spatial characteristics. By analyzing the audio signals from both the headset microphone and teleconference equipment, the system identifies and separates speech from nearby participants from speech from distant participants. This segmentation allows the system to maintain general noise cancellation while preserving the ability to distinguish between different speech sources.
Solution Approach 2:
The system uses feedback from audio signals received from both the headset and teleconference equipment to continuously identify and track the sources of speech. By comparing the audio signals and detecting similarities, the system can identify when nearby participants are speaking and adjust the audio processing in real-time to block their speech while maintaining distant participant speech. This feedback mechanism prevents loss of information about speech sources.
Data Source
Figure 1
Figure 2~4
Figure 5~6
AI summary
An apparatus of a teleconference system comprises a controller, the teleconference system being configured to enable communication between a plurality of users. The controller detects similarity between speech audio signals with a user device and at least one teleconference equipment of one or more other users. The controller causes, in response to the similarity, a change in audibility of the speech audio signals between the user device and the at least one teleconference equipment.