Voice Signal Collision Detection and Frequency Shift
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
In multi-party voice communication systems, talker collisions often occur, making it difficult for listeners to distinguish and understand multiple voice signals simultaneously, as existing technologies do not effectively enhance speech intelligibility in mixed voice signals.
Innovation Solution
A voice signal processing method that detects talker collisions and applies time or frequency shifts to one of the signals to make it more distinguishable, using a common time base for synchronization and processing only during collision intervals, with techniques like time-stretching, copying, or frequency-shifting to reduce overlap and improve intelligibility.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If multiple voice signals are mixed simultaneously, then the system supports multi-party communication, but speech intelligibility deteriorates due to talker collisions
Solution Approach 1:
The patent segments the mixed voice signal into multiple individual voice signals using source separation techniques. This allows the system to handle multi-party communication while maintaining speech intelligibility by processing each speaker's signal separately rather than mixing them together, thus resolving the talker collision problem.
Solution Approach 2:
The patent changes acoustic parameters such as pitch, timbre, and spatial position of individual voice signals to distinguish between simultaneous speakers. By modifying these parameters, the system maintains intelligibility in multi-party communication scenarios where multiple voices overlap in time.
2Loss of information
If spatial cues are added to separate speakers, then speech intelligibility improves, but system complexity increases
Solution Approach 1:
The patent dynamically adjusts spatial cues and signal processing parameters based on real-time detection of talker collisions. The system activates source separation and spatial processing only when collisions are detected, rather than continuously processing all signals, thus improving intelligibility while managing system complexity through adaptive operation.
Solution Approach 2:
The system uses feedback from collision detection to control the application of spatial cues and source separation. When talker collisions are detected, the system activates processing to separate and spatially position the signals; when no collisions occur, processing is reduced or suspended, optimizing the balance between intelligibility and complexity.
3Loss of information
If voice signals are processed to distinguish colliding talkers, then speech intelligibility improves, but processing time increases
Solution Approach 1:
The patent performs preliminary detection of talker collisions using simple energy-based metrics before applying complex source separation algorithms. This preliminary action allows the system to identify when processing is needed and prepare for intelligent processing, reducing overall processing time while maintaining speech intelligibility.
Solution Approach 2:
The system applies partial processing by focusing computational resources only on the specific time-frequency regions where talker collisions occur, rather than processing the entire signal continuously. This selective processing approach maintains intelligibility where needed while minimizing overall processing time and computational load.
Data Source
AI summary
From a plurality of received voice signals, a signal interval in which there is a talker collision between at least a first and a second voice signal is detected. A processor receives a positive detection result and processes, in response to this, at least one of the voice signals with the aim of making it perceptually distinguishable. A mixer mixes the voice signals to supply an output signal, wherein the processed signal(s) replaces the corresponding received signals. In example embodiments, signal content is shifted away from the talker collision in frequency or in time. The invention may be useful in a conferencing system.


