Voice Signal Collision Detection and Frequency Shift

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

In multi-party voice communication systems, talker collisions often occur, making it difficult for listeners to distinguish and understand multiple voice signals simultaneously, as existing technologies do not effectively enhance speech intelligibility in mixed voice signals.

Innovation Solution

A voice signal processing method that detects talker collisions and applies time or frequency shifts to one of the signals to make it more distinguishable, using a common time base for synchronization and processing only during collision intervals, with techniques like time-stretching, copying, or frequency-shifting to reduce overlap and improve intelligibility.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Adaptability or versatility

If multiple voice signals are mixed simultaneously, then the system supports multi-party communication, but speech intelligibility deteriorates due to talker collisions

Engineering Contradiction:
Improvemulti-party communication capabilityVSAvoidspeech intelligibility
Core Design Contradiction:
Adaptability or versatilityVSLoss of information

Solution Approach 1:

The patent segments the mixed voice signal into multiple individual voice signals using source separation techniques. This allows the system to handle multi-party communication while maintaining speech intelligibility by processing each speaker's signal separately rather than mixing them together, thus resolving the talker collision problem.

Inventive Principle:
Principle #1Segmentation

Solution Approach 2:

The patent changes acoustic parameters such as pitch, timbre, and spatial position of individual voice signals to distinguish between simultaneous speakers. By modifying these parameters, the system maintains intelligibility in multi-party communication scenarios where multiple voices overlap in time.

Inventive Principle:
Principle #35Parameter changes

2Loss of information

If spatial cues are added to separate speakers, then speech intelligibility improves, but system complexity increases

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidsignal processing complexity
Core Design Contradiction:
Loss of informationVSDevice complexity

Solution Approach 1:

The patent dynamically adjusts spatial cues and signal processing parameters based on real-time detection of talker collisions. The system activates source separation and spatial processing only when collisions are detected, rather than continuously processing all signals, thus improving intelligibility while managing system complexity through adaptive operation.

Inventive Principle:
Principle #15Dynamics

Solution Approach 2:

The system uses feedback from collision detection to control the application of spatial cues and source separation. When talker collisions are detected, the system activates processing to separate and spatially position the signals; when no collisions occur, processing is reduced or suspended, optimizing the balance between intelligibility and complexity.

Inventive Principle:
Principle #23Feedback

3Loss of information

If voice signals are processed to distinguish colliding talkers, then speech intelligibility improves, but processing time increases

Engineering Contradiction:
Improvespeech intelligibilityVSAvoidsignal processing time
Core Design Contradiction:
Loss of informationVSLoss of time

Solution Approach 1:

The patent performs preliminary detection of talker collisions using simple energy-based metrics before applying complex source separation algorithms. This preliminary action allows the system to identify when processing is needed and prepare for intelligent processing, reducing overall processing time while maintaining speech intelligibility.

Inventive Principle:
Principle #10Preliminary action

Solution Approach 2:

The system applies partial processing by focusing computational resources only on the specific time-frequency regions where talker collisions occur, rather than processing the entire signal continuously. This selective processing approach maintains intelligibility where needed while minimizing overall processing time and computational load.

Inventive Principle:
Principle #16Partial or excessive action

Data Source

PatentUS9502047B2Talker collisions in an auditory scene
Publication Date: 2016.11.22 DOLBY LABORATORIES LICENSING CORP
  • US9502047B2 patent drawing
  • US9502047B2 patent drawing
  • US9502047B2 patent drawing

AI summary

From a plurality of received voice signals, a signal interval in which there is a talker collision between at least a first and a second voice signal is detected. A processor receives a positive detection result and processes, in response to this, at least one of the voice signals with the aim of making it perceptually distinguishable. A mixer mixes the voice signals to supply an output signal, wherein the processed signal(s) replaces the corresponding received signals. In example embodiments, signal content is shifted away from the talker collision in frequency or in time. The invention may be useful in a conferencing system.