Voice Separation Using DUET Overlap Elimination
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Traditional voice separation methods, particularly the DUET algorithm, fail to effectively eliminate time-frequency point overlaps, leading to impure separated voices due to inclusion of another person's voice, which degrades the quality of speech recognition.
Innovation Solution
A method and system utilizing the DUET algorithm to eliminate overlapping time-frequency points by determining them based on a distance rule (|d1−d2|<d0/4) and eliminating these points to recover pure first and second sounds in the time domain.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If the DUET algorithm is used for voice separation, then the separation process is simplified and computationally efficient, but overlapping time-frequency points are incorrectly assigned to one of the voices, reducing separation purity
Solution Approach 1:
The patent segments the time-frequency plane into distinct regions using contour lines derived from the magnitude spectra of mixed signals. By dividing the time-frequency plane into multiple regions separated by contour lines, each region can be independently assigned to specific source signals, thereby resolving overlapping points and improving separation purity while maintaining algorithmic simplicity.
Solution Approach 2:
The patent introduces contour lines as an intermediary element between the mixed signals and the separation decision. These contour lines, generated from the magnitude spectra, serve as boundaries that mediate the assignment of time-frequency points to different sources, eliminating direct conflicts and improving separation accuracy without adding significant computational complexity.
2Productivity
If overlapping time-frequency points are assigned to one of the voices in the DUET algorithm, then the separation process is completed, but the separated voice includes another person's voice, degrading speech recognition quality
Solution Approach 1:
By segmenting the time-frequency plane using contour lines, the patent creates distinct regions that prevent misassignment of overlapping points. This segmentation ensures that each time-frequency point is correctly associated with its source signal, improving speech recognition accuracy without significantly impacting processing speed.
Solution Approach 2:
The patent performs preliminary generation of contour lines from the magnitude spectra before the actual separation decision is made. This preliminary action establishes clear boundaries in advance, preventing misassignment of overlapping points during the separation process and thereby improving the reliability of speech recognition.
3Adaptability or versatility
If traditional DUET algorithm is applied, then the voice separation can be implemented with available algorithms, but the separated voices are not pure enough due to inclusion of other voices in overlapping regions
Solution Approach 1:
The patent applies segmentation by introducing contour lines that divide the time-frequency plane into distinct regions. This segmentation approach works with the existing DUET algorithm framework, maintaining its adaptability and availability while significantly improving voice purity by preventing cross-contamination in overlapping regions.
Solution Approach 2:
The contour lines serve as an intermediary layer that enhances the traditional DUET algorithm. This intermediary structure maintains compatibility with the existing algorithm while improving separation purity, allowing the system to leverage available algorithms without compromising voice quality.
Data Source
AI summary
Aspects disclosed herein generally relate to a method and a system for improving voice separation by eliminating overlaps or overlapping points. The time-frequency points from the two recorded mixtures are separated by using a Degenerate unmixing estimation technique (DUET) algorithm. The method or system further eliminates the overlapping time-frequency points which belongs to neither of the original resources of sounds.


