Source Sound Separator Using Linear Combination Spectrum Analysis
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing voice separation technologies, such as SAFIA and methods by Kobayashi et al., face challenges in separating a target voice from multiple interfering sounds, leading to deteriorated sound quality, especially when there are plural noise sources.
Innovation Solution
A source sound separator system using multiple microphones, comprising a first and second spectrum generator for emphasizing the target sound, a third spectrum generator for suppressing the target sound, a phase generator, and a target sound separator to process received sound signals in the time or frequency domain for effective separation of target and interfering sounds from different directions.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If band selection method (SAFIA) is used to separate target voice from noise, then voice separation is achieved, but separation performance severely deteriorates when there are three or more sound sources
Solution Approach 1:
The patent divides the frequency spectrum into multiple bands and processes each band separately using band selection. For each frequency band, it identifies the microphone with maximum amplitude and selects that band's signal, thereby segmenting the complex separation problem into manageable frequency-specific tasks that can be optimized independently
Solution Approach 2:
The patent transitions from time-domain processing to frequency-domain processing by performing spectral analysis. This dimensional change allows the system to distinguish between sound sources based on their frequency characteristics rather than just temporal patterns, enabling better separation performance with multiple noise sources
2Measurement precision
If frequency characteristics are calculated to emphasize sound signals from respective sound sources, then voice separation is improved, but sound quality deteriorates after ultimate source sound separation due to interfering sounds
Solution Approach 1:
The patent extracts and removes interfering sound components from the processed signal. After performing band selection and emphasizing frequency characteristics, it specifically identifies and extracts interfering sounds that remain in the signal, then removes them to produce clean output without the harmful artifacts that would otherwise deteriorate sound quality
Solution Approach 2:
The patent converts the presence of interfering sounds into useful information for separation. By analyzing the frequency characteristics of interfering sounds and using them as reference, the system can better distinguish and emphasize the target voice frequencies, turning the harmful interfering signals into helpful separation cues
Data Source
AI summary
In a source sound separator, first and second target sound predominant spectra are generated respectively by first and second processing operations for linear combination for emphasizing the target sound, using received sound signals of two microphones arrayed at a distance from each other. A target sound suppressed spectrum is generated by processing for linear combination for suppression of the target sound, using the two received sound signals. Further, a phase signal containing a larger amount of signal components of the target sound and exhibiting directivity in the direction of the target sound is generated by processing of linear combination, using the two received sound signals. The target sound and the interfering sound are separated from each other using the first and second target sound predominant spectra, the target sound suppressed spectrum, and the phase signal.


