Audio Signal Processing Device for Reverberant Sound Source Separation
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal processing techniques, such as those described in Japanese Unexamined Patent Application Publication No. 2006-197552 and Alexander et. Al., face limitations in sound source separation due to reverberation and noise, leading to inaccurate calculation of peak coordinates and uniform distribution of histograms, which hinders the effective emphasis of desired audio signals.
Innovation Solution
An audio signal processing device and method that includes a frequency-domain conversion unit, relative value calculation unit, mask generation unit, and time-domain conversion unit, which calculate and apply a time-frequency mask to emphasize desired audio signals by comparing relative values to a predetermined threshold, thereby suppressing undesired signals and enhancing sound source separation.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If linear combination process or histogram clustering is used for sound source separation, then sound source separation can be performed, but reverberation causes histogram peaks to become dull and noise causes uniform distribution, making accurate peak coordinate calculation difficult
Solution Approach 1:
The patent changes the parameter used for sound source separation from amplitude ratio and phase difference to relative time delay. This parameter transformation allows accurate separation even in reverberant environments because relative time delay remains stable despite reverberation, unlike amplitude ratio which becomes unreliable when histogram peaks become dull
Solution Approach 2:
The patent replaces the histogram clustering method (which is sensitive to reverberation and noise) with a relative time delay calculation method based on cross-correlation. This substitution eliminates the problem of dull histogram peaks and uniform distribution caused by noise, achieving more robust sound source separation
2Reliability
If sound source separation is performed in reverberant environments using conventional methods, then some separation effect can be obtained, but the separation accuracy is insufficient due to reverberation components
Solution Approach 1:
The patent extracts the direct sound component from the reverberant signal by identifying the maximum correlation point in the cross-correlation function. This extraction method separates the direct sound (desired signal) from reverberation components, achieving reliable sound source separation even in highly reverberant environments like automobiles
Solution Approach 2:
The patent introduces cross-correlation as an intermediary tool to measure the similarity between signals from different microphones. The cross-correlation function acts as a mediator that can distinguish direct sound from reverberation, enabling accurate separation without being degraded by reverberation components
3Manufacturing precision
If conventional sound source separation methods are used, then processing can be performed, but the desired audio signal cannot be sufficiently emphasized due to inaccurate separation
Solution Approach 1:
The patent performs preliminary calculation of relative time delay for each frequency component before generating the time-frequency mask. This preliminary action establishes accurate reference values that guide the mask generation process, ensuring that the desired audio signal is precisely emphasized without inaccurate suppression of other components
Data Source
AI summary
An audio signal processing device includes: a frequency-domain conversion unit that generates a plurality of pieces of frequency-domain information from a plurality of audio input signals acquired at different positions; a relative value calculation unit that calculates, for each piece of frequency-domain information, a relative value between a time-frequency component included in one frequency-domain information and a time-frequency component included in another frequency-domain information; a mask generation unit that compares the relative value with an emphasized range set based on a relative value threshold stored in advance to generate a time-frequency mask that decreases a value of the frequency-domain information corresponding to the relative value which is outside the emphasized range; a mask multiplication unit that multiplies the time-frequency mask by the frequency-domain information to generate emphasized frequency-domain information; and a time-domain conversion unit that converts the emphasized frequency-domain information into an audio output signal indicated as being time-domain information.


