Acoustic Source Separation Using Stochastic De-Mixing Updates
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing blind source separation techniques are computationally expensive and produce poor results in noisy environments, particularly when dealing with multiple simultaneous speakers, as they rely on restrictive constraints and are not suitable for age-related hearing loss or automatic speech recognition systems.
Innovation Solution
A method for acoustic source separation that involves converting acoustic data into the time-frequency domain and constructing a multichannel filter using de-mixing matrices, where the gradient values are calculated from a stochastic selection of time-frequency data frames to reduce computational burden and improve separation efficiency.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If traditional blind source separation techniques are used, then source separation can be achieved, but the computational cost is very high
Solution Approach 1:
The patent segments the computation by processing frequency bins independently and using stochastic gradient descent to process batches of time-frequency frames rather than all frames at once. This divides the large computational problem into smaller, manageable segments that can be processed iteratively, reducing memory requirements and computational complexity at each step.
Solution Approach 2:
The patent applies partial action by using a subset (batch) of time-frequency frames for each gradient descent iteration rather than processing the complete dataset. This allows the algorithm to make progressive improvements to the demixing matrices through multiple iterations, achieving good separation quality with reduced computational burden per iteration.
2Measurement precision
If traditional blind source separation techniques are used, then source separation can be achieved, but the processing time is very long
Solution Approach 1:
The patent implements periodic action through iterative gradient descent, where the algorithm periodically updates the demixing matrices by processing batches of frames, evaluating the cost function, and adjusting parameters. This iterative periodic process allows the system to converge to a solution over multiple cycles, achieving high separation quality without requiring a single lengthy processing step.
Solution Approach 2:
By processing only a batch of frames at each iteration rather than the complete dataset, the patent reduces the processing time per iteration. Multiple iterations with partial data processing accumulate to achieve the same or better separation quality than a single pass through all data, significantly reducing total processing time.
3Reliability
If restrictive constraints are applied in blind source separation, then mathematical convergence can be improved, but the technique becomes unsuitable for certain problem classes
Solution Approach 1:
The patent changes the parameterization approach by using a general cost function formulation that does not require restrictive constraints on the mixing model or source statistics. By parameterizing the demixing matrices and optimizing them through gradient descent with batch processing, the method adapts to different problem classes (hearing aid, speech recognition, general BSS) without requiring problem-specific constraint modifications.
Solution Approach 2:
The patent introduces dynamics through the iterative gradient descent optimization process, where the demixing matrices are dynamically adjusted based on the gradient of the cost function computed from batches of data. This dynamic adaptation allows the system to converge reliably while maintaining versatility across different problem types, as the optimization process automatically adapts to the specific characteristics of each application.
Data Source
AI summary
A method for acoustic source separation comprises inputting acoustic data from a plurality of acoustic sensors, combined from a plurality of acoustic sources, converting the acoustic data to time-frequency domain data comprising time-frequency data frames, and constructing a multichannel filter for the time-frequency data frames to separate signals from the acoustic sources. The constructing comprises determining a set of de-mixing matrices (Wf) to apply to each time-frequency data frame to determine a vector of separated outputs (yft) by modifying each of the de-mixing matrices by a respective gradient value (G;G′) for a frequency dependent upon a gradient of a cost function measuring a separation of the sources by the respective de-mixing matrix. The respective gradient values for each frequency are each calculated from a stochastic selection of the time-frequency data frames.


