Audio Source Separation with Independent Classification and Time Gating
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing source separation systems often fail to accurately identify and extract specific audio sources, such as speech, due to limitations in classification and separation processes, leading to suboptimal quality and increased latency.
Innovation Solution
A method and apparatus that combines source separation with classification, using time gating and soft gating techniques to qualify extracted sources, reducing latency and improving accuracy by performing classification independently of separation and mitigating errors through multiple classifiers and residual signal processing.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If source separation is performed to extract target audio sources, then the extracted source quality is improved, but the accuracy of identifying the target source type deteriorates
Solution Approach 1:
The system segments the audio processing task into two independent parts: source separation to extract candidate sources, and independent audio classification to identify source types. This segmentation allows each component to optimize for its specific function without compromising the other.
Solution Approach 2:
The patent introduces an intermediary classification stage that acts as a mediator between source separation and final output. The classifier independently verifies the type of extracted sources, resolving the contradiction by adding a verification layer that doesn't interfere with the separation process.
2Measurement precision
If classification is performed on separated audio signals, then source type identification accuracy is improved, but processing latency increases
Solution Approach 1:
The system performs preliminary classification on the original mixture signal before separation, and simultaneously performs classification on separated signals. This preliminary action provides early source type information that can guide subsequent processing, reducing overall latency while maintaining accuracy.
Solution Approach 2:
The patent implements parallel processing where classification operates continuously on both the original mixture and separated signals simultaneously. This continuous parallel action eliminates sequential waiting time, reducing latency while maintaining high identification accuracy through multiple classification pathways.
3Reliability
If multiple classifiers are used to mitigate errors, then classification reliability is improved, but device complexity increases
Solution Approach 1:
The system implements feedback mechanisms where multiple classifiers provide their outputs, and a combination logic layer processes these outputs to produce the final classification result. This feedback-based ensemble approach improves reliability by cross-validating multiple independent classification decisions.
Solution Approach 2:
The patent employs a universal combination logic layer that can integrate outputs from multiple different classifiers. This multi-functional layer handles various classification scenarios and error conditions uniformly, improving reliability without proportionally increasing complexity through a standardized integration approach.
Data Source
AI summary
Computer-implemented methods and devices for combined audio separation and classification are provided. An estimated separated signal is time gated based on a determination of an audio classifier of, at least in part, the original mix of signals before separation. Combined separation, classification, and time gating of both the estimated signal and a residual signal are also provided.


