Source Separation Sound Quality Control
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing source separation methods for audio signals often degrade sound quality as they increase the attenuation of interfering signals, leading to a trade-off between desired separation and unwanted artifacts, and lack effective control over this trade-off.
Innovation Solution
An apparatus and method that use supervised learning to estimate sound quality and adapt parameters for signal processing, allowing for perceptually-motivated and signal-adaptive control over the trade-off between sound quality and separation, including direct estimation of post-processing parameters to ensure sound quality meets predefined requirements.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If the attenuation of interfering signals is increased to improve source separation, then the separation quality is improved, but the sound quality of the output signal deteriorates due to introduced artifacts
Solution Approach 1:
The patent applies parameter changes by using a quality parameter q that controls the trade-off between separation quality and sound quality. By adjusting this parameter, the system can adaptively change the processing characteristics to achieve optimal separation while maintaining acceptable sound quality and minimizing artifacts.
Solution Approach 2:
The patent implements feedback by using an estimated sound quality metric to continuously monitor the output and adjust the separation process. The quality parameter q is determined based on estimated sound quality, creating a closed-loop system that adapts to maintain sound quality while achieving separation.
2Object-generated harmful factors
If partial enhancement is used instead of total separation to maintain higher sound quality, then the sound quality is improved, but the attenuation of interfering signals is reduced
Solution Approach 1:
The patent applies dynamics by making the separation degree adjustable and adaptive rather than fixed. The quality parameter q allows the system to dynamically switch between partial enhancement and total separation modes, adapting to different requirements and maintaining optimal sound quality while achieving the desired level of interfering signal attenuation.
3Manufacturing precision
If aggressive source separation is applied to achieve complete separation, then the separation quality is improved, but listener fatigue increases due to high listening effort
Solution Approach 1:
The patent applies parameter changes by using the quality parameter q to control the aggressiveness of the separation process. By adjusting this parameter, the system can achieve adequate separation quality while avoiding overly aggressive processing that would introduce artifacts and cause listener fatigue, thus maintaining comfortable listening levels.
Data Source
Figure 1a
Figure 1b
Figure 2
AI summary
An apparatus for generating a separated audio signal from an audio input signal is provided. The audio input signal comprises a target audio signal portion and a residual audio signal portion. The residual audio signal portion indicates a residual between the audio input signal and the target audio signal portion. The apparatus comprises a source separator (110), a determining module (120) and a signal processor (130). The source separator (110) is configured to determine an estimated target signal which depends on the audio input signal, the estimated target signal being an estimate of a signal that only comprises the target audio signal portion. The determining module (120) is configured to determine one or more result values depending on an estimated sound quality of the estimated target signal to obtain one or more parameter values, wherein the one or more parameter values are the one or more result values or depend on the one or more result values. The signal processor (130) is configured to generate the separated audio signal depending on the one or more parameter values and depending on at least one of the estimated target signal and the audio input signal and an estimated residual signal, the estimated residual signal being an estimate of a signal that only comprises the residual audio signal portion.