Adaptive Mid-Side Extraction for Moving and Reverberant Audio Sources
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional methods for extracting mid and side audio signals from stereo audio signals fail to effectively separate non-center panned, moving, or reverberated audio sources, limiting their applicability in complex audio environments.
Innovation Solution
A method that extracts a target mid audio signal from a stereo audio signal by obtaining target panning and phase difference parameters for each time segment and frequency band, allowing for weighted sum and difference calculations to isolate the target audio source.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If conventional mid and side signal calculation is used, then center-panned sources can be effectively separated, but non-center panned, moving, or reverberated audio sources cannot be separated
Solution Approach 1:
The patent applies dynamics by making the mid/side signal calculation adaptive rather than static. The system dynamically adjusts the separation parameters based on the detected panning position and characteristics of the target audio source. This allows the separation technique to effectively handle center-panned, non-center panned, and moving sources by continuously adapting to the source's spatial and temporal properties.
Solution Approach 2:
The patent changes the parameters used in mid/side signal calculation from fixed values to variable parameters that depend on the target source characteristics. By modifying the calculation parameters based on detected panning position, phase information, and temporal properties, the system achieves reliable separation across diverse source types while maintaining adaptability to different audio scenarios.
2Device complexity
If basic mid and side signal extraction is used, then the method is simple and efficient, but it fails to separate general audio sources such as moving sources or sources with reverberation
Solution Approach 1:
The patent segments the audio signal processing into distinct stages: detection of target source characteristics, calculation of separation parameters based on those characteristics, and application of mid/side signal extraction with those parameters. This segmentation allows the system to maintain computational efficiency while incorporating sophisticated adaptation mechanisms for handling diverse audio sources.
Solution Approach 2:
The patent introduces an intermediary parameter calculation step that bridges the simple mid/side extraction and the complex separation requirements. By calculating intermediate parameters (such as optimal weighting factors and phase adjustments) based on detected source properties, the system enables enhanced versatility without directly complicating the core extraction operation.
3Measurement precision
If panning-specific extraction techniques are used, then stationary non-center panned sources can be targeted, but moving sources and sources with reverberation cannot be separated
Solution Approach 1:
The patent implements feedback by using the detected characteristics of the target audio source (panning position, phase information, temporal properties) to continuously adjust the mid/side signal separation parameters. This feedback loop enables the system to maintain high measurement precision for targeting specific source characteristics while simultaneously adapting to handle various source types including moving sources and those with reverberation.
Solution Approach 2:
The patent performs preliminary detection and parameter calculation before applying the mid/side signal extraction. By first identifying the target source characteristics and pre-calculating the optimal separation parameters, the system achieves precise targeting of specific source properties while being prepared to handle diverse source types through the pre-established adaptive framework.
Data Source
AI summary
The present disclosure relates to a method and audio processing arrangement for extracting a target mid (and optionally a target side) audio signal from a stereo audio signal. The method comprises obtaining (S1) a plurality of consecutive time segments of the stereo audio signal and obtaining (S2), for each of a plurality of frequency bands of each time segment of the stereo audio signal, at least one of a target panning parameter (Θ) and a target phase difference parameter (Φ). The method further comprises extracting (S3), for each time segment and each frequency band, a partial mid signal representation (211, 212) based on at least one of the target panning parameter (Θ) and the target phase difference parameter (Φ) of each frequency band and forming (S4) the target mid audio signal (M) by combining the partial mid signal representations (211, 212) for each frequency band and time segment.


