Separated Sound Source Synthesis Using Frequency-Azimuth Plane
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current methods for synthesizing separated sound sources from stereo audio signals struggle to accurately identify the actual azimuth of sound sources and improve sound quality, particularly in applications like object-based audio services and multi-channel upmixing.
Innovation Solution
A method and apparatus that generate spatial information using a frequency-azimuth plane by determining signal intensity ratios between left and right channel signals, identifying local maxima in energy distribution, and applying a probability density function, such as a Gaussian window function, to extract separated sound sources.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If the ADRess algorithm is used to establish an azimuth axis based on the ratio of left channel signal to right channel signal, then the synthesis process can be simplified, but the azimuth identification accuracy deteriorates because it does not reflect the actual azimuth
Solution Approach 1:
The patent transforms the azimuth representation from a simplified ratio-based parameter to an energy distribution-based parameter in the frequency-azimuth plane. By changing the fundamental parameter used to represent azimuth information, the system achieves both accurate azimuth identification and maintains computational efficiency through the established energy accumulation process.
Solution Approach 2:
The patent introduces a new dimensional representation by creating the frequency-azimuth plane, which adds the azimuth dimension to the frequency spectrum. This dimensional expansion allows simultaneous representation of frequency content and spatial distribution, resolving the contradiction between simplified processing and accurate azimuth identification.
2Ease of manufacture
If conventional methods are used to synthesize separated sound sources, then the processing can be completed with basic signal operations, but the sound quality deteriorates due to insufficient handling of spatial information
Solution Approach 1:
The patent introduces the frequency-azimuth plane as an intermediary representation that bridges the input stereo signal and the output separated sound sources. This intermediate structure organizes spatial and spectral information in a way that facilitates both processing and high-quality reconstruction, resolving the contradiction between processing simplicity and sound quality.
Solution Approach 2:
The patent applies different processing strategies to different regions of the frequency-azimuth plane. By identifying local maxima in energy distribution and applying targeted signal processing to specific azimuth-frequency regions, the system achieves high sound quality while maintaining overall processing efficiency through selective rather than universal processing.
3Measurement precision
If the energy distribution is calculated by accumulating energy for each azimuth in the frequency-azimuth plane, then the azimuth identification accuracy is improved, but the computational complexity increases
Solution Approach 1:
The patent performs preliminary organization of signal energy into the frequency-azimuth plane structure before azimuth identification. By pre-accumulating energy distribution in this structured representation, the system enables efficient azimuth detection through simple maximum search operations, reducing the computational complexity that would otherwise be required for accurate azimuth identification.
Data Source
AI summary
Provided is a method and apparatus for synthesizing a separated sound source, the method including generating spatial information associated with a sound source included in a frame of a stereo audio signal, and synthesizing a separated frequency-domain sound source from the frame of the stereo audio signal based on the spatial information, wherein the spatial information includes a frequency-azimuth plane representing an energy distribution corresponding to a frequency and an azimuth of the frame of the stereo audio signal.


