Stereo Audio Upmixing with Independent Channel Extraction
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio upmixing technologies face issues with positioning accuracy and high correlation between audio signals of different channels, leading to diminished spatial perception and degraded performance.
Innovation Solution
An audio upmixing method that involves feature extraction of stereophonic audio signals, independent extraction of left and right channel signals through separate audio output channels, and fusion of these signals to generate target channel audio signals with reduced correlation and improved positioning accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Measurement precision
If stereophonic audio signal is decoded to generate audio upmixing signal, then audio upmixing function is achieved, but positioning accuracy deteriorates and spatial perception is diminished
Solution Approach 1:
The patent divides the audio signal processing into separate left and right channel extraction paths, with each channel processed independently through dedicated neural networks. This segmentation allows for more precise control and positioning of individual audio elements, improving spatial accuracy in the upmixing output.
Solution Approach 2:
The patent transforms the stereophonic audio signal from a two-channel input into a multi-dimensional output by extracting multiple left and right channel signals that map to different spatial positions. This dimensional expansion enables more accurate spatial positioning and enhanced spatial perception in the upmixing result.
2Device complexity
If traditional decoding method is used for audio upmixing, then processing simplicity is maintained, but signal correlation between channels increases and performance degrades
Solution Approach 1:
The patent employs separate neural networks for processing left and right channels independently, reducing inter-channel correlation. This segmented processing approach maintains reasonable system complexity while significantly improving upmixing performance by preventing signal redundancy between channels.
Solution Approach 2:
The patent transforms the audio signal representation by extracting features in the frequency domain and applying neural network transformations. This parameter transformation from time-domain to frequency-domain processing enables better control over channel independence and correlation reduction.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
Disclosed are an audio upmixing method and an audio apparatus. The method includes: obtaining (202) a stereophonic audio signal, and performing feature extraction on the stereophonic audio signal to obtain a stereophonic audio feature; extracting (204) a plurality of left channel audio signals and a plurality of right channel audio signals from the stereophonic audio feature based on a plurality of audio output channels, where the plurality of audio output channels are independent of each other, and each of the plurality of audio output channels is configured to output the corresponding left channel audio signal or the corresponding right channel audio signal; fusing (206) the left channel audio signal and the right channel audio signal corresponding to one target channel to obtain first audio signals of a plurality of target channels, where each of the plurality of target channels corresponds to two of the plurality of audio output channels; and outputting (208) an audio upmixing signal of a target format based on the first audio signals of the plurality of target channels, where the target format corresponds to the plurality of target channels. An audio upmixing signal of a target format is output based on the first audio signals of the plurality of target channels, where the target format corresponds to the plurality of target channels. This method can improve an audio upmixing effect. (Fig. 1)