Stereo Upmixing via Frequency Bin Redistribution
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing technologies struggle to automatically generate surround sound from stereo signals, especially when only stereo recordings are available, and they often result in degradation of the original stereo signal during upmixing.
Innovation Solution
A method that transforms a stereo signal into an upmixed multi-channel time domain audio signal by using a short-time Fast Fourier Transform (s-t FFT) to generate frequency bins, identifying regions of interest in a two-dimensional positional distribution, and applying filtering functions to extract and redistribute audio components across additional channels.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Adaptability or versatility
If traditional upmixing methods are used to convert stereo to multi-channel, then surround sound capability is improved, but the original stereo signal quality degrades
Solution Approach 1:
The audio signal is segmented into frequency bins through FFT transformation, allowing independent processing of different frequency components. This enables selective redistribution of spectral content to maintain stereo integrity while adding surround channels.
Solution Approach 2:
The patent transitions from 2D stereo representation to multi-dimensional spatial distribution by mapping frequency bins to multiple output channels. This dimensional expansion allows surround sound creation without compromising the original stereo plane.
2Productivity
If automatic upmixing is applied to stereo recordings, then production efficiency is improved, but spatial accuracy deteriorates
Solution Approach 1:
The system automatically analyzes the stereo signal's spectral content and self-determines optimal channel distribution without manual intervention. The algorithm processes frequency bins autonomously to create spatially accurate multi-channel output.
Solution Approach 2:
The patent dynamically adjusts spatial parameters based on the input signal's frequency content. By changing panning and spatial distribution parameters according to spectral analysis, the system maintains accuracy while operating automatically.
3Measurement precision
If frequency-based spatial distribution is used, then sound localization is improved, but computational complexity increases
Solution Approach 1:
The computational task is segmented into discrete frequency bins, allowing efficient parallel processing. Each bin can be independently transformed and distributed, reducing overall computational burden while maintaining precise localization.
Solution Approach 2:
The system performs preliminary FFT transformation to create frequency representations before spatial distribution. This preliminary processing enables efficient subsequent channel mapping and reduces real-time computational complexity.
Data Source
AI summary
The present disclosure describes systems and methods for audio signal processing, and more specifically, techniques for automatically generating surround sound by extending stereo signals including a left and right channel to multi-channel formats in an unsupervised and content-independent or content-agnostic manner. In operation, a computing device may receive a stereo audio input signal containing two channels from a sound source. The computing device may transform the stereo audio input signal into an upmixed multi-channel time domain audio signal to create an immersive surround sound listening experience by wrapping the original stereo field to a higher number of speakers in the frequency domain. Based at least on the continuous mapping and the panning coefficient, the computing device may generate the upmixed multi-channel time domain audio signal.


