Low-Complexity Audio Downsampling With Cascaded Filterbanks
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio processing methods face high computational complexity and audio quality trade-offs, particularly in bandwidth extension applications, due to inefficient use of filterbanks and sampling rates, leading to unwanted artifacts and increased processing demands.
Innovation Solution
A method involving a cascaded placement of analysis and synthesis filterbanks with specific channel configurations to achieve low complexity processing while maintaining audio quality, using quadrature mirror filterbanks and spectral alignment to reduce computational complexity and improve perceptual quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Manufacturing precision
If complex modulated filter banks with high frequency resolution and high degree of oversampling are employed, then audio quality is improved, but computational complexity increases
Solution Approach 1:
The patent segments the audio signal processing into multiple stages: initial analysis filter bank, transposition stage, and synthesis filter bank. Each stage processes specific frequency ranges or transposition orders, allowing the system to achieve high audio quality through targeted processing rather than uniformly high complexity across all processing paths.
Solution Approach 2:
The patent applies partial processing by selectively processing only certain transposition orders (T=2, T=3, T=4) rather than all possible orders, and by using different processing complexities for different frequency bands. This allows the system to achieve sufficient audio quality without the excessive computational complexity of processing all components at maximum detail.
2Adaptability or versatility
If multiple analysis filter banks processing signals of different transposition orders are used, then bandwidth extension is achieved, but device complexity increases
Solution Approach 1:
The patent designs a universal processing framework where a single analysis filter bank output is reused for multiple transposition orders (T=2, T=3, T=4). The same filtered subband signals are processed through different transposition operations to generate multiple high-frequency bands, eliminating the need for separate analysis filter banks for each transposition order.
Solution Approach 2:
The patent merges the processing paths by combining multiple transposition operations (different transposition orders) into a unified synthesis stage. Instead of maintaining separate complete filter bank chains for each transposition order, the system merges the intermediate processing and combines results at the synthesis stage, reducing overall complexity.
3Manufacturing precision
If bandpass filters are applied to input signals to obtain non-overlapping power spectral densities, then spectral separation is achieved, but computational complexity increases
Solution Approach 1:
The patent performs preliminary spectral separation through the analysis filter bank stage, which divides the input signal into distinct subbands before transposition. This preliminary separation ensures that subsequent transposition operations work with already-separated spectral components, eliminating the need for additional bandpass filtering stages and reducing overall computational complexity.
Data Source
Figure 1~2
Figure 3~4
Figure 5A~5B
AI summary
A method for downsampling an audio signal, comprises: generating, using a first filter bank (102, 601, 2302), a plurality of subband signals from the audio signal, wherein a sampling rate of the subband signals is smaller than a sampling rate of the audio signal; performing a sample rate conversion using at least one synthesis filter bank (602-2, 2304) followed by an analysis filter bank (603-2, 2307) to obtain a sample rate converted signal, the at least one synthesis filter bank (602-2, 2304) having a number of channels different from a number of channels of the analysis filter bank (603-2, 2307); processing, using a time stretch processor (604-2, 2309), the sample rate converted signal to obtain a time stretched signal; and combining the time stretched signal and a low-band signal or a different time stretched signal.