Adaptive Spectrum-Time Converter for Audio Signal Coding
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio coding techniques, particularly those using the modified discrete cosine transform (MDCT), face challenges in efficiently encoding highly harmonic signals and stereo signals with significant phase shifts, leading to suboptimal energy compaction and inefficiencies in joint channel coding.
Innovation Solution
The proposed solution involves a signal-adaptive approach that switches between different transform kernels based on the instantaneous input characteristics. This includes using transform kernels with different symmetries for the MDCT and MDST operations, allowing for adaptive energy compaction and improved phase handling in stereo signals.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If traditional MDCT is used for audio coding, then general audio coding quality is maintained, but coding efficiency deteriorates for highly harmonic signals and stereo signals with phase shifts
Solution Approach 1:
The patent applies dynamics by making the transform kernel selection adaptive rather than fixed. The system dynamically switches between different transform kernels (MDCT, MDST, or both) based on signal characteristics detected in each frame, allowing optimal energy compaction for highly harmonic signals and phase-shifted stereo signals while maintaining general audio coding quality
Solution Approach 2:
The patent changes the parameter of transform kernel type based on signal properties. By detecting characteristics such as harmonic content and inter-channel phase differences, the system modifies the transform parameters (selecting different kernel types) to optimize energy compaction and coding efficiency for specific signal conditions
2Productivity
If MDCT is used for stereo signal coding, then channel coding is performed, but joint channel coding performance deteriorates for signals with 90 degree phase shifts
Solution Approach 1:
The system dynamically adapts the transform kernel selection based on detected inter-channel phase differences. When 90 degree phase shifts are detected, the system switches to MDST kernels which are specifically designed to handle such phase characteristics, thereby improving joint channel coding performance
Solution Approach 2:
The patent changes the transform kernel parameters to match the signal's phase characteristics. By detecting phase shift conditions and adjusting the kernel type accordingly (using MDST for phase-shifted signals), the system optimizes coding performance for stereo signals with significant phase differences
3Adaptability or versatility
If fixed transform kernel is used, then implementation is simple, but adaptability to different signal types deteriorates
Solution Approach 1:
The system introduces dynamic adaptability by switching between transform kernels based on signal characteristics. While this increases complexity compared to a fixed kernel approach, it enables the system to adapt to different signal types (highly harmonic signals, phase-shifted stereo signals, and general audio signals) to optimize coding efficiency
Solution Approach 2:
The patent implements feedback by detecting signal characteristics (harmonic content, phase differences) and using this information to guide kernel selection. This feedback mechanism enables adaptive transform kernel switching that optimizes performance for different signal conditions while managing implementation complexity through structured decision logic
Data Source
AI summary
A schematic block diagram of a decoder for decoding an encoded audio signal is shown. The decoder includes an adaptive spectrum-time converter and an overlap-add-processor. The adaptive spectrum-time converter converts successive blocks of spectral values into successive blocks of time values, e.g. via a frequency-to-time transform. Furthermore, the adaptive spectrum-time converter receives a control information and switches, in response to the control information, between transform kernels of a first group of transform kernels including one or more transform kernels having different symmetries at sides of a kernel, and a second group of transform kernels including one or more transform kernels having the same symmetries at sides of a transform kernel. Moreover, the overlap-add-processor overlaps and adds the successive blocks of time values to obtain decoded audio values, which may be a decoded audio signal.


