Audio Signal Processing Adaptive Coding Mode Selection
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Existing audio signal encoding methods often deteriorate sound quality by applying inappropriate coding modes or schemes that do not suit the audio properties, particularly in high frequency bands with high energy signals or harmonic components.
Innovation Solution
An audio signal processing method and apparatus that separately encodes pulses in specific frequency bands with high energy, such as percussion sounds, and harmonic tracks, using adaptive coding modes based on pulse and harmonic ratios, and employs modified discrete cosine transform to accurately extract and encode pulses, reducing bit requirements and improving sound quality.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Device complexity
If a generic coding mode is applied to all audio signals, then device complexity is reduced, but sound quality deteriorates for signals with high energy pulses or harmonics
Solution Approach 1:
The patent implements dynamic mode selection that adapts the coding scheme based on the audio signal characteristics. The system calculates pulse ratio and harmonic ratio for each frame and switches between generic mode, non-generic mode, and harmonic mode accordingly. This dynamic adaptation allows the system to maintain low complexity for generic signals while achieving high sound quality for signals with specific characteristics like high energy pulses or harmonics.
Solution Approach 2:
The patent changes coding parameters based on signal properties by calculating pulse ratio (ratio of pulse energy to total energy) and harmonic ratio (ratio of harmonic energy to total energy). These parameter changes trigger different coding modes: generic mode for low pulse/harmonic ratios, non-generic mode for high pulse ratio, and harmonic mode for high harmonic ratio. This parameter-driven approach resolves the contradiction by matching coding complexity to signal requirements.
2Manufacturing precision
If separate encoding of pulses and harmonics is implemented, then sound quality is improved, but device complexity increases
Solution Approach 1:
The patent segments the audio signal analysis into distinct components by separately calculating pulse ratio and harmonic ratio. This segmentation allows the system to identify and handle high energy pulses and harmonic components independently. The segmentation is implemented through separate energy calculation processes for pulses and harmonics, enabling targeted encoding strategies for each component type.
Solution Approach 2:
The system dynamically selects between different encoding approaches based on the calculated ratios. When pulse ratio exceeds a threshold, non-generic mode is activated for separate pulse encoding. When harmonic ratio exceeds a threshold, harmonic mode is activated for separate harmonic encoding. This dynamic selection mechanism allows separate encoding to be applied only when beneficial, avoiding unnecessary complexity for signals that don't require it.
3Manufacturing precision
If more bits are allocated for encoding pulses and harmonics, then sound quality is improved, but bit usage increases
Solution Approach 1:
The patent uses parameter-based mode selection to allocate bits efficiently. By calculating pulse ratio and harmonic ratio, the system determines the appropriate coding mode for each frame. This parameter-driven approach ensures that enhanced encoding with higher bit allocation is applied only when the signal characteristics warrant it, avoiding unnecessary bit consumption for generic signals that don't require special processing.
Solution Approach 2:
The system dynamically adjusts bit allocation based on signal characteristics. In non-generic mode, more bits are allocated for precise pulse encoding when pulse ratio is high. In harmonic mode, more bits are allocated for accurate harmonic representation when harmonic ratio is high. This dynamic bit allocation strategy improves sound quality for complex signals while maintaining efficient bit usage for simpler signals.
Data Source
AI summary
The present invention relates to a method for processing an audio signal, comprising: a step of performing a frequency conversion process on an audio signal to obtain a plurality of frequency transform coefficients; a step of selecting either a general mode or a non-general mode, on the basis of a pulse ratio, for the frequency transform coefficients having a high frequency band from among the plurality of frequency transform coefficients; and a step of performing, if the non-general mode is selected, the following steps: extracting a predetermined number of pulses from the frequency transform coefficients having the high frequency band, and generating pulse information; generating an original noise signal from the frequency transform coefficients having the high frequency band, excluding the pulses; generating a reference noise signal using the frequency transform coefficient having a low frequency band from among the plurality of frequency transform coefficients; and generating noise position information and noise energy information using the original noise signal and the reference noise signal.


