Audio Spectrum Shaping for Low-Energy Feature Retention
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Conventional audio encoding and decoding methods suffer from low encoding quality due to significant variations in statistical average energy distribution across frequency bands, leading to potential loss of spectral features, especially in regions with low energy, which affects the retention of spectral lines during processing.
Innovation Solution
The method involves shaping the whitened spectrum to increase spectral amplitudes in target frequency bands, reducing the dynamic range of energy distribution, thereby retaining more spectral features and improving encoding quality by determining a target frequency band based on factors like sampling rate, channel quantity, and encoding mode, and adjusting spectral values using gain factors to ensure consistent energy distribution.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Loss of information
If time-frequency transformation and whitening are performed on audio signal, then frequency-domain representation is obtained, but statistical average energy distribution range becomes large causing spectral feature loss
Solution Approach 1:
The patent applies spectral shaping to modify the statistical average energy distribution of the whitened spectrum. By adjusting spectral parameters through shaping filters, the energy distribution is transformed to reduce its dynamic range, thereby preventing spectral feature loss while maintaining the frequency-domain representation obtained through time-frequency transformation and whitening.
2Manufacturing precision
If spectral lines in low energy regions are processed by encoding neural network model, then encoding is completed, but spectral lines are likely to be lost resulting in low encoding quality
Solution Approach 1:
The patent performs spectral shaping as a preliminary action before the encoding neural network model processes the spectrum. By pre-adjusting the energy distribution to reduce dynamic range and enhance low-energy spectral regions, the spectral lines are better preserved during subsequent neural network encoding, thereby improving encoding quality and spectral line retention.
Data Source
AI summary
This application disclose an encoding method and apparatus, a decoding method and apparatus, a device, a storage medium, and a computer program, and belong to the field of encoding and decoding technologies. In embodiments of this application, a first whitened spectrum for media data is whitened to obtain a second whitened spectrum, and then encoding is performed based on the second whitened spectrum. A spectral amplitude of the second whitened spectrum in a target frequency band is greater than or equal to a spectral amplitude of the first whitened spectrum in the target frequency band. It can be learned that, in this solution, the spectral amplitude of the first whitened spectrum in the target frequency band is increased, so that a difference between statistical average energy of spectral lines for different frequencies in the obtained second whitened spectrum is small.


