Audio Signal Encoding Using Transient Block Grouping
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current audio data compression technologies face challenges in encoding transient state signals, resulting in low encoding quality and poor audio signal reconstruction effects due to the failure to effectively extract and transmit transient state features.
Innovation Solution
An audio signal encoding and decoding method that identifies transient state blocks within audio frames, groups and arranges their spectra based on transient state identifiers, and uses neural networks for encoding and decoding to improve encoding quality and reconstruction accuracy.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Quantity of substance
If traditional audio data compression technology is used to compress audio signals, then the data amount is reduced for transmission and storage, but the encoding quality deteriorates when the audio signal is in a transient state
Solution Approach 1:
The audio signal is divided into multiple blocks, and each block is independently analyzed to identify transient state characteristics. This segmentation allows the encoding system to apply different encoding strategies to different blocks based on their transient state identifiers, thereby maintaining high encoding quality for transient blocks while still achieving overall data compression.
Solution Approach 2:
Different encoding precision is applied to different blocks based on their transient state characteristics. Transient state blocks receive higher encoding precision to preserve their unique characteristics, while non-transient blocks use standard compression. This local quality approach ensures that encoding quality is maintained where it matters most while still achieving overall data reduction.
2Quantity of substance
If traditional audio encoding schemes are applied to transient state signals, then the data amount is reduced, but the audio signal reconstruction effect deteriorates
Solution Approach 1:
The system performs preliminary analysis on each audio block to identify transient state characteristics before the main encoding process. By detecting transient state identifiers in advance, the system can prepare appropriate encoding parameters and ensure that transient blocks are encoded with sufficient detail to maintain reconstruction accuracy, while still achieving overall data compression.
Solution Approach 2:
The encoding parameters are dynamically changed based on the transient state identifiers of different blocks. When a transient block is detected, the system adjusts encoding parameters to preserve more information, thereby improving reconstruction reliability for transient segments while maintaining data compression for the overall signal.
3Device complexity
If spectra of all blocks are encoded uniformly, then the encoding process is simple, but the encoding quality for transient state blocks deteriorates
Solution Approach 1:
The encoding process transitions from a static uniform approach to a dynamic adaptive approach. The system dynamically adjusts encoding strategies based on transient state identifiers detected in each block, allowing the encoding quality to adapt to the specific characteristics of each block while maintaining a relatively simple overall process structure.
Data Source
AI summary
Embodiments of this application disclose an audio signal encoding and decoding method, including: obtaining, based on spectra of M blocks of a current frame of a to-be-encoded audio signal, M transient state identifiers of the M blocks, where the M blocks include a first block, and a transient state identifier of the first block indicates that the first block is a transient state block, or indicates that the first block is a non-transient state block; obtaining group information of the M blocks based on the M transient state identifiers of the M blocks; performing grouping and arranging on the spectra of the M blocks based on the group information of the M blocks, to obtain a to-be-encoded spectrum of the current frame; encoding the to-be-encoded spectrum by using an encoding neural network to obtain a spectrum encoding result; and writing the spectrum encoding result into a bitstream.


