Data compression method based on arithmetic coding and related equipment
By introducing arithmetic coding and probabilistic prediction with a fixed variable-length prefix into the SZ compressor, the conflict between Huffman coding and LZ series algorithms in data structure processing is resolved, achieving more efficient floating-point data compression and improving the overall compression ratio and computational performance.
Patent Information
- Application Number
- CN202510965668.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-14
- Publication Date
- 2025-10-17
AI Technical Summary
In existing HPC scenarios, the combination of Huffman coding and LZ series algorithms in the SZ compressor has an inherent conflict, resulting in a decrease in compression performance and an inability to effectively exploit the compressibility of floating-point data.
A data compression method based on arithmetic coding is adopted, which combines fixed-length and variable-length prefixes for probability prediction. The bits in the quantization factor sequence are encoded through a probability model. The repetition pattern recognition capability of the LZ compression algorithm and the frequency difference of Huffman coding are utilized to avoid data structure corruption.
It achieves more efficient floating-point data compression, improves the overall compression rate while keeping computational overhead controllable, and is suitable for large-scale floating-point compression scenarios in high-performance computing.
Smart Images

Figure CN120811399A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] Embodiments of the present application relate to the field of data compression, and in particular, to a data compression method based on arithmetic coding and related equipment. BACKGROUND
[0002] With the wide application of high-performance computing (HPC) in scientific simulation, climate modeling, astrophysics, etc., the storage and transmission of massive floating-point data face great bandwidth and storage pressure. Therefore, efficient, lossless or approximately lossless compression of floating-point data has become an important technical direction to improve the overall input / output (I / O) performance.
[0003] At present, one of the widely used floating-point compressors in the HPC scenario is the SZ compressor. The compressor converts the original floating-point data into a series of quantization factors through a prediction error quantization mechanism, and further performs entropy coding on these quantization factors to improve the compression rate. The existing SZ system mainly uses two ways to exploit the compressibility of data:
[0004] (1) Entropy coding represented by Huffman coding: used to entropy code quantization factors by taking advantage of the difference in symbol frequency;
[0005] (2) Dictionary coding represented by LZ series compression algorithms (such as LZ77, LZ4): used to exploit the repetition of the same segment in the bit stream.
[0006] Although Huffman coding and LZ series algorithms each have good compression characteristics, their combined use has inherent conflicts. If Huffman coding is performed first and then the LZ algorithm is applied, since the code words generated after Huffman coding have the characteristics of indefinite length and are difficult to align, it will destroy the structural properties of the original data, making it difficult for the LZ series algorithm to identify repeated patterns at the byte granularity, and the compression effect is significantly reduced. Conversely, if LZ compression is used first and then Huffman coding is performed, it is difficult to accurately model the probability distribution of quantization factors at the granularity of quantization factors, making it difficult to fully play the role of entropy coding. SUMMARY
[0007] Based on the above problems, embodiments of the present application provide a data compression method based on arithmetic coding and related equipment, aiming to improve the compression rate of quantization factors while avoiding the incompatibility or destructive interference of traditional Huffman coding and LZ series algorithms in data processing granularity and coding format, thereby realizing more efficient floating-point data compression.
[0008] In a first aspect, the embodiments of the present application provide an arithmetic coding-based data compression method applied to an SZ compressor containing a probability model, the method comprising:
[0009] inputting to-be-compressed data into the SZ compressor for prediction processing and quantization processing, to obtain a quantization factor sequence composed of a plurality of quantization factors;
[0010] performing probability prediction on each bit in the quantization factor sequence according to the probability model, to determine the probability distribution of each bit; wherein the probability model further comprises using different length prefixes to obtain encoded bit data to perform probability prediction on the currently processed bit, and the different length prefixes comprise a fixed bit length prefix set based on an LZ compression algorithm and a variable length prefix set based on the number of quantization factor bits;
[0011] performing arithmetic coding on the quantization factor sequence according to the probability distribution of each bit, to generate compressed data.
[0012] In an embodiment, before the probability prediction on each bit in the quantization factor sequence using different length prefixes to determine the probability distribution of each bit, the method comprises:
[0013] determining the maximum value of the plurality of quantization factors;
[0014] determining the minimum number of bits k required for each quantization factor according to the most significant bit of the maximum value of the quantization factor;
[0015] re-encoding each quantization factor according to the minimum number of bits k, to generate a quantization factor sequence composed of a plurality of quantization factors with a length of k, to trigger the step of performing probability prediction on each bit in the quantization factor sequence using different length prefixes to determine the probability distribution of each bit.
[0016] In an embodiment, the probability prediction on each bit in the quantization factor sequence according to the probability model to determine the probability distribution of each bit comprises:
[0017] if the bit sequence in the quantization factor to which the currently processed bit belongs is greater than the length value of the fixed bit length prefix, then performing probability prediction on the currently processed bit using different length prefixes.
[0018] In an embodiment, the method further comprises:
[0019] obtaining the context probability distribution of the currently processed bit based on the fixed bit length prefix and the variable length prefix, respectively;
[0020] The obtained context probability distribution is fused to determine a final probability distribution of the bit currently processed.
[0021] In an embodiment, the fixed bit length prefix is an 8-bit prefix, and the value of the fixed bit length prefix is the coded bit data of the 8 continuous bits before the bit currently processed.
[0022] In an embodiment, the value of the variable length prefix is the coded bit data within the bit range of the quantization factor to which the bit currently processed belongs and before the bit currently processed.
[0023] In an embodiment, the number of the prefixes of different lengths is 2.
[0024] In a second aspect, the embodiments of the present application further provide a data compression device based on arithmetic coding, comprising:
[0025] a prediction and quantization unit configured to input the data to be compressed into the SZ compressor to perform prediction and quantization processing, and obtain a quantization factor sequence composed of a plurality of quantization factors;
[0026] a probability prediction unit configured to perform probability prediction on each bit in the quantization factor sequence according to the probability model to determine the probability distribution of each bit, wherein the probability model further comprises using prefixes of different lengths to obtain the coded bit data to perform probability prediction on the bit currently processed, and the prefixes of different lengths comprise a fixed bit length prefix set based on the LZ compression algorithm and a variable length prefix set based on the number of quantization factor bits;
[0027] an arithmetic coding unit configured to perform arithmetic coding on the quantization factor sequence according to the probability distribution of each bit to generate compressed data.
[0028] In a third aspect, the embodiments of the present application further provide a computer device, comprising:
[0029] a central processing unit, a memory, and an input and output interface;
[0030] The memory is a transitory storage memory or a persistent storage memory.
[0031] The central processing unit is configured to communicate with the memory and perform instruction operations in the memory to execute the method of any one of the above.
[0032] In a fourth aspect, the embodiments of the present application further provide a computer readable storage medium having a computer program stored thereon, and the computer program is executed by a processor to perform the method of any one of the above.
[0033] From the above technical solution can be seen, the embodiments of the application have the following advantages:
[0034] The arithmetic coding-based data compression method provided by the application introduces two types of prefixes with fixed length and variable length for probability prediction, the variable length prefix based on the number of bits of the quantization factor can mine the compressibility of the quantization factor in the symbol frequency, and the fixed bit length prefix can mine the repeated fragments before and after, so as to avoid the conflict between Huffman coding and LZ compression in data structure processing, realize more efficient compression of the quantization factor, improve the overall compression rate and keep the calculation overhead controllable, and be suitable for large-scale floating point number compression scenarios in high-performance computing. BRIEF DESCRIPTION OF DRAWINGS
[0035] In order to more clearly illustrate the technical solutions in the embodiments of the application or the prior art, the drawings needed to be used in the embodiments or prior art description will be briefly introduced below. Obviously, the drawings in the following description are only embodiments of the application, and for those skilled in the art, other drawings can be obtained without creative labor on the basis of the provided drawings.
[0036] Figure 1 A working flow diagram of an SZ compressor provided by the embodiments of the application;
[0037] Figure 2 A compression rate comparison diagram under different entropy coding strategies provided by the embodiments of the application;
[0038] Figure 3 An architecture diagram of an SZ quantization factor compressor based on arithmetic coding provided by the embodiments of the application;
[0039] Figure 4 A flow diagram of an arithmetic coding-based data compression method provided by the embodiments of the application;
[0040] Figure 5 A diagram showing problems existing in context modeling using fixed multi-length prefixes in a traditional arithmetic coding provided by the embodiments of the application;
[0041] Figure 6 A context modeling method diagram based on structure alignment and prefix truncation strategy provided by the embodiments of the application;
[0042] Figure 7 A comparison diagram showing the influence of using different prefix combinations and different prefix structures on the compression file size provided by the embodiments of the application;
[0043] Figure 8A write and read performance comparison chart of each compression scheme based on a multi-core parallel environment is provided for the embodiment of the present application.
[0044] Figure 9 A data compression device structure schematic diagram based on arithmetic coding is provided for the embodiment of the present application.
[0045] Figure 10 A computer device structure schematic diagram is provided for the embodiment of the present application. DETAILED DESCRIPTION
[0046] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments of the present application. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative labor fall within the scope of protection of the present application.
[0047] With the wide application of high-performance computing (HPC) in scientific simulation, climate modeling, astrophysics and other fields, the storage and transmission of massive floating-point data face great bandwidth and storage pressure. Therefore, efficient, lossless or approximately lossless compression of floating-point data has become an important technical direction to improve the overall I / O performance.
[0048] At present, one of the widely used floating-point compressors in the HPC scene is the SZ compressor. The workflow of the SZ compressor is as follows Figure 1As shown, scientific data first undergoes necessary data format conversion and data block partitioning in the preprocessing module. The prediction module then uses a specific prediction algorithm (such as linear regression or curve fitting) to predict the value of the current data point. The difference between the predicted value and the original value is output as the prediction error. The quantization module then converts each prediction error into an integer value, known as a quantization factor, within a specified error bound. Errors that are too large to be effectively quantized are marked as outliers. Furthermore, the Huffman Coding module performs entropy coding on the quantization factors. High-frequency values in the quantization factors are assigned shorter codewords, while low-frequency values are assigned longer codewords. The Huffman Coding module ultimately outputs Huffman-encoded codewords. The lossless compression module combines the Huffman-encoded codewords, outliers, and metadata (used to reconstruct the original data structure during decompression, including compression parameters, quantization step size, data size, and other information) and then performs overall compression again (for example, using algorithms such as LZ4 and zlib). The final compressed data file is written to disk.
[0049] Existing SZ compressors mainly use two methods to explore the compressibility of data:
[0050] (1) Huffman Coding is used to entropy encode the quantization factor by taking advantage of the difference in symbol frequency. For example, if the quantization factor 00010 appears 90 times and the quantization factor 11101 appears 10 times, then 1 bit is used to encode the quantization factor 00010 and 4 bits are used to encode the quantization factor 11101.
[0051] (2) Use LZ algorithm (such as LZ77, LZ4) to mine the repetitiveness of the same fragments in the bitstream.
[0052] While Huffman coding and the LZ family of algorithms each possess excellent compression properties, their combined use inherently conflicts. If Huffman coding is performed first and then the LZ algorithm is applied, the resulting codewords, which are variable in length and difficult to align, will destroy the repetitive patterns of the original data. This makes it difficult for the LZ algorithm to detect repetitive patterns at the byte granularity, significantly reducing the compression effect.
[0053] Conversely, if Huffman encoding is performed after LZ compression, the LZ compression treats the entire bit stream as a "byte sequence" and replaces the repeated part with "reference position + length". The data after LZ compression is replaced by the instruction marker (such as [offset, length]), which causes the input to be not quantization factors but data mixed with reference instructions and part of the original symbols when Huffman encoding is performed again. Huffman cannot recognize each quantization factor as a coding unit and cannot encode the probability distribution of the quantization factor at the granularity of the quantization factor. Huffman can only encode the data stream replaced by the instruction marker, so it is difficult to fully play the role of entropy encoding.
[0054] Reference is made to Figure 2 , Figure 2 For comparison of compression rates under different entropy encoding strategies, it can be seen from the schematic diagram that the selection of entropy encoding strategies has a significant impact on the compression rate on four different typical scientific data sets, namely CESM-ATM (Community Earth System Model - Atmosphere Model), EXAALT (Molecular Dynamics at the Exascale), Hurricane ISABEL (Hurricane ISABEL numerical simulation data set), and HACC (Hardware / Hybrid Accelerated Cosmology Code). On different data sets and different error limits, the compression rate of Huffman encoding followed by LZ compression (Huff+LZ4) is generally better than that of LZ followed by Huffman encoding (LZ4+Huff). It can be seen that the recognition and replacement of LZ4 for repeated fragments, the change of the original data structure, and the suppression of the recognition effect of the subsequent Huffman quantization factor, thereby suppressing the compression rate. The compression rate of the combination of the above two methods also proves that the entropy encoding strategy of Huffman combined with LZ has an inherent conflict in structure and granularity, and it is difficult to achieve optimal compression through linear combination. Especially under a relatively loose error limit (such as 1e-2), the compression rate of the combination of LZ compression algorithm and Huffman encoding is not much higher than that of single Huffman.
[0055] Based on this, the modeling strategy based on arithmetic encoding proposed in the present application can effectively unify the two types of compression characteristics, achieve more efficient context modeling without destroying the data structure, and thus improve the overall compression rate.
[0056] The following will be combined Figure 3An arithmetic coding based SZ quantization factor compressor for implementing the method provided by the embodiments of the present application is introduced and described. As shown in Figure 3 The arithmetic coding based SZ quantization factor compressor includes three-layer functional frameworks: a SZ compressor framework, a model-aware compression framework, and a context-based modeler framework. The three frameworks cooperate with each other to realize a complete process from input of original floating-point data to output of compressed data.
[0057] I. SZ compressor framework, which converts the data compression problem into a quantization factor compression problem. The applicant considers that after the original data to be compressed is processed by prediction and prediction error calculation, the original value is no longer saved, and the main data for subsequent processing is the quantization factor, which accounts for most of the bits in the compression result. In addition, because outliers are relatively few, the number of metadata is also extremely small. Therefore, whether the compression rate is high or not mainly depends on whether the quantization factor can be efficiently encoded and compressed.
[0058] The SZ compressor framework is responsible for the first half of the compression process of the original floating-point data, including prediction, quantization, and basic encoding, specifically including:
[0059] (1) Curve Fitting Predictor (Fitting Predictor): The original floating-point data is predicted according to the trend of the data to generate a predicted value sequence.
[0060] (2) Prediction Error Quantizator (Prediction Error Quantizator): The error between the original data and the predicted value is calculated, and the error is quantized to obtain quantization factors.
[0061] (3) Encoder (Encoder): The quantization factors and outliers (such as prediction failure points) are further encoded and processed.
[0062] II. Arithmetic coding based quantization factor compression module (Framework of MAC / Framework of Modeling-aware Compression)
[0063] This module replaces the traditional Huffman+LZ combination strategy to perform structure-aware modeling and arithmetic coding compression on the quantization factors, mainly including the following three core components:
[0064] (1) Structure-aware Preprocessor: The main task of the structure-aware preprocessor is to reduce the length of the input bitstream while preserving the integrity of the quantization factor boundaries. Without a structure-aware preprocessor, all quantization factors are usually stored as a fixed integer byte number, for example, each quantization factor occupies 4 bytes, which will inevitably introduce a large number of redundant high-order zeros. The structure-aware preprocessor statistically analyzes the number of valid bits of all quantization factors, determines the minimum number of bits required for the quantization factor, and re-encodes each quantization factor accordingly, representing it with the minimum number of bits, removing the invalid high-order zeros in the integer storage, thereby generating a structured bitstream, reducing the length of the input bitstream while maintaining clear quantization factor boundaries.
[0065] (2) Context-based Modeler:
[0066] This module is the core modeling module, which builds a probability prediction model to guide the arithmetic coding compression process. The modeler contains multiple submodules:
[0067] ①Prefix Extractor
[0068] According to the current bit to be encoded, two types of prefixes are extracted forward:
[0069] (1) Fixed-length prefix (e.g., 8 bits, used to exploit LZ-like compressibility)
[0070] (2) Variable-length prefix (equal to the quantization factor bit width, used to mine Huffman-like frequency differences)
[0071] ②Prefix Context Converter
[0072] Convert the prefix to a unique context index for looking up the corresponding probability distribution.
[0073] ③History Info(History information)
[0074] Record the frequency statistics of 0 / 1 in each context as a basis for prediction.
[0075] ④Probability Output
[0076] Based on the statistics of the current prefix context in the historical information, the probability distribution of the current bit being 0 or 1 is calculated.
[0077] ⑤Actual Bit Checker
[0078] Verify the actual value of the current bit and update historical information to continuously improve modeling accuracy.
[0079] (3) Confidence-based Encoder:
[0080] Mixer: used to fuse multiple predicted prefix results.
[0081] The final probabilities are passed to the arithmetic encoder to drive bit compression and output a continuous compressed bit stream.
[0082] SSE (Secondary Symbol Estimation): Performs fine-grained adjustments and probability corrections to the initial predicted probabilities through table lookup or function, improving the accuracy of the predicted probabilities and thus making the compression rate of arithmetic coding closer to the optimal level.
[0083] The SZ quantization factor compressor based on arithmetic coding provided in this application is jointly modeled by a context-based modeler and two types of prefixes, so that the SZ compressor retains both the ability of Huffman coding to utilize the frequency characteristics of the quantization factor and the ability of the LZ algorithm to recognize and process repetitive structures. The two are unified under the arithmetic coding framework to avoid structural conflicts, and ultimately can also improve the compression rate and maintain good computing performance and scalability.
[0084] It can be understood that the SZ quantization factor compressor based on arithmetic coding described in the embodiment of the present application is for the purpose of more clearly illustrating the technical solution of the embodiment of the present application, and does not constitute a limitation on the technical solution provided in the embodiment of the present application. Ordinary technicians in this field can know that with the evolution of system architecture and the emergence of new business scenarios, the technical solution provided in the embodiment of the present application is also applicable to similar technical problems.
[0085] The data compression method based on arithmetic coding provided by the embodiment of the present application is described in more detail below with reference to the accompanying drawings.
[0086] The embodiment of the present application provides a data compression method based on arithmetic coding, such as Figure 4 As shown, the method is applied to an SZ compressor including a probability model, and includes steps S410-S430.
[0087] S410: Inputting the data to be compressed into the SZ compressor for prediction processing and quantization processing to obtain a quantization factor sequence consisting of multiple quantization factors;
[0088] The data to be compressed refers to an array of floating-point numbers, such as scientific simulation data (climate temperature, particle coordinates, etc.) generated in HPC, which is usually continuous and has a certain trend of original values.
[0089] By using the fitting predictor in the SZ compressor, the value of the next data point is predicted by fitting the value change between adjacent data points, so as to minimize the difference between the original data and the predicted value, and reduce the complexity of subsequent entropy coding. For example, the original data is [100.2, 100.3, 100.5], and the predictor may predict the next point as [100.7].
[0090] Then the prediction error quantizer is used to calculate the error (difference) between the actual value and the predicted value, and the error is discretized according to the set error limit (such as relative error ≤1e-2), and the result is an integer, which is called quantization factor. For example, the error is 0.2, and the minimum unit allowed by the precision is 0.1, so the quantization factor is 0.2 / 0.1=2. The final output of this stage is a quantization factor sequence, and each bit of the sequence is coded at the bit level in the subsequent stage.
[0091] S420: Probability prediction of each bit in the quantization factor sequence according to the probability model to determine the probability distribution of each bit;
[0092] Step S420 mainly uses the probability model to predict whether each bit of each quantization factor is 0 or 1. The probability is then passed as input to the arithmetic encoder to compress the bits into shorter codes.
[0093] The probability model of the present application is actually a prefix-based context model. In simple terms, to predict whether the current bit is 0 or 1, the model mainly looks at the values of the previous few bits.
[0094] The probability model of the present application also includes using different length prefixes to obtain encoded bit data to predict the probability of the currently processed bit. Here, the different length prefixes include fixed bit length prefixes based on the LZ compression algorithm, and variable length prefixes based on the number of quantization factor bits;
[0095] The fixed length prefix actually combines the compressibility concept of the LZ compression algorithm's repeated pattern recognition, but in the embodiment of the present application, instead of finding repetitions with a sliding window, short local bit data is used to count the probability of 0 or 1 usually followed by the current fixed bit length prefix. For example: the prefix 10110011 has been followed by 0 many times, so the model will predict that the probability of 0 is higher than the probability of 1.
[0096] The variable length prefix combines the compressibility idea based on frequency difference in Huffman coding, that is, there may be a structural pattern between the bits within each quantization factor, for example, if the first bit of the previous complete quantization factor is 1, then the first bit of the current quantization factor is also more likely to be 1.
[0097] The probabilities obtained by calculating the two prefixes respectively are fused to output the probability distribution of the current processing bit, so that the two types of compressibility represented by the LZ series algorithm and Huffman coding can be considered, and an accurate probability model is provided for arithmetic coding.
[0098] S430: According to the probability distribution of each bit, the arithmetic coding is performed on the quantization factor sequence to generate compressed data.
[0099] The core principle of arithmetic coding will be explained below. Arithmetic coding is not to assign a separate code word to each symbol (as in Huffman coding), but to encode the entire bit stream as a decimal interval between 0 and 1. When encoding, the interval is gradually compressed according to the probability range of each bit. The closer the predicted probability is to the actual value of the bit (for example, P(0)=0.99), the fewer the bits required for coding, and the better the compression effect. Therefore, the more accurate the probability prediction, the higher the compression rate.
[0100] For each bit to be encoded, read its probability distribution (such as P(0)=0.75, P(1)=0.25), divide the current coding interval according to the probability distribution, and select the corresponding sub-interval as the new coding interval according to the actual bit being 0 or 1. Repeat the above steps to finally compress the entire bit sequence into an accurately represented numerical interval, thereby effectively generating compressed data. This process can fully utilize the context information extracted by the prefix modeling to achieve efficient data compression.
[0101] The arithmetic coding-based data compression method provided in the present application introduces fixed length and variable length prefixes for probability prediction. The variable length prefix based on the number of quantization factor bits can exploit the compressibility of quantization factors in symbol frequency, and the fixed bit length prefix can exploit repeated fragments before and after, thereby avoiding the conflict between Huffman coding and LZ compression in data structure processing while achieving more efficient compression of quantization factors, improving the overall compression rate and keeping the computational overhead controllable, and being suitable for large-scale floating point compression scenarios in high-performance computing.
[0102] In order to reduce the coding burden and improve the overall compression efficiency, the embodiments of the present application also perform preprocessing on the quantization factor sequence before compression. Specifically, in an embodiment, before the probability of each bit in the quantization factor sequence is predicted according to the probability model to determine the probability distribution of each bit, the method of the embodiments of the present application further includes: determining the maximum value in the plurality of quantization factors; determining the minimum number of bits k required for each quantization factor according to the most significant bit of the quantization factor of the maximum value; re-encoding each quantization factor according to the minimum number of bits k to generate a quantization factor sequence composed of a plurality of quantization factors with a length of k, triggering the step of performing probability prediction on each bit in the quantization factor sequence using a prefix of different lengths to determine the probability distribution of each bit.
[0103] In SZ compression, each floating point number becomes a quantization factor after prediction. These quantization factors are usually stored in the form of an integer int type (such as 32 bits), but there may be a large number of redundant high bits 0 or sign extension bits. For example, a quantization factor with a maximum value of 60 can actually be represented by 6 bits (111100 in binary).
[0104] If the bit-level modeling and arithmetic coding are directly performed on these integer quantization factors, not only the efficiency is low, but also the modeling resources are wasted. Therefore, before the probability of each bit in the quantization factor sequence is predicted using a prefix of different lengths to determine the probability distribution of each bit, the quantization factor is preprocessed:
[0105] Step 1: Determine the maximum value
[0106] Find the quantization factor with the largest value in the current quantization factor sequence. The maximum value represents the upper limit of the representation capability of the entire sequence.
[0107] Step 2: Determine the minimum number of bits k
[0108] According to the most significant bit (i.e., the position of the highest 1) required by the maximum value in binary representation, determine how many bits each quantization factor in the entire sequence "at least" needs to represent.
[0109] For example, the quantization factor is stored in the form of an integer int type (such as 32 bits), and the maximum quantization factor is 60 (decimal). Then, the binary representation of 32 bits is 00000000 00000000 00000000 00111100. It can be seen that the highest bit is actually in the 6th bit, which means that each quantization factor can actually be represented by 6 bits (111100), avoiding the redundant 0 in the original 32 bits.
[0110] Step 3: Re-encode the quantization factor
[0111] Each quantization factor is converted into a bit string of k-bit length, and the high-bit redundant 0s are discarded to form a continuous bit stream (each quantization factor occupies k bits) as the input for the next step of arithmetic coding.
[0112] Step 4: Trigger the next step of modeling and coding
[0113] After completing the re-encoding, prefix-based probability modeling and arithmetic coding can be performed (i.e., entering the aforementioned step S420).
[0114] In the embodiments of the present application, the originally redundant and lengthy integer quantization factors are represented by the minimum number of bits k, so that all quantization factors have a fixed bit width, which is beneficial to boundary control and prefix processing in subsequent modeling, and can also reduce the coding length, improve the compression rate, reduce the computational overhead, avoid redundant modeling, and otherwise the encoder will waste computational resources when analyzing invalid high bits.
[0115] In addition to the above problem of redundant 0s, please refer to Figure 5 Through research, the applicant has found that the existing arithmetic coding has the following two typical problems when processing quantization factor bit streams with boundary structures:
[0116] Problem 1: Invalid modeling—prefix crossing the boundary of quantization factor units:
[0117] Suppose each quantization factor only occupies 17 bits, and the 15th (counting from left to right) bit of "quantization unit 3" is being coded, then when using a fixed-length prefix (such as 24 bits / 3 bytes), 24 consecutive bits will be extracted from the bit stream; since each quantization factor only occupies 17 bits, therefore the 24-bit prefix contains 10 bits of "quantization unit 2" and 14 bits of "quantization unit 3".
[0118] This will cause the following problem: the 10-bit data of non-quantization factor 3 in the prefix has no statistical or contextual relevance to the current quantization unit 3, and including this part in the probability modeling will introduce "pseudo-context", thus reducing the prediction accuracy and easily misleading the model, affecting the compression efficiency.
[0119] Problem 2: Redundant modeling—effective parts of the prefix are repeatedly used:
[0120] In existing arithmetic coding, in order to improve the prediction accuracy, multiple-length prefixes (such as 2 bytes, 3 bytes, 4 bytes, etc.) are often used for parallel context modeling.
[0121] However, in a structured quantization factor bit stream, even if the prefix lengths are different, these prefixes may have the same effective segment in content, that is, most of their prefixes are overlapping.
[0122] It can be observed from Step 2 and Step 3 that, although the prefixes of different lengths extract different ranges, they all cover most of the bits in the same quantization unit (represented by the red horizontal line); in other words, the context fragments used for modeling are actually highly similar or identical;
[0123] For example, a 16-bit prefix is selected, which can be 1010101010101010 in particular;
[0124] A 24-bit prefix: it can be 000010101010101010101010 in particular;
[0125] Then the part actually used for prediction is most likely the same "1010101010...".
[0126] However, the key point is that the system maintains an independent statistical model for each prefix length, and since these prefixes carry repetitive information, it is equivalent to multiple models learning the same context repeatedly, resulting in redundant modeling. This kind of repetition not only fails to bring additional prediction gain, but also increases the computational burden and model complexity, thereby affecting the overall compression performance.
[0127] Based on the above problems, in a feasible embodiment, the probability of each bit in the quantization factor sequence is predicted according to a probability model to determine the probability distribution of each bit, including: if the bit sequence of the current processed bit in the quantization factor it belongs to is greater than the length value of the fixed bit length prefix, the probability of the current processed bit is predicted using prefixes of different lengths.
[0128] Before encoding each bit in a quantization factor, the system will determine the bit sequence of the current processed bit in the quantization factor it belongs to (i.e., the relative position of the bit from high to low within the quantization factor). If the bit sequence of the current bit is greater than the length value of the set fixed bit length prefix (e.g., 8 bits), it means that there are already enough number of encoded bits to form a prefix context. At this time, the system will enable the prefix-based context modeling strategy, i.e., using two types of prefixes of different lengths (fixed length prefix and variable length prefix) to extract context information, make probability prediction, and complete the arithmetic coding of the current bit based on the prediction result.
[0129] If the bit sequence of the current bit does not meet the above condition (i.e., it is not enough to support a complete fixed prefix length), the system does not enable prefix modeling, but only uses the default probability distribution or a simplified model to encode the bit, in order to reduce modeling errors and save computing resources.
[0130] Exemplarily, it is assumed that each quantization factor is 17 bits after the pre-processing of simplified redundancy 0, the length of the fixed bit prefix is 8 bits, and the current processing is the 3rd bit of a quantization factor:
[0131] Because the 3rd bit is <8 (the length of the fixed bit prefix), the complete 8-bit context cannot be extracted at this time, so the prefix modeling is not enabled;
[0132] When the 9th bit is reached, 8 bits can be used to form a prefix (for example, from the 1st to the 8th bit), at which time the prefix modeling strategy can be enabled to improve the prediction accuracy and compression rate.
[0133] Further reference can be made to Figure 6 , Figure 6 The flow provided by the present application is used for efficient probability modeling of the bit stream of the quantization factor, so as to balance the high compression rate and low calculation overhead. In the figure, three consecutive quantization factors (unit1, unit2, unit3) are taken as examples to illustrate the selection method and modeling method of the prefix (such as the two prefixes in Figure 6
[0134] In an embodiment, the fixed bit length prefix is an 8-bit prefix, and the value of the fixed bit length prefix is the continuous 8-bit coded bit data before the currently processed bit.
[0135] The fixed bit length prefix is one of the prefixes provided by the present application for the context-based modeler, and the length of the fixed bit length prefix remains constant throughout the compression process. In the present embodiment, the fixed bit length prefix is set to 8 bits (i.e., 1 byte). For details, reference can be made to the green equal-length solid line pointed to by the green dashed arrow in Figure 6 , and marked with the "8bit" label, indicating that the fixed bit length prefix is an 8-bit prefix.
[0136] When the system is ready to perform probability prediction (i.e., modeling) on a bit, it will continuously extract 8 bits from the left coded bit data as the current context information.
[0137] For example, it is assumed that the current processing is the 15th bit in the bit sequence, then: the system will extract 8 continuous bits from the 7th to the 14th coded bit, such as 10110010; and then calculate the probability distribution according to whether the next bit is usually 0 or 1 when 10110010 is the prefix in the history.
[0138] The prefix is used to simulate the modeling capability of the traditional LZ compression for "local repetition mode", which can quickly identify the common byte-level repetition in the data stream, has high modeling and coding efficiency, and is easy to implement in hardware or in parallel.
[0139] In an embodiment, the variable length prefix is a value of the coded bit data located before the current processing bit and within the bit range of the quantization factor in which the current processing bit is located.
[0140] The variable length prefix, as its name implies, is not fixed in length, and its maximum length can be adjusted according to the number of bits of the quantization factor after the removal of redundant 0s. For details, refer to the red equal-length solid line portion indicated by the red dashed arrow and labeled as “variable prefix” in Figure 6 For example, in Step 1, Step 2, and Step 3, the red solid line segment extracted from the left of the current bit to be encoded (i.e., the red bit in the quantization factor) is the corresponding variable length prefix.
[0141] There are two boundary conditions for the extraction of this prefix: it must be before the current processing bit, and it must be limited within the quantization factor in which the bit is located and cannot span to the previous quantization factor.
[0142] For example, refer to the bit stream in Figure 6 , assuming that the minimum number of bits after the removal of redundant 0s for the quantization factor is 17 bits, and if the current processing bit is the 16th bit of the red portion in Unit 3 in Step 2, the variable length prefix can be taken from the 15th bit forward to the first bit of Unit 3. In Step 2, even though the bits in Unit 1 and Unit 2 have been encoded, any bit in Unit 1 and Unit 2 cannot be used.
[0143] The variable length prefix can be truncated at the boundary of the quantization factor structure and will not take the data of the previous bit forward. This corresponds to the modeling of the “symbol frequency difference” in Huffman coding with a single quantization factor as the boundary, which helps to accurately model the probability law within the quantization factor and avoids misjudgment caused by the prefix crossing the boundary.
[0144] Please refer to Figure 7 , Figure 7 , which shows the final compressed file data size after using different prefix combinations for compression modeling on the CESM-ATM data set. The horizontal axis in the figure represents the prefix combination used, and the vertical axis is the total amount of compressed data in bytes, with the unit being bytes (the smaller the value, the better the compression effect). Figure 7 In the figure, the blue column represents the General Modeling (general modeling method, i.e., the traditional method), which does not consider the quantization factor structure boundary and uses the original bit stream of the data to be compressed for modeling. The yellow column represents the Structure-based Modeling (structure-aware modeling, i.e., the basic application method), which indicates that the quantization factor boundary can be perceived and aligned, and the variable length prefix is truncated according to the quantization factor boundary. Figure 7In the figure, the horizontal axis labels such as "10001" and "11001" indicate the enabled prefix combinations, representing, from left to right, whether 1-byte, 2-byte, 3-byte, 4-byte, and 6-byte prefixes are enabled. For example, "10001" indicates that 1-byte and 6-byte prefixes are used (5-byte prefixes are generally not used), and 2-, 3-, and 4-byte prefixes are not used; "11001" indicates that 1-, 2-, and 6-byte prefixes are used, but 3- and 4-byte prefixes are not used.
[0145] In traditional arithmetic coding, enabling more prefix combinations generally helps improve prediction accuracy, thereby increasing compression. However, each prefix requires its own probability model. As the number of prefixes increases, the computational effort increases linearly, but the compression effect may reach a stage of diminishing returns. It can be understood that if the number of prefixes is N, and the average computational effort per prefix is t (generally speaking, the computational effort for different prefixes is comparable), the final computational effort can be approximately equal to Nt + m, where m is a constant representing the modeling-related constant overhead in arithmetic coding that is independent of the number of prefixes. This includes, but is not limited to: the fixed computational effort for initializing the probability model (for example, creating a basic frequency table); conventional arithmetic coding logic outside the main encoding / decoding loop; and fixed metadata processing logic.
[0146] In order to further control the modeling computational burden while maintaining a significant improvement in the compression rate, in a feasible embodiment, the number of prefixes of different lengths is 2.
[0147] from Figure 7 As can be seen from the figure, as the number of enabled prefixes increases (horizontally from left to right), the compressed file size shows an overall downward trend; however, for the blue column (general modeling method), the compression effect is limited; for the yellow column (structure-aware encoding method provided by this scheme), after enabling the prefix combination "10001" ("10001" represents from left to right whether the 1-byte, 2-byte, 3-byte, 4-byte, and 6-byte prefixes are enabled. Because the prefix will be truncated at the quantization factor unit in structure-aware encoding (this scheme), and because the unit length of the quantization factor is usually at most 4 bytes, enabling the 6-byte prefix is equivalent to enabling the aforementioned variable-length prefix), the file size is already the lowest point among the compressed data sizes of all prefix combinations. Compared with using one prefix, the compressed data size is also significantly reduced. Continuing to add other prefixes (such as "11111") does not bring a significant improvement in the compression rate, but may waste computing resources due to model redundancy.
[0148] Therefore, this application only retains two prefixes: one is a fixed 8-bit length prefix set based on the LZ algorithm, which corresponds to the compressibility source of LZ-type compression; the other is a variable-length prefix set based on the quantization factor bit position, which corresponds to the compressibility source of Huffman-type compression. This covers the compression concepts of both types of algorithms, greatly reduces the prefix modeling computational load, and avoids prefix redundancy and boundary crossing problems.
[0149] like Figure 7 As shown, under "10001+ structure-aware modeling," the compression rate is near-optimal, eliminating the need for additional prefixes. Experimental results demonstrate that, using only two complementary prefixes (an 8-bit prefix and a unit quantization factor length prefix), the proposed structure-aware modeling strategy achieves compression comparable to or even better than multi-prefix modeling, while effectively controlling modeling complexity and system resource overhead. This strategy performs particularly well on the CESM-ATM dataset, demonstrating the effectiveness of the proposed solution in balancing compression rate and computational efficiency.
[0150] In one embodiment, the method further includes: obtaining the context probability distribution of the currently processed bit based on the fixed bit length prefix and the variable length prefix respectively; and fusing the obtained context probability distributions to determine the final probability distribution of the currently processed bit.
[0151] As mentioned above, the present application solution introduces two types of context prefix modeling methods, a fixed-length prefix (eg, 8 bits) and a variable-length prefix (ie, unit quantization factor length).
[0152] The specific operation process is as follows:
[0153] Separate modeling: For the current bit to be encoded, the system uses the two prefixes as context, queries their respective historical statistical information, and obtains two probability distribution results, for example: fixed bit length prefix prediction: P1(0)=0.7, P1(1)=0.3; variable prefix prediction: P2(0)=0.5, P2(1)=0.5.
[0154] In order to integrate the modeling capabilities of the two types of contexts, the system introduces a fusion mechanism; weighted averaging, confidence weighting, maximum entropy selection, etc. can be used to obtain the final probability distribution. For example, the probability distribution of the current bit is P_final(0)=0.6, P_final(1)=0.4. Finally, the probability distribution of the bit is sent to the arithmetic encoder, and the current bit is compressed according to the probability distribution.
[0155] Fusing two different prefixes can complement each other in capturing context. The combined prediction is more accurate, which improves the arithmetic coding compression rate and avoids the problem of redundant calculation of multiple prefixes in traditional methods.
[0156] The arithmetic coding based data compression method provided in the present application shows obvious performance advantages in the HPC scenario. Please refer to Figure 8 , Figure 8 The performance comparison experiment results of different compression schemes on the high-performance computing HPC platform are shown. The experiment is based on the CESM-ATM data set, and in the multi-core parallel environment (from 1024 cores to 8192 cores), the compression, write-out, decompression and read-in time performance of a plurality of compression schemes are compared.
[0157] Specifically, Figure 8 The different compression schemes contained in the present application are as follows:
[0158] orisz: traditional SZ compressor process (original SZ, traditional SZ compressor based on prediction-quantization-Huffman coding).
[0159] adt-fse-sz: SZ compressor based on adaptive data transcoding and finite state entropy algorithm (Adaptive Data Transcoding + Finite State Entropy with SZ, which is an improved version of SZ based on ADT and FSE algorithm).
[0160] mac-sz: SZ compression method based on arithmetic coding proposed in the present application (Model-based Arithmetic Coding with SZ, which uses context modeling and arithmetic coding to replace Huffman and LZ type compression algorithm to improve compression rate and parallel efficiency).
[0161] lpaq-sz: SZ compressor based on lightweight PAQ algorithm (Lightweight PAQ (Prediction by Partial Matching and Arithmetic Coding) with SZ, which is a lossless compression based on SZ compressor combined with LPAQ algorithm).
[0162] zfp: ZFP compressor, a lossy compression algorithm for floating point data.
[0163] none: no compression (i.e. the original data is directly written / read, used as a benchmark for comparison).
[0164] Figure 8 The upper half of the total time of the write stage, including compression time (orange) + write-out time (red). Figure 8 The lower half of the total time of the read stage, including decompression time (orange) + read-in time (red).
[0165] In HPC scenarios, the data generated by scientific simulations is huge in size, and I / O efficiency becomes a performance bottleneck. The goal of a compressor is to balance the computation overhead and I / O saving: the higher the compression rate, the less data written out, and the shorter the I / O time. But the more complex the compression algorithm, the longer the computation time required. Therefore, a trade-off between "compression time consumption" and "I / O saving brought by compression rate" is needed. The traditional Huffman encoding + LZ type compression method has the following problems in this scenario: limited compression rate, resulting in limited I / O reduction, and unable to fully utilize the computing power of more cores.
[0166] The scheme (mac-sz) of the present application replaces the original combination of Huffman encoding and LZ type compression method with arithmetic encoding, and performs more fine-grained bit-level modeling on the quantization factor, while introducing a fixed bit length prefix (8-bit prefix) and a variable length prefix. By controlling the number of prefixes to be 2, linear expansion of parallel computing overhead is avoided. Thanks to this, the compression rate of the present application is significantly improved without significantly increasing the compression time, thereby greatly saving the I / O time.
[0167] For details, see Figure 8 , Figure 8 The upper half shows that, at 1024 cores, the compression time consumption of mac-sz and orisz is close; as the number of cores grows (2048→4096→8192), at 4096 cores and above, the total compression + write-out time of mac-sz is lower than orisz for the first time. At 8192 cores, mac-sz is even better than all the comparison schemes, so the improved compression rate of the scheme of the present application makes the disk write-out data less, and in a high-core environment, the overall write-out time is reduced and the throughput performance is improved.
[0168] Figure 8 In the reading stage of the lower half, mac-sz is slightly slower at 1024 and 2048 cores; but at 8192 cores, its decompression + reading total time is lower than that of original SZ and most schemes, the decompression time is well controlled, and there is no significant increase in overhead due to the use of arithmetic encoding, and based on the foregoing experimental results, it is known that the file size obtained according to the method provided by the present application is smaller, so the reading I / O time is greatly reduced.
[0169] It can be seen from the experiment that the mac-sz, that is, the data compression method based on arithmetic coding provided in the application, can ensure the improvement of compression rate and show better compression / decompression and I / O throughput performance in a high core number scenario, which verifies the feasibility and effectiveness of the application scheme in HPC application. The specific advantages of the data compression method based on arithmetic coding in the application are as follows: the compression rate is improved significantly (the quantization factor compression rate is improved by 10%-28%, and the file compression rate is improved by up to 14%); it is more advantageous in a high concurrency scenario, especially when the core number is greater than 4096, the total write and read time consumption decreases obviously. The compression time consumption growth is controllable, the calculation and I / O overhead are balanced and optimal in a multi-core architecture, which is more suitable for modern HPC cluster deployment, and has broad application potential in scientific simulation, climate, astronomy and other big data scenarios.
[0170] Figure 8 The total write and read time consumption of the MAC-SZ compression method proposed in the application and other mainstream compression schemes under different core number configurations is shown. In a parallel computing environment of 1024 cores to 8192 cores, as the core number increases, the overall I / O throughput performance is continuously improved due to the significant improvement of the compression rate of the method, and the compression+write performance is better than that of the original SZ when the core number is greater than 4096. The experimental results fully verify the practical application value and expansion potential of the application scheme in a high-performance computing cluster.
[0171] In order to realize the data compression method based on arithmetic coding in the application embodiment, the application embodiment further provides a data compression device based on arithmetic coding, as shown in Figure 9 The device comprises:
[0172] A prediction and quantization unit 901 is configured to input the to-be-compressed data into the SZ compressor for prediction processing and quantization processing, so as to obtain a quantization factor sequence composed of a plurality of quantization factors.
[0173] A probability prediction unit 902 is configured to perform probability prediction on each bit in the quantization factor sequence according to the probability model, so as to determine the probability distribution of each bit; wherein the probability model further comprises using different length prefixes to obtain the encoded bit data to perform probability prediction on the currently processed bit, and the different length prefixes comprise a fixed bit length prefix set based on the LZ compression algorithm, and a variable length prefix set based on the number of quantization factor bits.
[0174] An arithmetic coding unit 903 is configured to perform arithmetic coding on the quantization factor sequence according to the probability distribution of each bit, so as to generate compressed data.
[0175] Based on the hardware implementation of the above program modules, and in order to implement the arithmetic coding-based data compression method provided in the embodiments of the present application, the embodiments of the present application further provide a computer device, as shown in Figure 10 The computer device 1000 includes:
[0176] A central processing unit 1001, a memory 1002, and an input / output interface 1003.
[0177] The memory 1002 is a transitory storage memory or a persistent storage memory.
[0178] The central processing unit 1001 is configured to communicate with the memory 1002 and execute instruction operations in the memory 1002 to perform any of the above arithmetic coding-based data compression methods.
[0179] Of course, in actual application, each component in the computer device 1000 is coupled together through a bus system 1004. It can be understood that the bus system 1004 is used to realize the connection and communication between the components. The bus system 1004 includes a data bus, a power supply bus, a control bus, and a state signal bus. However, in order to clearly illustrate, all kinds of buses are marked as the bus system 1004 in the Figure 10
[0180] The memory 1002 in the embodiments of the present application is used to store various types of data to support the operation of the computer device 1000. Examples of these data include any computer programs used to operate on the computer device 1000.
[0181] It can be understood that when the processor in the above computer device executes the computer program, the functions of each unit in the above corresponding device embodiments can also be implemented, which will not be described here. Exemplarily, the computer program can be divided into one or more modules / units, which are stored in the memory and executed by the processor to complete the embodiments of the present application. One or more modules / units can be a series of computer program instruction segments capable of completing a specific function, which are used to describe the execution process of the computer program in the computer device. For example, the computer program can be divided into the units in the above computer device, and each unit can realize the specific functions as described above with respect to the corresponding computer device.
[0182] The computer device can be a desktop computer, a notebook computer, a palm computer, a cloud server, and the like. The computer device can include, but is not limited to, a processor and a memory. Those skilled in the art can understand that the processor and the memory are merely examples of the computer device, and do not constitute a limitation on the computer device, and the computer device can include more or fewer components, or combine certain components, or different components, for example, the computer device can also include an input / output device, a network access device, a bus, and the like.
[0183] The processor can be a central processing unit (CPU), and can also be other general-purpose processors, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, and the like. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, and the like. The processor is a control center of the computer device, and connects all parts of the computer device through various interfaces and lines.
[0184] The memory can be used to store computer programs and / or modules, and the processor realizes various functions of the computer device by running or executing the computer programs and / or modules stored in the memory, and calling data stored in the memory. The memory can mainly include a program storage area and a data storage area, wherein the program storage area can store an operating system, at least one application required by a function, and the like; and the data storage area can store data created according to use of the terminal, and the like. In addition, the memory can include a high-speed random access memory, and can also include a non-volatile memory, for example, a hard disk, a memory, a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, at least one disk storage device, a flash memory device, or other volatile solid-state memory device.
[0185] The embodiment of the present application further provides a computer readable storage medium, which stores a computer program, and the computer program is executed by a processor to perform the data compression method based on arithmetic coding.
[0186] The embodiment of the present application further provides a computer program product, which stores a computer program / instruction, and the computer program / instruction is executed by a processor to implement the arithmetic coding-based data compression method described in the first aspect of the embodiment of the present application or any specific implementation manner of the first aspect.
[0187] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working processes of the system, the device and the unit described above can refer to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0188] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic, and the division of the units is merely a logical function division, and there can be another division manner in actual implementation, for example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between the units can be indirect couplings or communication connections through some interfaces, devices or units, and can be electrical, mechanical or in other forms.
[0189] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, that is, can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments.
[0190] In addition, each functional unit in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware, or in the form of software functional units.
[0191] The integrated unit, if implemented in the form of a software function unit and sold or used as an independent product, can be stored in a computer readable storage medium. Based on such understanding, the technical solutions of the present application essentially or the part that contributes to the prior art or the whole or part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a plurality of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present application. The aforementioned storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, Read-Only Memory), a random access memory (RAM, Random Access Memory), a magnetic disk or an optical disk, and various media that can store program codes.
Claims
1. A data compression method based on arithmetic coding, characterized in that: Applied to an SZ compressor including a probability model, the method comprises: Inputting the data to be compressed into the SZ compressor for prediction processing and quantization processing to obtain a quantization factor sequence consisting of multiple quantization factors; Probability prediction is performed on each bit in the quantization factor sequence according to the probability model to determine the probability distribution of each bit; wherein the probability model further comprises obtaining encoded bit data using prefixes of different lengths to perform probability prediction on the currently processed bit, the prefixes of different lengths including a fixed bit length prefix set based on the LZ compression algorithm and a variable length prefix set based on the number of bits in the quantization factor; According to the probability distribution of each bit, arithmetic coding is performed on the quantization factor sequence to generate compressed data.
2. The method according to claim 1, characterized in that Before performing probability prediction on each bit in the quantization factor sequence according to the probability model to determine the probability distribution of each bit, the method includes: determining a maximum value among a plurality of said quantization factors; Determining the minimum number of bits k required for each quantization factor based on the most significant bit of the quantization factor of the maximum value; Each of the quantization factors is re-encoded according to the minimum number of bits k to generate a quantization factor sequence consisting of multiple quantization factors of length k, and a step of performing probability prediction on each bit in the quantization factor sequence using prefixes of different lengths to determine the probability distribution of each bit is triggered.
3. The method according to claim 1, characterized in that The performing probability prediction on each bit in the quantization factor sequence according to the probability model to determine the probability distribution of each bit includes: If the position order of the currently processed bit in the quantization factor to which it belongs is greater than the length value of the fixed bit length prefix, prefixes of different lengths are used to perform probability prediction on the currently processed bit.
4. The method according to claim 3, characterized in that The method further comprises: Based on the fixed bit length prefix and the variable length prefix, respectively, the context probability distribution of the currently processed bit is obtained; The obtained context probability distributions are fused to determine the final probability distribution of the currently processed bit.
5. The method according to any one of claims 1 and 3, characterized in that The fixed bit length prefix is an 8-bit prefix, and the value of the fixed bit length prefix is the continuous 8-bit encoded bit data before the currently processed bit.
6. The method according to any one of claims 1 and 3, characterized in that The value of the variable-length prefix is: the encoded bit data that is located before the currently processed bit and belongs to the bit range of the quantization factor where the currently processed bit is located.
7. The method according to any one of claims 1 and 3, characterized in that The number of prefixes of different lengths is 2.
8. A data compression device based on arithmetic coding, characterized in that: Applied to SZ compressors involving probabilistic models, including: A prediction and quantization unit, configured to input the data to be compressed into the SZ compressor for prediction processing and quantization processing, and obtain a quantization factor sequence consisting of a plurality of quantization factors; a probability prediction unit, configured to perform probability prediction on each bit in the quantization factor sequence according to the probability model to determine a probability distribution of each bit; wherein the probability model further comprises obtaining encoded bit data using prefixes of different lengths to perform probability prediction on the currently processed bit, the prefixes of different lengths including a fixed bit length prefix set based on the LZ compression algorithm and a variable length prefix set based on the number of bits in the quantization factor; An arithmetic coding unit is used to perform arithmetic coding on the quantization factor sequence according to the probability distribution of each bit position to generate compressed data.
9. A computer device, characterized in that: include: CPU, memory and input / output interfaces; The memory is a transient storage memory or a persistent storage memory; The central processing unit is configured to communicate with the memory and execute instructions in the memory to perform the method according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is performed.