Quality-guaranteeing compression method based on XOR coding small error

By introducing a small error-healthy compression method based on XOR encoding in time series data compression, the error control and calculation efficiency problems in floating-point data compression are solved, and efficient compression ratio and stable performance are achieved.

CN120017072APending Publication Date: 2025-05-16NINGBO INST OF TECH ZHEJIANG UNIV ZHEJIANG
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411890710.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-20
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

When facing floating-point data, it is difficult to ensure error control and computational efficiency while retaining high compression ratios, especially in lossy compression. How to effectively utilize error tolerance without losing critical information is a technical challenge.

Method used

A small error-healthy compression method based on XOR encoding is proposed. Through the SXC and SMC methods, the design within the maximum error limit is ensured that each decoded data point is within the predefined error tolerance, while taking into account efficient compression ratio and execution efficiency.

Benefits of technology

Achieve better space savings with smaller error tolerance, SMC performs excellently in compression ratio and performance stability, and provides optimal compression and decompression speeds, demonstrating strong reliability and stability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120017072A_ABST
    Figure CN120017072A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and provides a quality guarantee compression method based on XOR coding small errors, which comprises an SXC method and an SMC method. The SXC method comprises the steps that if the content outside a variable region [v-, v +] of one double-precision value v is the same as the corresponding part of the other double-precision value u, v'is set to be v, the erasing digit mask is set to be j, and erasing is not conducted again; otherwise, for the variable region, setting the corresponding part with u to be smaller than v-or the corresponding part with u to be larger than v +, setting v'to be v +, setting the erasure digit mask to be j-1, needing to re-erase, and obtaining an XOR result xorstored = u + v '; the SMC method comprises the following steps: directly selecting v + and setting the mantissa of the v + to be zero. The efficient compression ratio and the efficient execution efficiency can be considered at the same time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a quality-preserving compression method based on XOR coding with small errors. Background Art

[0002] With the rapid development of the Internet of Things (IoT) and big data, the scale of time series data has grown exponentially, placing higher demands on data compression technology. Especially in the fields of medicine and finance, the rapid growth of time series data has become a prominent feature of the big data era, and how to effectively manage and store this data has become increasingly important. In order to reduce storage and transmission costs while retaining the core features of the data, compression technology plays a vital role. Existing time series data compression methods can be divided into two main categories: lossy compression and lossless compression.

[0003] Lossy compression techniques compress data by allowing a certain degree of error. Common methods include maximum error bound (L∞-bound) compression. This method ensures that the error of each data point does not exceed the predefined tolerance range. Typical lossy compression techniques include wavelet compression and piecewise linear approximation (PLA). Wavelet compression decomposes time series data into coefficients at multiple scales and achieves compression by selectively retaining some coefficients, while the PLA method approximates time series data by using line segments.

[0004] In recent years, Sim-Piece and its extended version Mix-Piece, as new lossy compression methods, have been improved based on the PLA method. They optimize the compression ratio and execution efficiency by grouping similar paragraphs, and have shown excellent compression performance.

[0005] In contrast, lossless compression technology completely retains the original data without losing any information, ensuring that the original data can be fully restored after decompression. Common lossless compression methods include incremental encoding, which encodes the difference between consecutive data points rather than the absolute value. The Gorilla method is a typical lossless compression method that encodes floating-point time series data through XOR operations to achieve efficient compression. Chimp proposed an improved solution that further improved the compression and decompression speed by optimizing the processing of leading zeros in XOR operations. Recently, Elf / Elf+ proposed to optimize the trailing zeros after XOR operations through preprocessing, thereby improving the compression ratio and execution efficiency.

[0006] Although existing compression technologies have achieved certain results in applications in different fields, they still have some shortcomings when facing floating-point time series data. Existing compression methods often find it difficult to ensure error control and computational efficiency during the compression process while retaining a high compression ratio. Especially when it comes to lossy compression, how to effectively use error tolerance without losing key information remains a technical challenge. Summary of the invention

[0007] The content of the present invention is to provide a quality compression method based on XOR coding with small errors, which can be designed under the background of maximum error limit to ensure that each decoded data point is within a predefined error tolerance while taking into account high compression ratio and execution efficiency.

[0008] According to the present invention, a quality-preserving compression method based on XOR coding with small errors includes an SXC method and an SMC method;

[0009] The SXC method is:

[0010] If the variable area of ​​a double value v[v - ,v + ] is the same as the corresponding part of another double-precision value u, set v' to v, the number of bits to be erased mask to j, and do not erase again;

[0011] Otherwise, for the variable region, the corresponding part of u is less than v - , or the corresponding part of u is greater than v + , set v' to v + , the number of bits to be erased is mask j-1, and it needs to be erased again, and the XOR result is

[0012] The SMC method is:

[0013] Directly select v + And set its mantissa to zero.

[0014] Preferably, the variable region [v - ,v + ]=[v-δ, v+δ], δ is the error tolerance.

[0015] Preferably, re-erasing is: when restoring the compressed value from the XOR result, the non-contributing part appears again in the restored value, and the same mask is used again for erasing.

[0016] As a preference, compression is followed by encoding, which is the process of xoring the result to be stored. stored Become a flag bit and some bits;

[0017] When encoding using the SXC method, the number of leading zeros CLZ is expressed as follows:

[0018]

[0019] When |v|<δ, v'=0 is selected; subsequent erase and re-erase operations will also result in xor stored is 0, which is the same as the XOR result when v=u; therefore, set mask to 63 to distinguish these cases, and use 2 flag bits to indicate these cases;

[0020] We choose to use 21 as the offset value to represent the length of the mantissa, i.e. mask-21, which requires 5 bits to represent. Then, we use 1 bit to indicate whether a re-erase has occurred. Next, we write xor stored The number of leading zeros CLZ in the middle, and then write the rest, i.e. the center bit;

[0021] When encoding using the SMC method, for large error tolerance, a bias value of 37 is used, i.e., mask-37, which requires 4 bits to represent; for small error tolerance, the original bias value of 21 is retained, which requires 5 bits to represent; in addition, the re-erase flag is removed, and SMC does not need to store re-erase. When the storage of the trailing zeros and re-erase is considered at the same time, the space required by SMC does not exceed the space required by SXC, shortening the storage of the trailing zeros and flags.

[0022] The beneficial effects of the present invention are as follows:

[0023] The present invention is suitable for processing streaming spatiotemporal data and the error after processing is within a specified range, that is, it is suitable for application in real-time scientific computing and analysis scenarios. The present invention ensures that each decoded data point is within a predefined error tolerance while taking into account efficient compression ratio and execution efficiency.

[0024] Compared with existing excellent lossy and XOR-based lossless floating-point compression methods, the SXC method achieves better space savings than many methods at a smaller error tolerance. At the same time, the SMC method performs well in compression ratio at a smaller error tolerance and maintains stable high performance. In addition, SMC always provides the best compression and decompression speed. In addition to the impressive compression ratio, stable mean absolute error (MAE) and root mean square error (RMSE) values, SXC and SMC also demonstrate strong reliability and stability. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 This is a flow chart of a quality-preserving compression method based on XOR encoding with small errors in an embodiment;

[0026] Figure 2 A schematic diagram of erasing and re-erasing in an embodiment;

[0027] Figure 3 is a schematic diagram of a variable region in an embodiment;

[0028] Figure 4 Schematic diagram of the coding scheme of SXC in the embodiment;

[0029] Figure 5 Schematic diagram of the coding scheme of SMC in the embodiment;

[0030] Figure 6 It is a schematic diagram of the compression ratio ratio of the four methods of SXC, Sim-Piece+, Mix-Piece and Elf+ to SMC in the embodiment;

[0031] Figure 7 These are the MAE and RMSE-MAE images of the four lossy compression methods: SXC, SMC, Sim-Piece+, and Mix-Piece in the embodiments. DETAILED DESCRIPTION

[0032] In order to further understand the content of the present invention, the present invention is described in detail in conjunction with the accompanying drawings and embodiments. It should be understood that the embodiments are only for explaining the present invention and are not intended to limit it.

[0033] Example

[0034] like Figure 1 As shown, this embodiment provides a quality-preserving compression method based on XOR coding with small errors, which includes an SXC method and an SMC method;

[0035] The SXC method is:

[0036] If the variable area of ​​a double value v[v - ,v + ] is identical to the corresponding part of another double precision value u, that is, any v∈[v - ,v + ] and i<=j, b i (u) = b i (v), where b i (u), b i (v) refers to the i-th bit in the binary representation of u and v, i and j represent the i-th and j-th bits in the binary representation; set v' to v, and the number of bits to be erased mask is j (from right to left), and no re-erasure is performed; so that for i <j,b i (v') = b i (u). Therefore, for i <j,

[0037] Otherwise, for the variable region, the corresponding part of u is less than v -, or the corresponding part of u is greater than v + , set v' to v + , the number of bits to be erased is mask j-1, and it needs to be erased again, and the XOR result is That is, for i ≥ j, b i (u) i (v - ) or b i (u)>b i (v + ). Set v'=v+, mask=j-1, re-erase=True. Therefore, (Erase), b i (xor stored )=“100.....0” for i≥j and its length is the shortest, that is, 1.

[0038] The SMC method is:

[0039] When the exponent remains unchanged, the SXC method selects the shortest mantissa part, but adds an additional 1-bit overhead to indicate whether re-erasure is needed. When analyzing the SXC strategy, it was observed that the two choices for the decimal part differed by only 1 bit. The space saved by directly selecting the slightly longer mantissa part and removing the re-erasure flag offsets the additional cost of the longer decimal part. This adjustment not only simplifies the operation, but also significantly reduces the computational overhead, thereby saving a lot of time.

[0040] At least one value with a mantissa of all '0' can be easily identified when the exponent changes. To achieve a shorter XOR mantissa, the complexity of determining the optimal choice between them should be avoided. Instead, by directly choosing v + The process is simplified by setting its mantissa to zero. This method simplifies the operation while maintaining efficiency.

[0041] Due to the limitation of computer representation of floating point numbers, many decimals cannot be represented exactly in binary form, unlike decimal integers. For example, the binary representation of 0.9 actually corresponds to the decimal value 0.90000000000000000222… It is considered equal to 0.9 only within a certain precision. Otherwise, it cannot be accurately represented with limited binary bits. However, within a given error tolerance, we can intentionally introduce representation errors of floating point numbers to achieve the desired floating point value. For a positive double precision floating point value v = 3.15 and an error tolerance of 0.01, a simple way is to set some bits in its binary representation to 0, thereby reducing the value of v.

[0042] ​This way we can represent 3.15 as 3.140625 given an error of 0.01, which requires far fewer bits in the binary representation. This is the basic prototype of the "erase" operation we want to introduce.

[0043] like Figure 2 As shown, the erase operation converts v1=3.15 to v1'=3.140625 by setting several consecutive bits to '0' at the end of the binary double precision representation. The part set to 0 is considered to have no contribution to the XOR operation, and the length of the set 0 is called "mask". For double precision values, mask≤52.

[0044] When restoring the compressed value from the XOR result, the part that did not contribute appears again in the restored value, so the same mask is used again to erase it, which is called "re-erasing".

[0045] For a given double precision value v and an error tolerance δ, Figure 3 As shown, the selectable value range of v is [v - ,v + ] = [v-δ, v+δ] is called the variable region, δ is the error tolerance. It is clear to identify the interval [v - ,v + ] is the (longest) changing digit (or shortest significant digit) within ].

[0046] Let bin(v) be the binary representation of v.

[0047]

[0048] where b i (v)∈{0,1} for i=-t,···,h. Let To meet b -j (v - ⊕v + )=1 and For i>j The first leftmost bit of . Then the feasible region F(v,δ) is [2 -j -1, 2 -j ], which shows that [v - ,v + ] The shortest significant digit usually has j digits.

[0049] The technical issues of SXC and SMC are stated as follows:

[0050] SXC chooses v'∈[v - ,v +Make the significant digits of u the shortest. Let γ represent the number of the shortest significant digits obtained. Assume that the feasible region F(v,δ) = [2 -j -1, 2 -j . For i > j, we set b i (v’) = b i (v); b j (v') = 1; for i < j, we set b i (v') = b i (u). Then, without considering the exponent change, the shortest number of significant digits γ1 of u is at least the second shortest. That is:

[0051]

[0052] For the exponent change, set b j (v') = 0 for 1 ≤ j ≤ 52. Then, traverse [Exp(v - ) + 1, Exp(v + )], and make Exp(v') any number in this interval. In this case, the shortest significant bits γ2 of u. The shortest XOR result can be obtained from γ1 and γ2.

[0053] SMC selects v' ∈ [v - , v + that has the shortest significant bits and can be efficiently implemented with the minimized flag bits. From the obtained feasible region F(v,δ) = [2 -j -1, 2 -j , b i (v') = b i (v) for i > j, b j (v') = 1, and b i (v′) = 0 for i < j. It is easy to find v' with the shortest significant bits from [v - , v + .

[0054] After compression, encoding is performed. Encoding is to turn the result xor stored to be stored into flag bits and some bits;

[0055] When encoding by the SXC method, as Figure 4 shown, this embodiment improves the encoding scheme of Chimp to better adapt to the characteristics of the XOR values generated by the SXC method. Specifically, for the representation of leading zeros, this concept is extended from Chimp and further improved to handle the additional cases generated by the variable number of leading zeros.

[0056] The representation of the number of leading zeros CLZ is as follows:

[0057]

[0058] When |v|<δ, v'=0 is selected; subsequent erase and re-erase operations will also result in xor stored is 0, which is the same as the XOR result when v=u; therefore, mask is set to 63 to distinguish these cases, and 2 flag bits are used to indicate these cases.

[0059] When xor stored ≠0, under experimental conditions, the accuracy of the data is usually no more than 10 bits, while the accuracy of the newly generated numbers by SXC is always kept below 8 bits. Based on the research results, it is determined that these data representations require a maximum of 31 significant digits in the mantissa. Therefore, 21 is chosen as the offset value to represent the length of the mantissa, that is, mask-21, which requires 5 bits to represent; then, 1 bit is used to indicate whether re-erasure occurs; next, write xor stored The number of leading zeros CLZ is written, and then the remainder, which is the center bit, is written.

[0060] When encoding the SMC method, Figure 5 As shown in the figure, in order to limit the length of the XOR valid digit, for large error tolerance, the offset value is 37, that is, mask-37, which requires 4 bits to represent; for small error tolerance, the original offset value is retained as 21, which requires 5 bits to represent; in addition, the re-erase flag bit is removed, because SXC is in the case of xor stored ≠ 0, in which case the trailing zero length of the XOR valid number generated by SMC is at most 1 bit shorter than that of SXC. However, SMC does not need to store "re-erasure", and when the storage of trailing zeros and re-erasures is considered at the same time, the space required by SMC does not exceed that required by SXC, shortening the storage of trailing zeros and flag bits.

[0061] experiment

[0062] SXC and SMC are compared with the lossy floating point compression methods Sim-Piece and Mix-Piece, as well as the lossless method Elf. Since Sim-Piece and Elf have optimized their methods in their latest versions, eventually named Sim-Piece+ and Elf+ respectively, for the sake of fairness, these enhanced versions will be used for the comparison.

[0063] Among the latest maximum error bound compression methods, Sim-Piece+ and Mix-Piece have attracted much attention due to their excellent performance. In order to obtain more concise and convincing conclusions from the experimental evaluation, this embodiment only selected these two methods for comparative testing instead of many other existing methods. This embodiment also selected Elf+, a recent lossless compression method that has performed well in compression ratio, as a benchmark for comparative testing.

[0064] Experimental Results

[0065] The compression ratio comparison results are as follows: Figure 6 As shown in Table 1. When the error tolerance is small, SMC and SXC are in an absolute advantage. When the error tolerance increases to 0.1%, the compression rate of SMC is 0.152, slightly higher than Mix-Piece's 0.149, but still the second best. As a lossless compression method, Elf+ performs similarly to Sim-Piece+ on some data sets, and even surpasses some lossy methods at a smaller error tolerance. Figure 6 Among them, (1) the partial error tolerance is 0.005%, (2) the partial error tolerance is 0.01%, (3) the partial error tolerance is 0.05%, 1) the partial error tolerance is 0.1%.

[0066] SMC and Mix-Piece have obvious advantages on specific datasets, while Sim-Piece+, Mix-Piece and Elf+ are more sensitive to the characteristics of the dataset. In contrast, the compression ratios of SXC and SMC are relatively stable.

[0067] The trends of MAE and RMSE-MAE are as follows: Figure 7 As shown in (1) and (2) in . For SMC and SXC, the MAE values ​​are similar and remain relatively stable under different error tolerances, indicating that the error distribution after compression is relatively uniform. Figure 7 (2) in Fig. 2 shows the impact of large errors on the error distribution. The curves of SMC and SXC are relatively stable, indicating that they have little impact on extreme errors.

[0068] At error tolerances of 0.22% and 3.85%, Mix-Piece and Sim-Piece+ surpass SMC and SXC in terms of MAE, while the MAE of Sim-Piece+ and Mix-Piece increases rapidly when it exceeds 3.85%. This trend is consistent with the change in the average compression ratio, as PLA-based methods usually achieve better compression ratios at higher MAE values.

[0069] In general, SMC and SXC show high stability under different error tolerances. Figure 7The curves in show that the compression performance of these two methods is not easily affected by extreme errors and has strong generalization ability. Therefore, these two methods are suitable for a wider range of application scenarios. Although Mix-Piece is more stable than Sim-Piece+, both are still limited under smaller error tolerance and are more susceptible to extreme errors. Therefore, SMC and SXC perform better in tasks that require stability and robustness, while Sim-Piece+ and Mix-Piece may have certain advantages in specific scenarios with smaller error tolerance, but their robustness is weaker.

[0070] Table 1 Average compression ratio of five compression methods (the smaller the better).

[0071]

[0072]

[0073] Table 2 Average compression time per 500 data of five compression methods (the smaller the better), in μs.

[0074]

[0075] Table 3 Average decompression time per 500 data of five compression methods (the smaller the better), in μs.

[0076]

[0077] Tables 2 and 3 show the average compression and decompression time of the five compression methods under error tolerances ranging from 0.005% to 0.1%. In terms of compression time, SMC always shows the shortest compression time, outperforming other methods, and this trend remains stable under larger error tolerances. In contrast, Mix-Piece and Sim-Piece+ have significantly longer compression times, while the advantage of SXC gradually weakens and is eventually surpassed by Sim-Piece+. As a lossless method, Elf+ has a longer compression time (e.g., 25.51μs at δ=0.005%). In terms of decompression, SMC and SXC maintain high efficiency, and the decompression time is significantly lower than Sim-Piece+, Mix-Piece, and Elf+. For example, at δ=0.005%, the decompression time of SMC is 8.41μs, while Mix-Piece is 59.52μs and Elf+ is 18.25μs.

[0078] The decompression time of SMC and SXC remains stable under different error tolerances, indicating their high efficiency in decompressing data. The decompression time of Sim-Piece+ and Mix-Piece is longer, especially when the error tolerance increases. Although Elf+, as a lossless compression method, has a better decompression time, its performance is still affected by the trade-off between lossless compression and compression ratio.

[0079] In general, SMC and SXC are the most efficient methods in compression and decompression.

[0080] The present invention and its embodiments are described schematically above, and the description is not restrictive. The drawings show only one embodiment of the present invention, and the actual structure is not limited thereto. Therefore, if a person skilled in the art is inspired by it and designs a structural method and an embodiment similar to the technical solution without creativity without departing from the purpose of the invention, they shall all fall within the protection scope of the present invention.

Claims

1. A quality-preserving compression method based on XOR coding with small errors, characterized in that: Including SXC method and SMC method; The SXC method is: If the variable area of ​​a double value v[v - ,v + ] is the same as the corresponding part of another double-precision value u, set v' to v, the number of bits to be erased mask to j, and do not erase again; Otherwise, for the variable region, the corresponding part of u is less than v - , or the corresponding part of u is greater than v + , set v' to v + , the number of bits to be erased is mask j-1, and it needs to be erased again, and the XOR result is xor stored =u⊕v'; The SMC method is: Directly select v + And set its mantissa to zero.

2. The quality-preserving compression method based on XOR coding with small errors according to claim 1 is characterized in that: Variable region [v - ,v + ]=[v-δ, v+δ], δ is the error tolerance.

3. The quality-preserving compression method based on XOR coding with small errors according to claim 2, characterized in that: Re-erasing means that when restoring the compressed value from the XOR result, the part that did not contribute appears again in the restored value and is erased again using the same mask.

4. The quality-preserving compression method based on XOR coding with small errors according to claim 3 is characterized in that: After compression, encoding is performed. Encoding is to convert the result to be stored into xor stored Become a flag bit and some bits; When encoding using the SXC method, the number of leading zeros CLZ is expressed as follows: When encoding with the SXC method, when |v|<δ, v'=0 is selected; subsequent erase and re-erase operations will also result in xor stored is 0, which is the same as the XOR result when v=u; therefore, set mask to 63 to distinguish these cases, and use 2 flag bits to indicate these cases; We choose to use 21 as the offset value to represent the length of the mantissa, i.e. mask-21, which requires 5 bits to represent. Then, we use 1 bit to indicate whether a re-erase has occurred. Next, we write xor stored The number of leading zeros CLZ in the middle, and then write the rest, i.e. the center bit; When encoding using the SMC method, for large error tolerance, a bias value of 37 is used, i.e., mask-37, which requires 4 bits to represent; for small error tolerance, the original bias value of 21 is retained, which requires 5 bits to represent; in addition, the re-erase flag is removed, and SMC does not need to store re-erase. When the storage of the trailing zeros and re-erase is considered at the same time, the space required by SMC does not exceed the space required by SXC, shortening the storage of the trailing zeros and flags.