Mass hydrology and water conservancy data compression method and system based on small-scale stationary characteristics
By distinguishing between small-scale and large-scale hydrological and water conservancy data, and employing iterative compression and normal distribution approximation, as well as integer matrix storage and encoding of small-scale data, the problem of insignificant compression effect in existing technologies is solved, achieving efficient data compression and precision control.
Patent Information
- Application Number
- CN202511014684.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-23
- Publication Date
- 2025-11-14
AI Technical Summary
Existing compression algorithms are not effective in compressing massive hydrological and water conservancy data with small-scale stationary characteristics. They fail to effectively utilize the scale characteristics of hydrological and water conservancy data, resulting in poor compression performance.
By constructing a compression method for massive hydrological and water conservancy data based on small-scale stationary features, we distinguish between small-scale and large-scale data, adopt iterative compression and normal distribution approximation to represent small-scale data, and use integer matrix storage and fixed-length encoding or Huffman encoding for compression.
It achieves significant compression of massive hydrological and water conservancy data while controlling data accuracy loss, improving compression efficiency, and is suitable for efficient storage and transmission of watershed data.
Smart Images

Figure CN120956281A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of massive data compression technology, and more specifically, to a method and system for compressing massive hydrological and water conservancy data based on small-scale stationary characteristics. Background Technology
[0002] With the further advancement of monitoring of watershed hydrological and water conservancy data, the amount of such data in watershed areas is constantly increasing. This includes information such as water level, temperature, humidity, water vapor evaporation, and rainfall. Watershed hydrological and water conservancy data plays an important role in watershed ecological protection and flood control.
[0003] In the field of massive hydrological and water resources data compression technology, especially for massive watershed spatiotemporal data, data compression strategies directly affect data analysis, data application, and data transmission. Traditional compression algorithms, such as the LZ series and Huffman coding, are mainly designed for general data. When faced with massive hydrological and water resources data that have small-scale stationary characteristics, their compression effect is generally not significant because they do not differentiate the scale of the hydrological and water resources data and do not utilize the small-scale stationary characteristics of the hydrological and water resources data.
[0004] Some targeted compression algorithms, such as linear prediction residual coding, although they pay attention to the different properties of hydrological and water conservancy data, fail to take into account the relative characteristics of hydrological and water conservancy data at different scales. They do not focus on the stationary characteristics that exist at small scales relative to large scales, and they do not combine traditional optimized coding methods for integers. Therefore, their compression effect is not significant. Summary of the Invention
[0005] This invention addresses the technical problems existing in the prior art by providing a method and system for compressing massive hydrological and water conservancy data based on small-scale stationary features. By constructing a data approximation representation of the original data with controllable accuracy loss, the information entropy of small-scale regions in the data stream is effectively reduced.
[0006] According to a first aspect of the present invention, a method for compressing massive hydrological and water conservancy data based on small-scale stationary features is provided, comprising: Step S1: Obtain hydrological and water conservancy data, determine the watershed dataset, and input the time vector dataset of hydrological and water conservancy data for the specific watershed; Step S2: Distinguish between small-scale and large-scale hydrological and water conservancy data. While focusing on analyzing the changing characteristics of large-scale data, iterative compression is used for small-scale data. Step S3: Calculate the mean and standard deviation of the small-scale dataset, and approximate the small-scale data using a normal distribution; Step S4: Calculate the deviation of each data point in the small-scale dataset from the mean, and approximate the deviation by taking an integer number of resolution standard deviations, and construct an integer matrix for storage; Step S5: Distribute all small-scale data within a preset range, and use multiple bits to sequentially encode each integer in the integer matrix to compress the data, or use Huffman coding to compress the data according to the different characteristics of different integer distribution frequencies.
[0007] Based on the above technical solution, the present invention can also be improved as follows.
[0008] Optionally, the hydrological and water conservancy data includes: latitude, longitude, and time; the time vector dataset for inputting specific watershed hydrological and water conservancy data includes: Based on the analysis of the longitude and latitude of the basin, input the specific longitude and latitude time vector dataset.
[0009] Optionally, the hydrological and water conservancy data time vector is divided into: hourly scale, daily scale, weekly scale, monthly scale, and yearly scale; the distinction between small-scale and large-scale hydrological and water conservancy data includes: Small scale and large scale are relative terms. Short-term data is considered small-scale, and long-term data is considered large-scale. For example, daily-scale data is considered small-scale relative to annual-scale data.
[0010] Optionally, when focusing on analyzing the changing characteristics of large-scale data, the iterative compression of small-scale data includes: When focusing on the weekly scale variation characteristics of data, the daily scale data is approximated by the average value every 24 hours, while the hourly data is compressed using an iterative compression method. When focusing on the characteristics of monthly data changes, the weekly data is approximated by the average value every 7 days, while an iterative compression method is used to compress the daily data.
[0011] Optionally, the deviation of each data point in the small-scale dataset from the mean is expressed as follows:
[0012] In the formula, Represented as data The discrete values of the distribution, that is, the data Mapped to ; To take the integer closest to x; It is data from a small-scale dataset in the input data. For average value, Here, k represents the standard deviation and the resolution.
[0013] Optionally, the step of approximating the deviation using an integer number of resolution standard deviations includes: For any data All can be mapped to ,use Approximate representation ; For any small-scale dataset, the average value needs to be stored. Standard deviation and corresponding .
[0014] Optionally, the step of using multiple bits to sequentially encode each integer in the integer matrix to compress the data, or using Huffman coding to compress the data based on different integer distribution frequency characteristics, includes: If the decoding difficulty is classified as difficult, then Huffman coding is used to compress the data based on the different frequency characteristics of different integer distributions, resulting in the most significant compression effect; if the decoding difficulty is classified as easy, then multiple bits are used to sequentially encode each integer in the integer matrix at a fixed length, resulting in a relatively low compression ratio.
[0015] Optionally, when evaluating the compression ratio of fixed-length coding and Huffman coding, the peak signal-to-noise ratio (PSNR) can be used to analyze the accuracy loss caused by data compression.
[0016] Optionally, the step of sequentially encoding each integer in the integer matrix with multiple bits to compress the data includes the following steps: Step S5-1: Based on the actual small-scale hydrological and water conservancy data, all are distributed in It is clear that all integers N in the integer matrix are distributed in , For resolution; Step S5-2: Calculate the number of bits required for fixed-length encoding. ; Step S5-3: Utilize Each bit encodes an integer N, with the highest bit being the sign bit, which encodes the sign information: 0 for positive numbers and 1 for negative numbers. Step S5-4: For any integer N, determine its sign bit; Step S5-5: Using the method of removing the sign bit The bits store the binary code of its absolute value.
[0017] According to a second aspect of the present invention, a massive hydrological and water conservancy data compression system based on small-scale stationary features is provided, comprising: The hydrological and water conservancy data acquisition module is used to acquire hydrological and water conservancy data and input a time vector dataset of hydrological and water conservancy data for a specific watershed. The hydrological and water conservancy data classification module is used to distinguish between small-scale and large-scale hydrological and water conservancy data. When focusing on analyzing the changing characteristics of large-scale data, iterative compression is used for small-scale data. The data storage module is used to calculate the mean and standard deviation of the small-scale dataset and approximate the small-scale data using a normal distribution; calculate the deviation of each data point in the small-scale dataset from the mean and approximate the deviation using an integer number of resolution standard deviations, and construct an integer matrix for storage; The small-scale data compression module is used to distribute all small-scale data within a preset range and use multiple bits to sequentially encode each integer in the integer matrix to compress the data, or use Huffman coding to compress the data according to the different characteristics of the frequency of different integer distributions.
[0018] The technical effects and advantages of this invention are as follows: This invention provides a method and system for compressing massive hydrological and water conservancy data based on small-scale stationary features. By efficiently utilizing the small-scale stationary features of massive hydrological and water conservancy data and through spatiotemporal scale analysis, it addresses the stationary features of small-scale data and the relatively large fluctuations in data at large scales. This ensures that the loss of data accuracy remains controllable while significantly compressing massive hydrological and water conservancy data, thus solving the problem that traditional compression algorithms are not effective for large-scale watershed data.
[0019] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures pointed out in the description, claims and drawings. Attached Figure Description
[0020] Figure 1 A flowchart illustrating the method for compressing massive hydrological and water conservancy data based on small-scale stationary features provided in an embodiment of the present invention. Figure 2 A flowchart of fixed-length encoding provided for embodiments of the present invention; Figure 3 Distribution map of hourly water vapor data in 1982 provided for embodiments of the present invention; Figure 4 The peak signal-to-noise ratio range and corresponding number of days in 1982 are provided for embodiments of the present invention. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] Understandably, given the deficiencies in the background technology, this invention proposes a method for compressing massive hydrological and water conservancy data based on small-scale stationary characteristics, specifically as follows: Figure 1 As shown, it includes the following steps: Step S1: Obtain hydrological and water conservancy data, determine the watershed dataset, and input the time vector dataset of hydrological and water conservancy data for the specific watershed; It should be noted that hydrological and water conservancy data, such as NetCDF files, are generally three-dimensional tensors. Each data point is determined by three parameters: longitude, latitude, and time. Hydrological and water conservancy data can be regarded as a three-dimensional grid. Based on the longitude and latitude of the watershed being analyzed, a time vector dataset with specific longitude and latitude is input.
[0023] Step S2: Distinguish between small-scale and large-scale hydrological and water conservancy data. While focusing on analyzing the changing characteristics of large-scale data, iterative compression is used for small-scale data. It should be noted that hydrological and water conservancy data time vectors can generally be divided into: hourly, daily, weekly, monthly, and yearly scales. Generally speaking, small and large scales are relative; shorter periods are considered small, and longer periods are considered large. For example, daily data relative to yearly data; the daily scale can be considered small, while the yearly scale can be considered large. Small-scale changes are relatively stable; when focusing on analyzing the changing characteristics of large-scale data, the average value is used. It approximates small-scale data, achieving significant data compression by sacrificing controllable accuracy.
[0024] Then, iterative compression is applied to small-scale data. For example, when focusing on the weekly scale variation characteristics of the data, the daily scale data is approximated by the 24-hour average, while the hourly data is compressed using iterative compression methods. When focusing on the monthly scale variation characteristics of the data, the weekly scale data can be approximated by the 7-day average, while the daily data is compressed similarly using iterative compression methods.
[0025] Step S3: Calculate the mean of the small-scale dataset. Standard deviation And it uses a normal distribution to approximate the small-scale data; When focusing on analyzing large-scale change characteristics, the time scale of small-scale data is also considered. , The average value can be obtained. Standard deviation
[0026]
[0027]
[0028] In the formula, This is the average value. Standard deviation; The number of data points in the small-scale dataset; It is data from a small-scale dataset within the input data.
[0029] Based on the relatively stable characteristics of small-scale hydrological and water conservancy data in actual watersheds, small-scale data (especially for daily and weekly scales) are regarded as a combination of baseline data and floating data, and then a normal distribution is used to approximate their representation.
[0030] Step S4: Calculate the deviation of each data point in the small-scale dataset from the mean, and approximate the deviation by taking an integer number of resolution standard deviations, and construct an integer matrix for storage; In this embodiment, based on the characteristic of small-scale concentration in actual watershed hydrological and hydraulic data, the deviation of each data point in the small-scale dataset from the average value is calculated. This is represented as follows:
[0031] In the formula, Represented as data The discrete values of the distribution, that is, the data Mapped to ; To take the integer closest to x; It is data from a small-scale dataset in the input data. For average value, Here, k represents the standard deviation and the resolution.
[0032] The step of approximating the deviation with an integer number of resolution standard deviations and constructing an integer matrix for storage includes: Step 4-1, for any data All can be mapped to , that is, use Approximate representation ; Step 4-2: For any small-scale dataset (such as daily 24-hour data), the average value needs to be stored. Standard deviation and corresponding That's all; Step 4-3: Store by constructing an integer matrix. For example, for compression of daily-scale data, one can utilize... An integer matrix stores the data corresponding to each of the 24 hours of a day. .
[0033] Step S5: Distribute all small-scale data within a preset range, and use multiple bits to sequentially encode each integer in the integer matrix to compress the data, or use Huffman coding to compress the data according to the different characteristics of different integer distribution frequencies.
[0034] In this embodiment, the preset range is distributed in , A value of 3 is generally taken as an empirical value. In actual hydrological and water conservancy data, small-scale watershed data are all distributed in... middle, Generally, an empirical value of 3 is taken, and the vast majority of data are distributed in... The result conforms to the normal distribution law, verifying the rationality of step S3.
[0035] The method of compressing data by sequentially encoding each integer in an integer matrix with multiple bits of fixed length, or by using Huffman coding to compress data based on different frequency characteristics of different integer distributions, includes: If a relatively high decoding difficulty is permissible, Huffman coding can be used to compress the data based on the different frequency characteristics of different integer distributions, resulting in the most significant compression effect. If a relatively low decoding difficulty is desired, multiple bits can be used to sequentially encode each integer in the integer matrix at a fixed length, resulting in a relatively low compression ratio.
[0036] In this embodiment, data compression is achieved using fixed-length coding and Huffman coding, as shown below: Fixed-length encoding: Two bits are used to encode the four integers 0, 1, 2, and 3. Then, the highest bit is used as the sign bit, meaning three bits are used to encode -3, -2, -1, 0, 1, 2, and 3 sequentially as 111, 110, 101, 000 (100), 001, 010, and 011. The fixed-length encoding flowchart is shown below. Figure 2 (Pick The step of sequentially encoding each integer in the integer matrix with multiple bits to compress the data includes the following steps: Step S5-1: Based on the actual small-scale hydrological and water conservancy data, all are distributed in It is clear that all integers N in the integer matrix are distributed in , For resolution; Step S5-2: Calculate the number of bits required for fixed-length encoding. ; Step S5-3: Utilize Each bit encodes an integer N, with the highest bit being the sign bit, which encodes the sign information: 0 for positive numbers and 1 for negative numbers.
[0037] Step S5-4: For any integer N, determine its sign bit; Step S5-5: Using the method of removing the sign bit The bits store the binary code of its absolute value.
[0038] Huffman coding: Based on the different frequencies of each integer, a Huffman tree is constructed to further compress the data.
[0039] The compression ratios of the two encoding methods are evaluated, and the accuracy loss caused by data compression is analyzed using Peak Signal-to-Noise Ratio (PSNR). The following uses the hourly water vapor data from 1982 as an example to evaluate the compression ratio and accuracy loss of the two encoding methods:
[0040]
[0041]
[0042]
[0043]
[0044]
[0045]
[0046] The first part is storage. The first part is the number of bytes required, and the second part is the number of bytes required to store the discrete value matrix; the compression ratio is 10.6408.
[0047] Using Huffman coding, the example encoding obtained by constructing a Huffman tree is (0:00; 1:01; -1:10; -2:110; 2:1110; -3:11110; 3:11111).
[0048] Without sacrificing precision, Huffman coding can achieve a compression ratio of 14.7459, which is a significant improvement over fixed-length coding.
[0049] Specific examples: such as Figure 3 The data shown is hourly water vapor data for 1982. Distribution map, based on hourly water vapor data from 1982 (data source: European Centre for Medium-Range Weather Forecasts, data type: NetCDF), focusing on the Yangtze River basin from the Three Gorges to Chenglingji, including important stations such as Yichang. Construct an integer matrix, and the integers and their frequencies are shown in Table 1.
[0050] like Figure 4 The table shows the peak signal-to-noise ratio (PSNR) data range and its corresponding number of days (distortion rate assessment). It can be seen that although the data exhibits some distortion, the accuracy loss for most data is manageable (PSNR > 20). This further verifies the effectiveness of the normal distribution approximation. Furthermore, to further improve the PSNR, the resolution can be lowered. This refines the data range, thereby improving the peak signal-to-noise ratio (PSNR) and reducing the data distortion rate.
[0051] It should be noted that the massive hydrological and water conservancy data compression method and system based on small-scale stationary features described in the embodiments of the present invention are applicable to the efficient storage and transmission of massive hydrological and water conservancy data in the Yangtze River Basin, as well as to preprocessing scenarios for large-scale basin data.
[0052] According to a second aspect of the present invention, a massive hydrological and water conservancy data compression system based on small-scale stationary features is provided, comprising: The hydrological and water conservancy data acquisition module is used to acquire hydrological and water conservancy data and input a time vector dataset of hydrological and water conservancy data for a specific watershed. The hydrological and water conservancy data classification module is used to distinguish between small-scale and large-scale hydrological and water conservancy data. When focusing on analyzing the changing characteristics of large-scale data, iterative compression is used for small-scale data. The data storage module is used to calculate the mean and standard deviation of the small-scale dataset and approximate the small-scale data using a normal distribution; calculate the deviation of each data point in the small-scale dataset from the mean and approximate the deviation using an integer number of resolution standard deviations, and construct an integer matrix for storage; The small-scale data compression module is used to distribute all small-scale data within a preset range and use multiple bits to sequentially encode each integer in the integer matrix to compress the data, or use Huffman coding to compress the data according to the different characteristics of the frequency of different integer distributions.
[0053] It is understood that the massive hydrological and water conservancy data compression system based on small-scale stationary features provided by the present invention corresponds to the massive hydrological and water conservancy data compression method based on small-scale stationary features provided in the foregoing embodiments. The relevant technical features of the massive hydrological and water conservancy data compression system based on small-scale stationary features can be referred to the relevant technical features of the massive hydrological and water conservancy data compression method based on small-scale stationary features, and will not be repeated here.
[0054] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
[0055] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0056] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
[0057] Finally, it should be noted that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A method for compressing massive hydrological and water conservancy data based on small-scale stationary characteristics, characterized in that: Includes the following steps: Step S1: Obtain hydrological and water conservancy data, and input the time vector dataset of hydrological and water conservancy for the specific watershed; Step S2: Distinguish between small-scale and large-scale hydrological and water conservancy data. While focusing on analyzing the changing characteristics of large-scale data, iterative compression is used for small-scale data. Step S3: Calculate the mean and standard deviation of the small-scale dataset, and approximate the small-scale data using a normal distribution; Step S4: Calculate the deviation of each data point in the small-scale dataset from the mean, and approximate the deviation by taking an integer number of resolution standard deviations, and construct an integer matrix for storage; Step S5: Distribute all small-scale data within a preset range, and use multiple bits to sequentially encode each integer in the integer matrix to compress the data, or use Huffman coding to compress the data according to the different characteristics of different integer distribution frequencies.
2. The method for compressing massive hydrological and water conservancy data based on small-scale stationary characteristics according to claim 1, characterized in that, The hydrological and water conservancy data includes: latitude, longitude, and time; the time vector dataset for inputting specific watershed hydrological and water conservancy data includes: Based on the analysis of the longitude and latitude of the basin, input the specific longitude and latitude time vector dataset.
3. The method for compressing massive hydrological and water conservancy data based on small-scale stationary characteristics according to claim 1, characterized in that, The time vectors of the hydrological and water conservancy data are divided into: hourly scale, daily scale, weekly scale, monthly scale, and yearly scale; The distinction between small-scale and large-scale hydrological and water conservancy data includes: Small scale and large scale are relative terms. Short-term data is considered small-scale, and long-term data is considered large-scale. For example, daily-scale data is considered small-scale relative to annual-scale data.
4. The method for compressing massive hydrological and water conservancy data based on small-scale stationary characteristics according to claim 1, characterized in that, The iterative compression of small-scale data when focusing on analyzing the changing characteristics of large-scale data includes: When focusing on the weekly scale variation characteristics of data, the daily scale data is approximated by the average value every 24 hours, while the hourly data is compressed using an iterative compression method. When focusing on the characteristics of monthly data changes, the weekly data is approximated by the average value every 7 days, while an iterative compression method is used to compress the daily data.
5. The method for compressing massive hydrological and water conservancy data based on small-scale stationary characteristics according to claim 1, characterized in that, The deviation of each data point in the small-scale dataset from the mean is expressed as follows: In the formula, Represented as data The discrete values of the distribution, that is, the data Mapped to ; To take the integer closest to x; It is data from a small-scale dataset in the input data. For average value, Here, k represents the standard deviation and the resolution.
6. The method for compressing massive hydrological and water conservancy data based on small-scale stationary characteristics according to claim 5, characterized in that, The method of approximating the deviation using an integer number of resolution standard deviations includes: For any data All can be mapped to ,use Approximate representation ; For any small-scale dataset, the average value needs to be stored. Standard deviation and corresponding .
7. The method for compressing massive hydrological and water conservancy data based on small-scale stationary characteristics according to claim 1, characterized in that, The method of compressing data by sequentially encoding each integer in an integer matrix with multiple bits of fixed length, or by using Huffman coding to compress data based on different frequency characteristics of different integer distributions, includes: If the decoding difficulty is classified as difficult, then Huffman coding is used to compress the data based on the different frequency characteristics of different integer distributions, resulting in the most significant compression effect; if the decoding difficulty is classified as easy, then multiple bits are used to sequentially encode each integer in the integer matrix at a fixed length, resulting in a relatively low compression ratio.
8. The method for compressing massive hydrological and water conservancy data based on small-scale stationary characteristics according to claim 7, characterized in that, When evaluating the compression ratio of fixed-length coding and Huffman coding, peak signal-to-noise ratio (PSNR) is used to analyze the accuracy loss caused by data compression.
9. The method for compressing massive hydrological and water conservancy data based on small-scale stationary characteristics according to claim 1, characterized in that, The step of sequentially encoding each integer data in an integer matrix using multiple bits with a fixed length includes the following steps: Step S5-1: Based on the actual small-scale hydrological and water conservancy data, all are distributed in It is clear that all integers N in the integer matrix are distributed in , For resolution; Step S5-2: Calculate the number of bits required for fixed-length encoding. ; Step S5-3: Utilize Each bit encodes an integer N, with the highest bit being the sign bit, which encodes the sign information: 0 for positive numbers and 1 for negative numbers. Step S5-4: For any integer N, determine its sign bit; Step S5-5: Using the method of removing the sign bit The bits store the binary code of its absolute value.
10. A massive hydrological and water conservancy data compression system based on small-scale stationary characteristics, characterized in that: include: The hydrological and water conservancy data acquisition module is used to acquire hydrological and water conservancy data and input a time vector dataset of hydrological and water conservancy data for a specific watershed. The hydrological and water conservancy data classification module is used to distinguish between small-scale and large-scale hydrological and water conservancy data. When focusing on analyzing the changing characteristics of large-scale data, iterative compression is used for small-scale data. The data storage module is used to calculate the mean and standard deviation of small-scale datasets and to approximate the small-scale data using a normal distribution. Calculate the deviation of each data point from the mean in the small-scale dataset, and approximate the deviation using an integer number of resolution standard deviations, then construct an integer matrix to store the data. The small-scale data compression module is used to distribute all small-scale data within a preset range and use multiple bits to sequentially encode each integer in the integer matrix to compress the data, or use Huffman coding to compress the data according to the different characteristics of the frequency of different integer distributions.