A Lossy Compression Method for Time-Series Data Combining Prediction and Coding
Through the combination of preprocessing, predictive compression and coding compression, the problem of insufficient trend correlation in lossy compression of timing data is solved, and the lossy compression effect of timing data with high compression ratio and low time overhead is achieved.
Patent Information
- Application Number
- CN202211036392.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-08-28
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-08-28
AI Technical Summary
The existing lossy compression algorithms for time series data are insufficient in taking into account the trend correlation between time series data, resulting in large errors between the decompressed data and the original data, and at the same time, the compression ratio is low and the time overhead is high.
Combining preprocessing, predictive compression and coding compression methods, by fusing data segments with the same trend, using the Holt exponential trend model and introducing linear fitting to optimize prediction accuracy, and using the BitCompress algorithm for encoding compression.
Lossy compression of timing data with high compression ratio and low time overhead is achieved, reducing the impact of timing data trend changes on prediction results, improving prediction accuracy and optimizing overall time complexity.
Smart Images

Figure CN115567058B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to a lossy compression method for time series data combining prediction and coding, belonging to the technical field of lossy compression of time series data. Background Art
[0002] Existing data compression algorithms can be classified into lossless compression and lossy compression according to whether errors are allowed. Lossless compression is used when the decompressed data is required to be exactly the same as the original data. For example, the compression of in-memory data and disk data. According to the development status of existing lossless compression technologies, an ordinary file can be compressed to 25%-50% of its original size. Lossless data compression can generally be implemented using two different mechanisms. One is entropy coding-based compression, such as Huffman coding and arithmetic coding. The other is dictionary-based compression, including LZ77, LZ78, and LZW, which are called the LZ series algorithms. Huffman coding belongs to a greedy algorithm. It needs to construct a Huffman tree before encoding, and then encode the data according to the code value 0 or 1 and the Huffman tree from top to bottom. Since the coding length is variable, both compression and decompression are very time-consuming. The LZ algorithm belongs to adaptive dictionary technology and is the most popular lossless compression algorithm when the prior statistical characteristics of the data are unknown. The LZ algorithm has been adopted by many compression format standards, such as Zip, GNUZip, and Zlib. These algorithms use a sliding window to check the input sequence. Its principle is to check whether the sequence to be compressed appears in the previously input data. If so, a pointer is used to point to the repeated string. The dictionary is a set of character sequences dynamically generated according to the sequence to be encoded. Lossy compression means compressing data within the allowable accuracy error, and the decompressed and restored data is not necessarily exactly the same as the original data. Currently, such compression algorithms are mainly applied to picture compression, audio compression, video compression, and some data for big data analysis, because selectively discarding some of the data will not cause an overall judgment of the data meaning by the program or people. In many cases, sacrificing some data to increase the compression ratio is the best choice. For example, through a blurred photo of a cat and a blurred photo of a dog, one can still distinguish between a cat and a dog.
[0003] For lossy compression, there are mainly three categories: deep learning-based algorithms, decomposition-based algorithms, and piecewise approximation-based algorithms. Traditional deep learning-based compression algorithms require a large amount of data as input to train the prediction model. The training of the prediction model takes a lot of time and it is difficult to meet the latency requirements of industrial applications. Decomposition-based algorithms include Discrete Wavelet Transform (DWT), Discrete Fourier Transform (DFT), and Singular Value Decomposition (SVD). Souza et al. applied the SVD algorithm to data compression. This algorithm first constructs a numerical matrix based on the original data, then decomposes the numerical matrix into three matrices, and finally only preserves a part of the rows and columns of these 3 matrices. When the original data is needed, the SVD algorithm restores the original data by multiplying the row and column information of the 3 preserved matrices. However, this algorithm will lose a large amount of key information. For Discrete Wavelet Transform and Discrete Fourier Transform, they have the same disadvantages as the Singular Value Decomposition method. Piecewise approximation-based algorithms include Piecewise Aggregate Approximation (PAA), Adaptive Piecewise Constant Approximation (APCA), and Piecewise Linear Approximation (PLA). PAA approximates the data by dividing the sequence into equal-length parts and re-encoding the average value of these parts. Then, these average values can be effectively indexed in a low-dimensional space. This method may miss some important information and sometimes lead to inaccurate results for time series data. The APCA algorithm divides the time series into a series of segments with variable lengths, and each segment is represented by the average value of the segment and the time metric value at the rightmost end of the segment. Compared with PAA, although the APCA method overcomes the limitation of the fixed size of the sliding window, its space complexity is twice that of the PAA algorithm. The PLA method first segments the time series and then uses a linear function to perform piecewise approximation on each segment. Although the above compression algorithms can achieve a high compression ratio, the compression time is very long. A compression algorithm based on polynomial approximation has been proposed by relevant researchers, and this algorithm is based on the least squares method. The time complexity of the least squares method is O(n 3 ). The relevant literature proposed the Slope Intercept Residual Compression Algorithm SIRCS. This algorithm applies a linear regression model to the slope, intercept, and residual decomposition of multiple data streams and combines tree mapping technology, but the time overhead of the SIRCS algorithm is very large. The latest lossy compression algorithm for time series data, LFZip, consists of a predictor, a quantizer, and an encoder. In the first stage, the Normalized Least Mean Square (NLMS) algorithm is used to estimate the data points. In the second stage, the quantizer is used to quantize the error between the predicted value and the true value. In the third stage, a compressor based on the BWT algorithm is used for the final encoding compression. The overall time complexity of LFZip is O(n 2 ).
[0004] In summary, existing prediction-based compression solutions lack consideration of the trend correlation between time-series data, resulting in a large error between the decompressed data and the original data. Although existing coding-based compression solutions can guarantee the accuracy of the decompressed data, most of them have the problems of low compression ratio and high time overhead. Summary of the Invention
[0005] Aiming at reducing the storage overhead, the present invention provides a lossy compression method for time-series data that combines prediction and coding, and realizes a lossy compression algorithm for time-series data with a high compression ratio and low time overhead.
[0006] The present invention provides a compression algorithm that combines prediction and coding, abbreviated as TPCFT. Through three parts: preprocessing, prediction compression, and coding compression, a lossy compression algorithm for time-series data with a high compression ratio and low time overhead is formed;
[0007] Preprocessing reduces the impact of the trend change of time-series data on the prediction result by fusing data segments with the same trend according to the trend correlation between adjacent data segments of time-series data;
[0008] Prediction compression is based on the Holt exponential trend model and improves its prediction accuracy by introducing linear fitting to the model. At the same time, the error calculation process of the prediction part is optimized, thereby reducing the overall time complexity of the algorithm;
[0009] The coding compression part adopts the BitCompress algorithm, which is a typical coding compression algorithm.
[0010] Preprocessing is performed by merging adjacent data segments with the same time trend in the sliding window into a new data segment, so as to reduce the impact of data trend change on the compression effect;
[0011] Assume that the data segment S = {s1, s2,..., s n-1 , s n}, the step size is w. When the step size is w, the trend of the data is represented as k, and its expression is (1.1).
[0012]
[0013] The trends of two consecutive data segments calculated by the above formula (1.1) are k1 and k2 respectively. If k1 and k2 satisfy the following formula (1.2), the two data segments can be fused into one data segment.
[0014] k1·k2≥0 (1 . 2)
[0015] If k1 and k2 do not satisfy formula (1.2), then two adjacent data segments will not be fused; the fused data segment is used as the input of the prediction compression algorithm. The fusion operation is to avoid the destruction of the correlation of the entire data segment when a certain data segment is split. Data segments with different trends will not be fused into the same data segment, thereby reducing the impact of the trend change of the data segment on the prediction result.
[0016] The prediction compression is a prediction compression algorithm based on the Holt exponential trend model, and the improvements based on the Holt exponential trend model include prediction method optimization, time overhead optimization, and linear fitting optimization.
[0017] Optimization 1: The optimization of the prediction method includes introducing linear fitting in the prediction stage to obtain the overall trend of the data segment to be processed. First, perform linear fitting on the original data to obtain a linear model; then select a certain time point as the segmentation point, and use the data before this segmentation point as the input of the Holt exponential trend model to obtain the first prediction sequence; then use the linear model to generate the data after this segmentation point, and use the generated data as the input of the Holt exponential trend model to obtain the second prediction sequence; the final prediction result takes the average value of the corresponding positions of the two prediction sequences; under the condition that there are equal probability errors in the above two prediction sequences, selecting the average value of the two prediction sequences can minimize the error.
[0018] The Holt exponential trend model is formalized as equations (1.3), (1.4), (1.5):
[0019]
[0020] l t = α·y t +(1 - α)·(l t-1 + b t-1 ) (1.4)
[0021]
[0022] In the above three formulas, b t represents the growth rate trend at time t, l t represents the horizontal trend, α represents the horizontal smoothing parameter, β represents the trend smoothing parameter, represents the predicted value at t + h;
[0023] Let R = {R1, R2,..., R n-1 , R n} be the original data segment, then the fitting sequence D = {D1, D2,..., D n-1 , D n} is generated by the linear fitting model represented by formula (1.6).
[0024]
[0025]
[0026] where x i represents the distance from the i-th point to the first point in the data segment R, y1 is the actual value of the first point in R, and y n is the actual value of the last data point in R. k can be calculated according to formula (1.7), which is the output of the linear fitting model; The reverse form of the sequence D is denoted as RD = {D n , D n-1 , …, D1, D0}, P = {P1, P2, …, P i-1} are the predicted values. When P i is used as the segmentation point, {R1, R2, …, R i-2 , R i-1} and {D n , D n-1 , …, D i+2 , D i+1} will be used as the input of the algorithm; For these two segments of data, the predicted sequences of the algorithm are P a and P b , respectively. The final prediction result can be formalized as formula (1.8),
[0027] P final = μ1·P a + μ2·P b (1.8)
[0028] where μ1 and μ2 are weight parameters;
[0029]
[0030] where R k represents the actual value, P k represents the predicted value, and n represents the lengths of the data segments P and R;
[0031] If the predicted values in P and the actual values in R satisfy equation (1.9), it is concluded that {R1, R2, …, R i-1} can be used to predict {R i , R i+1 , …, R n}; Finally, store {R1, R2, …, R i-1};
[0032] If the relationship between the average error and the threshold does not satisfy formula (1.9), the predicted value P iwill be replaced by the actual value at the corresponding position in R, and the current splitting point becomes P i+1 , repeat the above operations until a splitting point that meets the error requirement is found, and incorporate the data before this splitting point into the result set.
[0033] Optimization 2: Time overhead optimization. By introducing recursive error calculation during the prediction process and formalizing it as formula (1.10),
[0034]
[0035] where E i represents the average error between the original values {R i , R i+1 , …, R n} and the predicted values {P i , P i+1 , …, P n}, and P i is a possible splitting point.
[0036] Optimization 3: The linear fitting optimization mentioned above minimizes the overall error of the fitting result. The error between the fitting result and the true value can be formalized as equation (1.11),
[0037]
[0038] where x i represents the distance from the i-th point to the 1st point in the data segment R, y1 is the actual value of the 1st point in R, and n is the sequence length of the data segment P;
[0039] Equations (1.12) and (1.13) are obtained by equivalent transformation of equation (1.11):
[0040]
[0041]
[0042] Formula (1.14) is obtained by equivalent substitution, replacing Similarly, is replaced by , and can be replaced by Φ;
[0043]
[0044] According to equations (1.11), (1.12), (1.13), and (1.14), it can be concluded that the error error is a quadratic function of the linear fitting model parameter k. According to the optimization principle, it can be inferred that there must be an optimal k that minimizes the error between the predicted data and the original data;
[0045]
[0046] The following inequalities can be derived from equations (1.14) and (1.15):
[0047] φ≥0 (1.16)
[0048] Based on (1.14), (1.15), and (1.16), the error reaches its minimum value if and only if k satisfies formula (1.17).
[0049]
[0050] The error reaches its minimum value.
[0051] Encoding compression further compresses the data obtained from the prediction compression algorithm through a lossless compression algorithm based on variable-length coding, namely BitCompress. This includes first encoding each data bit by bit using a variable-length coding rule, then concatenating the encodings, and finally combining multiple encodings.
[0052] After prediction compression and encoding compression, the overall compression ratio can be formalized as formula (1.18).
[0053]
[0054] where the size of the original data is n and the size of the final compressed data is i.
[0055] The decompression process of the algorithm first applies encoding compression to decompress the compressed data to obtain the decompressed data, i.e., the decompression output of encoding compression. Then, the decompression output of encoding compression is used as the input of the prediction compression algorithm for prediction decompression. Finally, the output of the prediction compression algorithm is the restored data.
[0056] The beneficial effects of the present invention are as follows: by fusing data segments with the same trend, the influence of the trend change of time-series data on the prediction result is reduced. The prediction compression part is based on the Holt exponential trend model and improves its prediction accuracy by introducing linear fitting to the model. At the same time, the error calculation process of the prediction part is optimized, thereby reducing the overall time complexity of the algorithm. BRIEF DESCRIPTION OF THE DRAWINGS
[0057] Figure 1 It is a flowchart of compression and decompression of the present invention.
[0058] Figure 2 It is a schematic diagram of optimizing the prediction method.
[0059] Figure 3 It is a system framework diagram.
[0060] Figure 4 It is a system interaction timing diagram.
[0061] Figure 5 It is a system function diagram.
[0062] Figure 6 It is the encoding and decoding flowchart of the BitCompress algorithm. Specific implementation manner
[0063] To enable those skilled in the art to better understand the technical content of the present invention, the following further describes the preferred embodiments of the present invention with reference to the attached Figures 1 to 5 The present invention relates to a compression algorithm combining prediction and coding, abbreviated as TPCFT. Through three parts: preprocessing (Data Fusion), predictive compression (PCBH), and coding compression (BitCompress), a lossy compression algorithm for time-series data with a high compression ratio and low time overhead is formed;
[0064] The preprocessing reduces the impact of the trend change of time-series data on the prediction result by fusing data segments with the same trend according to the trend correlation between adjacent data segments of time-series data;
[0065] The predictive compression is based on the Holt exponential trend model and improves its prediction accuracy by introducing linear fitting to the model. At the same time, the error calculation process of the prediction part is optimized, thereby reducing the overall time complexity of the algorithm;
[0066] The coding compression part adopts the BitCompress algorithm, which is a typical coding compression algorithm.
[0067] The compression process and decompression process of the TPCFT algorithm are as Figure 1 shown. The time-series data collected by the sensor is always in a fluctuating state. In this article, the alternating appearance of the rising trend and falling trend of the data is defined as fluctuation. If these original data are directly used as the input of the predictive compression algorithm, it will inevitably lead to inaccurate prediction results and ultimately poor compression effects. Therefore, the preprocessing is performed by merging adjacent data segments with the same time trend in the sliding window into new data segments, thereby reducing the impact of data trend changes on the compression effect; the same trend means that two adjacent data segments both have a rising trend or a falling trend. After data fusion, the original data segments are transformed into several new data segments with the same trend. The overall goal of the preprocessing part is to improve the compression ratio of the TPCFT algorithm through data fusion;
[0068] Assume the data segment S = {s1, s2,..., s n-1 , s n}, where the step size is w. When the step size is w, the trend of the data is represented by k, and its expression is (1.1).
[0069]
[0070] The trends of two consecutive data segments calculated by the above formula (1.1) are k1 and k2 respectively. If k1 and k2 satisfy the following formula (1.2), then the two data segments can be fused into one data segment.
[0071] k1·k2≥0 (1 . 2)
[0072] If k1 and k2 do not satisfy formula (1.2), then two adjacent data segments will not be fused; the fused data segment is used as the input of the prediction compression algorithm. The fusion operation is to avoid the destruction of the correlation of the entire data segment when a certain data segment is split. Data segments with different trends will not be fused into the same data segment, thus reducing the impact of the change in the data segment trend on the prediction result.
[0073] For example, A = {S1,..., S w} and B = {S w+1 ,..., S 2w} are two adjacent data segments; k A , k B are calculated from data segment A and data segment B according to formula (1.2) respectively. If k A and k B satisfy formula (1.2), then the data fusion operation can be adopted to fuse A and B into a data segment C = {S1,..., S 2w}, and then data segment C will be used as the input of the second part of the TPCFT algorithm; if k A and k B do not satisfy formula (1.2), then data segments A and B will be used as the input of the second part of the TPCFT algorithm respectively.
[0074] The data collected by the sensor changes continuously over time, and most of these data are correlated. Traditional lossless compression algorithms do not consider data correlation, so the compression ratios they achieve are low, and they cannot effectively compress such data, and some data do not need to be stored in a lossless manner. Traditional lossy compression algorithms have problems such as long time consumption and may lose a large amount of useful information. Based on the above problems, prediction compression is a prediction compression algorithm based on the Holt exponential trend model, and the improvements based on the Holt exponential trend model include prediction method optimization, time overhead optimization, and linear fitting optimization.
[0075] Optimization 1: Due to acquisition errors or fluctuations in the acquired data itself, the data will change within a certain time interval. If only the data before the predicted data is used as input, the result will be inaccurate. If the overall trend of the data segment can be obtained, the prediction accuracy of the algorithm will be improved. The optimization of the prediction method includes introducing linear fitting in the prediction stage to obtain the overall trend of the data segment to be processed. First, perform linear fitting on the original data to obtain a linear model; then select a certain time point as the segmentation point, and use the data before this segmentation point as the input of the Holt exponential trend model to obtain the first prediction sequence; then use the linear model to generate the data after this segmentation point, and use the generated data as the input of the Holt exponential trend model to obtain the second prediction sequence; the final prediction result takes the average value of the corresponding positions of the two prediction sequences; under the condition that there are equal-probability errors in the above two prediction sequences, selecting the average value of the two prediction sequences can minimize the error; next, a formal description of the prediction mechanism of the algorithm will be given.
[0076] The Holt exponential trend model is formalized as equations (1.3), (1.4), and (1.5):
[0077]
[0078] l t = α·y t +(1 - α)·(l t-1 + b t-1 ) (1.4)
[0079]
[0080] In the above three formulas, b t represents the growth rate trend at time t, l t represents the horizontal trend, α represents the horizontal smoothing parameter, β represents the trend smoothing parameter, represents the predicted value at t + h; according to the recommended settings in relevant literature research, set α to 0.5, β to 0.01, and h to 1.
[0081] As Figure 2 shown, let R = {R1, R2,..., R n-1 , R n} be the original data segment, then the fitting sequence D = {D1, D2,..., D n-1 , D n} is generated by the linear fitting model represented by formula (1.6), and this operation corresponds to step 5 of algorithm 2;
[0082]
[0083]
[0084] where x i represents the distance from the i-th point to the first point in the data segment R, y1 is the actual value of the first point in R, and y n is the actual value of the last data point in R. k can be calculated according to formula (1.7), which is the output of the linear fitting model; as Figure 2 shown, the reverse form of sequence D is represented as RD = {D n , D n-1 , …, D1, D0}, P = {P1, P2, …, P i-1} is the predicted value. When P i is regarded as the segmentation point, {R1, R2, …, R i-2 , R i-1} and {D n , D n-1 , …, D i+2 , D i+1} will be used as the input of the algorithm; for these two segments of data, the predicted sequences of the algorithm are P a and P b respectively. The final prediction result can be formalized as formula (1.8),
[0085] P final = μ1·P a + μ2·P b (1.8)
[0086] where μ1 and μ2 are weight parameters; generally, to minimize the mean error, both are set to 0.5;
[0087]
[0088] where R k represents the actual value, P k represents the predicted value, and n represents the lengths of the data segments P and R;
[0089] If the predicted value in P and the actual value in R satisfy equation (1.9), it is concluded that {R1, R2, …, R i-1} can be used to predict {R i , R i+1 , …, R n}; finally, {R1, R2, …, R i-1} is stored;
[0090] If the relationship between the mean error and the threshold does not satisfy formula (1.9), the predicted value P i in P will be replaced with the actual value at the corresponding position in R, and the current segmentation point becomes P i+1, Repeat the above operations until a splitting point that meets the error requirement is found, and incorporate the data before this splitting point into the result set.
[0091] Optimization 2: Time overhead optimization. Since the PCBH algorithm continuously tries to find the most suitable splitting point, for each possible splitting point, the algorithm must predict the data after this splitting point and then calculate the average error between the predicted value and the actual data. This makes the calculation process a very time-consuming operation because it makes the overall time complexity of the algorithm O(n 2 ), where n is the sequence length of the original data;
[0092] After analysis, recalculating the error for each possible splitting point involves a large number of redundant operations, thus increasing the overall time overhead of the algorithm. To reduce the time overhead, a recursive error calculation is introduced during the prediction process and formalized as formula (1.10),
[0093]
[0094] where E i represents the average error between the original values {R i , R i+1 , …, R n} and the predicted values {P i , P i+1 , …, P n}, and P i is the possible splitting point.
[0095] Optimization 3: The linear fitting optimization mentioned above. When a linear model fits data, the fitting accuracy changes with the change of parameters. Therefore, it is necessary to select a most suitable parameter. If the parameter is selected by iterative attempts, its time complexity is uncertain, which will undoubtedly increase the time overhead of the algorithm. Therefore, this scheme is not appropriate. After analysis, the parameter of the linear model and the fitting error satisfy a quadratic function relationship. Thus, the problem of finding the optimal parameter is transformed into a convex optimization problem. The time complexity of solving this convex optimization problem is O(1). Therefore, this scheme is feasible. To minimize the overall error of the fitting result, the error between the fitting result and the true value can be formalized as formula (1.11),
[0096]
[0097] where x i represents the distance from the i-th point to the 1st point in the data segment R, y i is the actual value of the i-th point in R, and n is the sequence length of the data segment P;
[0098] Equations (1.12) and (1.13) are obtained by performing equivalent transformations on Equation (1.11):
[0099]
[0100]
[0101] Formula (1.14) is obtained by equivalent substitution, replacing with the symbol φ Similarly, being replaced, can be replaced by Φ;
[0102]
[0103] According to Equations (1.11), (1.12), (1.13), and (1.14), it can be concluded that the error is a quadratic function of the linear fitting model parameter k. According to the optimization principle, it can be inferred that there must be an optimal k that minimizes the error between the predicted data and the original data;
[0104]
[0105] From Equations (1.14) and (1.15), the following inequality can be derived,
[0106] φ ≥ 0 (1.16)
[0107] Based on (1.14), (1.15), and (1.16), when and only when k satisfies Formula (1.17),
[0108]
[0109] the error reaches its minimum value.
[0110] After compressing the data using the PCBH algorithm, the correlation between the data weakens. At this time, if the prediction-based compression algorithm is used again or no further processing is performed, the overall compression ratio cannot be significantly improved.
[0111] Encoding compression further compresses the data obtained from the prediction compression algorithm through a lossless compression algorithm based on variable-length coding, namely BitCompress. This includes first encoding each data bit by bit using a variable-length coding rule, then performing encoding concatenation, and finally combining multiple encodings; compared with traditional solutions, this algorithm eliminates the dependence on parameters by introducing an encoding mapping table, thus effectively improving the overall compression ratio of the algorithm again.
[0112] BitCompress can make full use of the free bits of the binary bits of numerical data to store other numerical data, thereby achieving the purpose of compressing data. The so-called free bits refer to the binary bits that are continuously 0 starting from the highest bit when the data is represented in binary form. The coding mapping table is shown in Table 1:
[0113] Original number 0 1 2 3 4 5 6 7 8 9 Coding unit 000 001 010 011 100 101 1100 1101 1110 1111
[0114] In this algorithm, the maximum number of bits in a data sequence is called the standard number of bits. Therefore, to ensure the consistency of encoding, the data needs to be standardized. Standardization means that for data that does not meet the standard number of bits, it is padded with 0s in front of the data to make it reach the standard number of bits. The standard number of bits determines the number of encoding units after each data is encoded.
[0115] Figure 6 The encoding and decoding processes of the BitCompress algorithm are introduced. For the data sequence in the figure, since the maximum value of the entire data sequence is 478, its standard number of bits is 3. According to the coding mapping table, after standardizing the data 45, it becomes 045. The encoding unit corresponding to the number 0 is 000, the encoding unit corresponding to the number 4 is 100, and the encoding unit corresponding to the number 5 is 101. After encoding the data 45, 3 encoding units are obtained. The same operation is performed on other data in the same sequence. To ensure the correctness of the decoding process, the standard number of bits of the entire encoding sequence needs to be marked during encoding. This algorithm achieves this purpose by designing a leading bit. The design rules of the leading bit are shown in Table 2.
[0116] Leading bit 00 01 10 11 Standard number of bits 1 2 3 4
[0117] After the data sequence is compressed, a binary sequence with a length of 64 bits is finally obtained. For Figure 6 the data sequence in, if it is directly stored in the time series database, it will occupy 48 bytes of storage space, and only 8 bytes are occupied after compression.
[0118] The decoding process is the inverse operation of the encoding process. First, the leading bits are read. The leading bits are 10, indicating that every three encoding units represent one data. For example, when decoding the data 45, first judge the first two bits of the encoding unit 000. After judgment, its first two bits are 00. According to the encoding rule, the number represented by this encoding unit must be an integer within the interval [0, 5]. If the first two bits are 11, then according to the encoding rule, it can be inferred that the number represented by this encoding unit must be an integer within [6, 9]. The decoding process for the subsequent encoding units 100 and 101 is the same. In summary, the decoding result of 000100101 is 045. After normalization, the final result 45 is obtained, which is the same as the value before encoding. Normalization means removing the 0s filled during the standardization in the encoding process.
[0119] After predictive compression and encoding compression, the overall compression ratio can be formalized as formula (1.18),
[0120]
[0121] where the size of the original data is n and the size of the final compressed data is i.
[0122] As Figure 1 shown, the decompression process of the algorithm first applies encoding compression (BitCompress) to decompress the compressed data to obtain the decompressed data, that is, the decompression output of BitCompress; then the decompression output of BitCompress is used as the input of the predictive compression algorithm (PCBH algorithm) for predictive decompression; finally, the output of the PCBH algorithm is the restored data. To meet the needs of practical applications, the compressed data should be stored in the database in a fixed format for easy query. Therefore, the table format of the compressed data in the database is designed. When querying the data, the decompression algorithm can restore the data according to the data table information.
[0123] The storage format of the data in the database is shown in Table 3, where timestamp represents the time when the data is collected. The parameter field represents the linear fitting model parameter k of the predictive compression for this segment of data, and the tag field is used to identify the data segment where the data is located, that is, the data with the same tag field belongs to the same data segment before compression and should also belong to the same data segment after restoration.
[0124] Table 3 Compressed data format
[0125]
[0126] The following introduces the process of querying data in the database. First, the database system queries data according to the timestamp, and the query result is a data block containing multiple records. The database engine merges and groups the queried data according to the tag field, and the data in the same group is used as an independent compressed data block to perform the decompression process as Figure 1 shown.
[0127] The performance of the method of the present invention is analyzed from two aspects of time complexity and compression ratio below.
[0128] The time complexity analysis process is divided into three steps, corresponding to three parts of the algorithm compression process respectively:
[0129] The first step analyzes the time complexity of preprocessing data fusion. Assume that the length of the original data sequence is n. In the data fusion process, for each data point in the sequence, only the result needs to be calculated according to formula (1.1). The overall time complexity of this part is obviously O(n);
[0130] The second step uses the prediction compression algorithm to compress data. Assume that the length of the sequence obtained from the data fusion part is n. For each data point, first calculate the error between the predicted value and the actual value according to formula (1.9). Then compare the error value with the threshold preset according to actual requirements. If the error value is greater than the threshold, select the next segmentation point and repeat the above operation until a segmentation point that meets the requirements is found. Since only the update operation needs to be performed according to the error value of formula (1.10), that is, the new error value can be obtained from the old error value within a constant time complexity. Therefore, the conclusion is drawn that the time complexity of this operation is still O(n);
[0131] The third step uses the BitCompress algorithm for compression. This algorithm is a coding compression scheme based on bit form. The whole only needs to traverse the original data once. For data that does not meet the standard number of bits, fill 0 in front of the data to make it reach the standard number of bits. The standard number of bits determines the number of coding units after each data is encoded. Therefore, the time complexity of the whole BitCompress algorithm is O(n);
[0132] Let the size of the data after being compressed by the prediction compression bag algorithm be m, and the relationship between m and u satisfies formula (2.1),
[0133] m + u = n (2.1)
[0134] After the above analysis, formula (1.2) is obtained, where the left side of the formula represents the overall time complexity of the TPCFT algorithm, and the right side represents the overall time complexity of the LFZip algorithm,
[0135] O(m 2) + O(m + u) ≤ O((m + u) 2 ) = O(n 2 ) (2.2)
[0136] The decompression process time complexity of the two algorithms is the same as the compression process time complexity analyzed above. Thus, the conclusion can be drawn that the overall time complexity of the TPCFT algorithm is lower than that of the LFZip algorithm.
[0137] Analysis of the compression ratio proves that the TPCFT algorithm is superior to the LFZip algorithm in terms of the compression ratio. Let the compression ratio of LFZip be cr, and the size of the original data segment be set to n. The size of the original data after being compressed by the LFZip algorithm can be formalized as (2.3).
[0138]
[0139] Let the compression ratio of the predictive compression bag algorithm be ct, and the size of the original data segment be set to n. The size of the original data after being compressed by the predictive compression algorithm can be formalized as (2.4).
[0140]
[0141] According to the above formulas (2.3) and (2.4), formula (2.5) is derived.
[0142]
[0143] where size F is the size of the original data after being compressed by the TPCFT algorithm, n is the size of the original data, cr is the compression ratio of LFZip, and ct is the compression ratio of PCBH. Formula (2.5) is converted into formula (2.6) through equivalent substitution.
[0144]
[0145] The left side of formula (2.6) represents the size of the original data after being compressed by the TPCFT algorithm, and the right side represents the size of the original data after being compressed by the LFZip algorithm. Through the above analysis, a clear conclusion can be drawn in this section: the compression ratio achieved by the TPCFT algorithm is higher than that of the LFZip algorithm.
[0146] Since the overall design of the system needs to meet the software design principles of low coupling and high cohesion, the overall architecture of the TPCFT-DB system is divided into three layers: the interface presentation layer, the application logic layer, and the database layer. The system architecture is as Figure 3 shown.
[0147] Figure 4Shows the interaction sequence diagram between each functional class inside the application logic layer and between the application logic layer and the other two layers. The following gives the detailed functional information of each class:
[0148] (1) The MainCall class is responsible for parsing commands from the interface presentation layer. If it is a compression command, this class will load the Compressor class. If it is a decompression command, this class will load the Decompressor class;
[0149] (2) The Compressor class is responsible for compressing data. The Compressor class contains the TPCFT compression algorithm. This class first converts the data to be compressed into a standard format; then this class executes the compression algorithm; finally, the Compressor class calls the API command of the database layer to write the compressed data into InfluxDB;
[0150] (3) The Decompressor class is responsible for decompressing data. This class first queries the data according to the query conditions passed from the interface presentation layer and encapsulates the queried data; then the Decompressor class calls the decompression algorithm to decompress the data; finally, the Decompressor class passes the decompressed data to the Interface class;
[0151] (4) The Interface class is responsible for displaying data. This class first obtains the data display panel of the interface presentation layer for the research on the edge real-time compression method for time series data; then this class fills the data and related attributes into the data display panel; finally, the Interface class renders and refreshes the data display panel.
[0152] The overall function of the system is as Figure 5 shown. The user selects each branch function, database management, data table management, user management, query data, and write data through the main interface. Database management, data table management, and user management directly interact with the database. Query data and write data interact with the database through the compression storage module. Finally, the database returns the execution result to the user.
Claims
1. A lossy compression method for time-series data combining prediction and coding, characterized in that: It forms a lossy compression algorithm for time-series data with a high compression ratio and low time overhead through three parts: preprocessing, predictive compression, and coding compression; Preprocessing reduces the impact of the trend change of time-series data on the prediction result by fusing data segments with the same trend according to the trend correlation between adjacent data segments of time-series data; Predictive compression is based on the Holt exponential trend model and improves its prediction accuracy by introducing linear fitting to this model. At the same time, the error calculation process of the prediction part is optimized, thereby reducing the overall time complexity of the algorithm; and the improvement based on the Holt exponential trend model includes prediction method optimization, time overhead optimization, and linear fitting optimization; The Holt exponential trend model is formalized as equations (1.3), (1.4), and (1.5): l t = α·y t + (1 - α)·(l t-1 + b t-1 ) (1.4) In the above three formulas, b t represents the growth rate trend of time t, l t represents the horizontal trend, α represents the horizontal smoothing parameter, β represents the trend smoothing parameter, represents the predicted value at t + h; Let \(R = \{R_1, R_2, \ldots, R\) n-1 , R n \}\) be the original data segment, then the fitting sequence \(D = \{D_1, D_2, \ldots, D\) n-1 , D n \}\) is generated by the linear fitting model represented by formula (1.6). where x i represents the distance from the i-th point to the first point in the data segment R, y1 is the actual value of the first point in R, and y n is the actual value of the last data point in R. k can be calculated according to formula (1.7), which is the output of the linear fitting model; the reverse form of the sequence D is denoted as RD = {D n , D n-1 , …, D1, D0}, P = {P1, P2, …, P i-1} are the predicted values. When P i is regarded as the segmentation point, {R1, R2, …, R i-2 , R i-1} and {D n , D n-1 , …, D i+2 , D i+1} will be used as the input of the algorithm; for these two segments of data, the predicted sequences of the algorithm are P a and P b , and the final prediction result can be formalized as formula (1.8). P final = μ1·P a + μ2·P b (1.8) where μ1 and μ2 are weight parameters; where R k represents the actual value, P k represents the predicted value, and n represents the lengths of the data segments P and R; If the predicted value in P and the actual value in R satisfy Equation (1.9), then the conclusion is drawn that {R1, R2, …, R i-1} can be used to predict {R i , R i+1 , …, R n}; finally, store {R1, R2, …, R i-1}; If the relationship between the average error and the threshold does not satisfy formula (1.9), the predicted value P in P i will be replaced by the actual value at the corresponding position in R, and the current splitting point becomes P i+1 , and the above operations are repeated until a splitting point that meets the error requirement is found, and the data before this splitting point is incorporated into the result set; For decompression, the decompression process of the algorithm first applies coding compression to decompress the compressed data to obtain the decompressed data, that is, the decompression output of coding compression; then the decompression output of coding compression is used as the input of the predictive compression algorithm for predictive decompression; finally, the output of the predictive compression algorithm is the restored data.
2. A lossy compression method for time-series data combining prediction and coding according to claim 1, characterized in that: Preprocessing is performed by merging adjacent data segments with the same time trend in the sliding window into a new data segment, so as to reduce the impact of data trend change on the compression effect; Suppose the data segment S = {s1, s2,..., s n-1 , s n}, the step size is w. When the step size is w, the trend of the data is represented as k, and its expression is (1.1), The trends of two consecutive data segments calculated by the above formula (1.1) are k1 and k respectively 2; If k1 and k2 satisfy the following formula (1.2), the two data segments can be merged into one data segment k1·k2≥0(1.2) If k1 and k2 do not satisfy formula (1.2), then two adjacent data segments will not be fused; the fused data segment is used as the input of the predictive compression algorithm. The fusion operation is to avoid the destruction of the relevance of the entire data segment when a certain data segment is split. Data segments with different trends will not be fused into the same data segment, thus reducing the impact of data segment trend change on the prediction result.
3. A lossy compression method for time-series data combining prediction and coding according to claim 1, characterized in that: The optimization of the prediction method includes introducing linear fitting in the prediction stage to obtain the overall trend of the data segment to be processed. First, perform linear fitting on the original data to obtain a linear model; then select a certain time point as the segmentation point, and use the data before this segmentation point as the input of the Holt exponential trend model to obtain the first prediction sequence; then use the linear model to generate the data after this segmentation point, and use the generated data as the input of the Holt exponential trend model to obtain the second prediction sequence; the final prediction result takes the average value of the corresponding positions of the two prediction sequences; under the condition that there is an equal probability error in the above two prediction sequences, selecting the average value of the two prediction sequences can minimize the error.
4. A lossy compression method for time-series data combining prediction and coding according to claim 1, characterized in that: For time overhead optimization, a recursive error calculation is introduced during the prediction process and formalized as formula (1.10), Among them, E i represents the average error between the original values {R i , R i+1 , …, R n} and the predicted values {P i , P i+1 , …, P n}, where P i is a possible splitting point.
5. A lossy compression method for time series data combining prediction and coding according to claim 1, characterized in that: Linear fitting optimization is performed to minimize the overall error of the fitting result. The error between the fitting result and the true value can be formalized as Equation (1.11). where x i represents the distance from the i-th point to the 1st point in the data segment R, and y i is the actual value of the i-th point in R, and n is the sequence length of the data segment P; Equations (1.12) and (1.13) are obtained by performing equivalent transformations on Equation (1.11). Equation (1.14) is obtained through equivalent substitution, replacing with the symbol φ Similarly, is replaced, and can be replaced by Φ; Based on Equations (1.11), (1.12), (1.13), and (1.14), it can be concluded that the error error is a quadratic function of the linear fitting model parameter k. According to the optimization principle, it can be inferred that there must be an optimal k that minimizes the error between the predicted data and the original data. From Equations (1.14) and (1.15), the following inequality can be derived. φ ≥ 0 (1.16) Based on (1.14), (1.15), and (1.16), the error reaches the minimum value if and only if k satisfies Equation (1.17). The error reaches the minimum value.
6. A lossy compression method for time series data combining prediction and coding according to claim 1, characterized in that: Encoding compression is performed by further compressing the data obtained from the prediction compression algorithm through a lossless compression algorithm based on variable-length coding, namely BitCompress, including first encoding each data bit by bit using a variable-length coding rule, then performing encoding splicing, and finally combining multiple encodings. After predictive compression and coding compression, the overall compression ratio can be formalized as Equation (1.18). Where the size of the original data is n and the size of the final compressed data is i.
Citation Information
Patent Citations
Lossless compression method for bridge health monitoring sensor data
CN111836045A
Methods and Computer Program Products for Compression of Sequencing Data
US20140093881A1