A data compression method and system based on a hybrid compression algorithm
Through the hybrid compression algorithm, the prediction model and segmentation processing technology are used to solve the problem of dictionary expansion and low compression ratio of the differential LZW algorithm in high-frequency timing data compression, and efficient data compression and decoding reconstruction are achieved.
Patent Information
- Application Number
- CN202510552638.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-29
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-04-29
AI Technical Summary
The existing differential LZW algorithms have problems such as infinite dictionary expansion, difficulty in adapting to strong random data and low compression rates in high-frequency timing data compression.
Using a hybrid compression algorithm, the residual sequence is generated through the prediction model, the standard deviation segments of the sliding window are used and the low variance segments are clustered to generate dictionary indexes, frequency domain transformation and quantization to process high variance segments, and combined with lossless coding seed data, an overall compressed data packet is formed.
Significantly reduce data redundancy, improve compression rate, ensure that the decoding end can accurately reconstruct the original data, adapt to the compression effect of dynamic characteristic data, and meet the data recovery needs of high-precision monitoring scenarios.
Smart Images

Figure CN120074539B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of time series measurement data compression, and in particular, to a data compression method and system based on a hybrid compression algorithm. Background Art
[0002] With the booming development of modern data centers and cloud platforms, real-time monitoring devices have been widely used, which can support the collection of monitoring indicators of various devices, such as CPU, memory, disk I / O, network traffic, etc. Specifically, real-time monitoring devices usually collect various measurement data at high frequencies and continuously, generating a huge-scale time series measurement data stream, which often has the characteristics of large data volume, high sampling frequency, and high correlation between data. However, there is usually a strong similarity or trend correlation between the data values at adjacent time points, forming a large amount of data redundancy, resulting in extremely high direct storage and transmission costs.
[0003] Currently, the commonly used data compression algorithm for time series measurement data streams is the differential LZW (Lempel-Ziv-Welch) algorithm, which stores the repeated segments in the input data in the form of dictionary "atoms" by constructing and continuously updating the data dictionary in real time. When the subsequent data appears the same as the existing pattern in the dictionary, the corresponding index value is used to replace the original data segment to achieve data compression.
[0004] However, since the differential LZW algorithm continuously adds new data patterns to the dictionary during operation, as the data stream increases and the data patterns change, the dictionary size will grow rapidly. Especially when the data source runs for a long time, the dictionary often becomes extremely large, and the storage cost of the dictionary will gradually account for an increasing proportion, even offsetting the storage savings brought by compression to a certain extent. In addition, the differential LZW algorithm compresses based on the existing repeated patterns in the dictionary, and for measurement data with frequent pattern changes or large random fluctuations, the data pattern change speed exceeds the dictionary update speed, making it difficult to capture effective common patterns in the dictionary. Summary of the Invention
[0005] This application provides a data compression method, system, storage medium, computer program product, and electronic device based on a hybrid compression algorithm, so as to at least solve the problems of infinite dictionary expansion, difficulty in adapting to strongly random data, and low compression ratio in the differential LZW algorithm for high-frequency time series data compression in the current related technologies.
[0006] In a first aspect, an embodiment of the present application provides a data compression method based on a hybrid compression algorithm, which is applied to the edge device encoding end. The method includes: obtaining a measured monitoring parameter sequence corresponding to the current sampling time period and historical measured monitoring seed data corresponding to a historical time window, and calling a first prediction model to determine a predicted monitoring parameter sequence, and calculating a residual sequence between the measured monitoring parameter sequence and the predicted monitoring parameter sequence; the predicted monitoring parameter sequence is a parameter sequence obtained by the first prediction model predicting the monitoring parameters within the current sampling time period according to the historical measured monitoring seed data; calculating the standard deviation of the residual sequence in each sliding window, so as to divide the residual sequence into a low-variance segment and a high-variance segment according to a standard deviation threshold, and generating corresponding segmentation boundary flags; performing clustering processing on the residual values of the low-variance segment to generate corresponding multiple clustering cluster features, and generating corresponding dictionary indexes according to each of the clustering cluster features; performing frequency domain transformation on the residual values of the high-variance segment to generate a corresponding frequency domain coefficient distribution, and generating a corresponding frequency domain compression code through quantization processing and entropy coding; performing lossless coding on the historical measured monitoring seed data to generate a corresponding seed data compression code, and packing and combining the segmentation boundary flags, the dictionary indexes, the frequency domain compression code, and the seed data compression code, so as to obtain measured monitoring compressed data for the measured monitoring parameter sequence.
[0007] Second aspect, an embodiment of the present application provides a data compression system based on a hybrid compression algorithm, which is deployed at the edge device encoding end. The system includes: an acquisition unit, configured to acquire the measured monitoring parameter sequence corresponding to the current sampling time period and the historical measured monitoring seed data corresponding to the historical time window, and call the first prediction model to determine the predicted monitoring parameter sequence, and calculate the residual sequence between the measured monitoring parameter sequence and the predicted monitoring parameter sequence; the predicted monitoring parameter sequence is the parameter sequence obtained by the first prediction model predicting the monitoring parameters within the current sampling time period according to the historical measured monitoring seed data; a residual sequence segmentation unit, configured to calculate the standard deviation of the residual sequence in each sliding window, and divide the residual sequence into a low variance segment and a high variance segment according to the standard deviation threshold, and generate corresponding segmentation boundary flags; a low variance segment compression unit, configured to perform clustering processing on the residual values of the low variance segment to generate corresponding multiple clustering cluster features, and generate corresponding dictionary indexes according to each of the clustering cluster features; a high variance segment compression unit, configured to perform frequency domain transformation on the residual values of the high variance segment to generate corresponding frequency domain coefficient distributions, and generate corresponding frequency domain compression encodings through quantization processing and entropy encoding; an encoding and packaging combination unit, configured to perform lossless encoding on the historical measured monitoring seed data to generate corresponding seed data compression encodings, and package and combine the segmentation boundary flags, the dictionary indexes, the frequency domain compression encodings, and the seed data compression encodings, so as to obtain the measured monitoring compression data for the measured monitoring parameter sequence.
[0008] Third aspect, there is provided an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the data compression method based on the hybrid compression algorithm according to any embodiment of the present application.
[0009] Fourth aspect, an embodiment of the present application provides a storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, it implements the steps of the data compression method based on the hybrid compression algorithm according to any embodiment of the present application.
[0010] Fifth aspect, an embodiment of the present application provides a computer program product, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, they implement the steps of the data compression method based on the hybrid compression algorithm according to any embodiment of the present application.
[0011] Through a data compression method and system based on a hybrid compression algorithm provided by the present application, at least the following technical effects can be produced:
[0012] (1) Using the prediction model constructed using historical monitoring seed data, the current sampled monitoring parameters are first predicted, thereby extracting the residual between the actual collected data and the predicted data. Since the monitoring data at adjacent time points are highly correlated, the residual value is small, which greatly reduces the redundancy of the data in the subsequent compression process and improves the overall compression rate.
[0013] (2) The residual sequence is divided into low variance segments and high variance segments by calculating the standard deviation through a sliding window. For the low variance segment, the clustering method is used to achieve efficient compression, and the clustering features are used to generate dictionary indexes, so as to control the growth of the dictionary size while capturing repeated data patterns, reducing the storage and maintenance costs of the dictionary; while for the high variance segment, frequency domain transformation, quantization processing and entropy coding are used to effectively deal with situations where the data pattern changes frequently or the random fluctuations are large, and to improve the compression effect of dynamic characteristic data.
[0014] (3) Clear segment boundary markers are used to distinguish between low-variance segments and high-variance segments, so that the decoder can accurately restore data processed in different ways. The dictionary index, frequency domain compression coding, and lossless coding of historical seed data are then packaged and combined to form an overall compressed data packet, so that the system can ensure the compression effect without losing key information when facing monitoring data, ensuring that the decoder can process historical seed data through the prediction model and reversely reconstruct the original data with the residual, thereby achieving the reversibility of the overall compression process and ensuring the integrity of the data, meeting the data recovery requirements of high-precision monitoring scenarios.
[0015] Through this technical solution, a hybrid compression algorithm is introduced at the edge device encoding end, and the historical measured monitoring seed data is used for predictive modeling, which effectively reuses the trend information of the time series data and takes the residual between the actual measured value and the predicted value as the compression object. In addition, a residual sequence standard deviation segmentation mechanism is introduced, and only the residual values of the low variance segment are clustered and a limited clustering cluster index is generated. The dictionary size is effectively controlled through lightweight dictionary generation, and frequency domain compression is used for data with drastic changes to extract frequency features, effectively capturing the hidden patterns of non-stationary data. As a result, a hybrid compression algorithm is deployed in resource-constrained edge devices, and the reliance on full analysis of real-time data is reduced through prior knowledge modeling, realizing real-time compression processing and transmission optimization on the edge side. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0017] Figure 1 The figure shows a flowchart of an example of a data compression method based on a hybrid compression algorithm according to an embodiment of the present application;
[0018] Figure 2 The figure shows an operation flowchart of an example of updating a prediction model in an edge network according to an embodiment of the present application;
[0019] Figure 3 The figure shows a schematic structural connection diagram of an example of an adaptive multi-scale attention error prediction model according to an embodiment of the present application;
[0020] Figure 4 The figure shows an operation schematic diagram of an example of encoding high-variance segment residual data according to an embodiment of the present application;
[0021] Figure 5 The figure shows a schematic structural diagram of an example of a data compression system based on a hybrid compression algorithm according to an embodiment of the present application;
[0022] Figure 6 It is a schematic structural diagram of an embodiment of an electronic device of the present application. Detailed implementation manners
[0023] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Apparently, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.
[0024] It should be noted that in the current related technologies, some experts and scholars have proposed various data compression methods to compress time-series measurement data streams, mainly including data compression methods based on coding, data compression methods based on function approximation, and data compression methods based on dictionaries.
[0025] Specifically, data compression methods based on encoding reduce redundancy through specific encoding strategies, such as common Run-length encoding (RLE), Delta encoding, etc. For example, Run-length encoding is good at compressing consecutive repeated data; Delta encoding is suitable for dealing with situations where data changes smoothly. However, these single encoding methods usually can only effectively compress specific types of data features and are difficult to adapt to multiple redundancy characteristics in data simultaneously. When the data contains both periodic trends and random fluctuations, the compression rate of such single encoding methods is usually limited. Although some research reports have proposed ways to use multiple encodings in combination (such as algorithms like Sprintz, RLBE, etc.), they mainly simply combine multiple existing encoding methods, making it difficult to fully utilize the internal laws of the data, and there are still problems with insufficient compression efficiency.
[0026] Data compression methods based on function approximation attempt to use mathematical functions to fit data segments or overall trends to achieve efficient data compression. For example, Chebyshev Polynomial Approximation (CPA) uses polynomials to approximately represent data; Discrete Fourier Transform (DFT) and Discrete Cosine Transform (DCT), etc., map data to the frequency domain and use a few frequency domain coefficients to achieve data compression. However, these methods generally have problems such as difficulty in selecting fitting functions and high computational complexity. In addition, the function approximation process for complex data often requires a large amount of trial and error and analysis, making it difficult to efficiently meet the real-time compression requirements of large-scale, multi-dimensional time-series data in practical applications.
[0027] Data compression methods based on dictionaries achieve compression by constructing data dictionaries and are also currently widely used compression methods. A typical method is differential LZW, which extracts typical repeated patterns from time-series data, constructs corresponding "atom" dictionaries, and uses these atoms for subsequent data representation to reduce redundancy. However, such dictionary compression methods have the following deficiencies: (1) The initial construction cost of the dictionary is relatively high, requiring a large amount of computing resources and storage space; (2) During long-term operation, the dictionary size is prone to unlimited expansion, causing difficulties in retrieval and maintenance; (3) For data with strong randomness or frequent pattern changes, its compression efficiency drops significantly, making it difficult to meet the requirements of application scenarios with strong real-time performance.
[0028] It should be understood that the purpose of the above description of the current related technologies is only to facilitate the public's better understanding of the inventive spirit and motivation of this application and is not regarded as a limitation of this application. In addition, the technical solutions described in the above current related technologies are not prior art and may also be unpublished technical solutions, such as those under research or in the laboratory stage.
[0029] In the technical solution of this application, for the processing of the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information involved, etc., it complies with the provisions of relevant laws and regulations and does not violate public order and good customs.
[0030] Figure 1 The flowchart of an example of the data compression method based on the hybrid compression algorithm according to an embodiment of this application is shown.
[0031] Regarding the execution subject of the method of the embodiment of this application, it can be any controller or processor with computing or processing capabilities. Specifically, it can be implemented by an edge node or a compression module in an edge node. The data compression method based on the hybrid compression algorithm can efficiently process time-series monitoring data with high sampling frequency, large data volume, and high correlation by combining a prediction model and multiple compression techniques. While ensuring data integrity, it achieves a significant compression effect and is suitable for the storage and transmission requirements of large-scale real-time monitoring data in modern data centers and cloud platforms.
[0032] In some examples, it can be integrated and configured in an electronic device or a terminal in a software, hardware, or software-hardware combination manner, and the types of terminals or electronic devices can be diverse, such as mobile phones, tablets, or desktop computers, etc.
[0033] As Figure 1 As shown, in step S110, an actual measured monitoring parameter sequence corresponding to the current sampling time period and historical actual measured monitoring seed data corresponding to the corresponding historical time window are obtained, and a first prediction model is called to determine a predicted monitoring parameter sequence, and the residual sequence between the actual measured monitoring parameter sequence and the predicted monitoring parameter sequence is calculated.
[0034] Here, the predicted monitoring parameter sequence is the parameter sequence obtained by the first prediction model predicting the monitoring parameters within the current sampling time period based on the historical actual measured monitoring seed data. It should be understood that the parameter types of the actual measured monitoring parameters can be diverse, such as CPU utilization rate, memory occupancy, disk I / O, network traffic, etc., which are not limited here for the time being.
[0035] In some implementation manners, within the current sampling time period, the device operation state data is obtained through a real-time monitoring module to form an actual measured monitoring parameter sequence; at the same time, the historical actual measured monitoring seed data within the corresponding historical time window is extracted from the device historical data storage. Preprocessing measures (such as data filtering, normalization, and outlier removal) can be adopted during the acquisition process to improve data quality.
[0036] It should be noted that during sampling, historical measured monitoring seed data and a fixed prediction model are used to generate a sequence of predicted monitoring parameters, and then the residuals are calculated by comparing with the measured data obtained from the current sampling. By using the prediction model to capture the changing trend of the data in advance, the difference between the actual collected data and the predicted data is usually small, resulting in less redundant information in the residual data itself and improving the overall compression efficiency. In some cases, both the edge device encoding end and the data center decoding end have the same historical seed data and prediction model parameters so that both sides can generate a consistent prediction sequence. In addition, the time lengths of the historical time window and the current sampling time period can be predefined in the system and can be adjusted according to system resources or business scenario requirements.
[0037] The prediction model can adopt a variety of time series prediction models, such as autoregressive models, LSTM models, or other time series prediction algorithms, which are not limited here for the time being. Specifically, the prediction model takes historical measured monitoring seed data as input and predicts the changing trend of the monitoring parameters within the current sampling time period, thereby obtaining a sequence of predicted monitoring parameters. Furthermore, a difference operation is performed on the data at the corresponding positions of the measured monitoring parameter sequence and the predicted monitoring parameter sequence to obtain a residual sequence. Thus, by using the prediction model to perform trend modeling on time series data and extracting the prediction residuals to replace the original data as the compression object, the trend redundancy information in the original data can be significantly reduced, making the residual sequence have a lower information entropy and higher compressibility.
[0038] In step S120, calculate the standard deviation of the residual sequence in each sliding window to divide the residual sequence into a low-variance segment and a high-variance segment according to a standard deviation threshold, and generate corresponding segmentation boundary flags.
[0039] In some embodiments, a sliding window with a fixed length is used to traverse the residual sequence, so that the data is divided into continuous small segments, and the length of the window can be set according to the sampling frequency and actual compression requirements. Within each sliding window, calculate the standard deviation of this segment of data to quantify the data fluctuation within the window. Specifically, according to a fixed standard deviation threshold or an adaptive standard deviation threshold, the residual sequence is divided into a low-variance segment (i.e., the segment of the residual sequence that does not exceed the standard deviation threshold, which indicates the part with less data fluctuation and higher data redundancy) and a high-variance segment (i.e., the segment of the residual sequence that exceeds the standard deviation threshold, which indicates the part with greater data fluctuation and richer information). At the same time, generate corresponding segmentation boundary flags for positioning during subsequent differential encoding and decoding for data reconstruction.
[0040] In step S130, perform clustering processing on the residual values of the low-variance segment to generate corresponding multiple clustering cluster features, and generate corresponding dictionary indexes according to each clustering cluster feature.
[0041] In some embodiments, for the residual data of the low-variance segments divided, lightweight clustering algorithms such as K-Means, Mini-Batch K-Means, or density clustering can be used to perform clustering analysis on the residual values, and several cluster centers are extracted as representative features. According to the representative features of each cluster, corresponding dictionary index values are assigned to form an index dictionary that can be mapped during decoding.
[0042] It should be noted that the low-variance segment data itself has a high degree of repeatability and similarity. Through clustering, the repeated patterns can be summarized into limited cluster centers, and multiple similar data can be replaced by a dictionary index, thereby greatly reducing the data storage volume. Based on the index generated by the clustering method, a simplified dictionary is created to represent the similar patterns in the low-variance data segment, which can effectively avoid the problem of the dictionary continuously expanding due to storing one by one in the traditional differential LZW algorithm. Instead, only a relatively small and dynamically updated dictionary is needed to manage the clustering index, thereby reducing the storage and maintenance costs of the dictionary.
[0043] In step S140, the residual values of the high-variance segments are subjected to a frequency-domain transformation to generate corresponding frequency-domain coefficient distributions, and corresponding frequency-domain compressed encodings are generated through quantization processing and entropy coding.
[0044] In some embodiments, frequency-domain transformation processing is applied to the high-variance segment residual data, such as the Fast Fourier Transform (FFT) or the Discrete Cosine Transform (DCT), to convert the time-domain data into a frequency-domain coefficient distribution. The obtained frequency-domain coefficients are quantized, that is, the original floating-point data is represented using a finite numerical range to reduce data redundancy. Furthermore, the quantized data is further compressed by using an entropy coding algorithm (such as Huffman coding or arithmetic coding) to generate frequency-domain compressed coding data.
[0045] Thus, for the high-variance segment residual data, since its changes are complex and it is difficult to capture commonalities directly using time-domain clustering, the main spectral features of the signal can be extracted through frequency-domain transformation, which is convenient for identifying its energy distribution law. By jointly applying quantization and entropy coding, while maintaining the main information of the data, the number of bits required for data storage can be greatly reduced, and the utilization rate of the transmission bandwidth is optimized.
[0046] In step S150, the historical measured monitoring seed data is losslessly encoded to generate corresponding seed data compressed encodings, and the segment boundary flags, dictionary indexes, frequency-domain compressed encodings, and seed data compressed encodings are packaged and combined to obtain the measured monitoring compressed data for the measured monitoring parameter sequence.
[0047] In some implementations, a lossless encoding algorithm (such as the LZ77 lossless compression algorithm or other lossless technologies) is used for the historical measured monitoring seed data to ensure that the seed data does not generate information loss during the encoding and subsequent decoding process. Thus, by performing lossless encoding on the seed data, it is ensured that key reference data is not lost during the compression process, thereby ensuring that the original monitoring parameters can be accurately restored during decoding.
[0048] Subsequently, the segment boundary markers, the dictionary index generated by the low variance segment, the frequency domain compression coding of the high variance segment, and the lossless coding of the seed data are packaged and combined according to the preset format to form the final measured monitoring compressed data. The packaging format may include header information, the length of each coding segment, a check code, etc., so that the decoding end can perform information recognition and data verification.
[0049] Through the embodiments of the present application, residual compression, cluster indexing and frequency domain coding are utilized to significantly reduce the data volume as a whole, thereby reducing the burden on network transmission bandwidth and central storage devices. By clustering low-variance segment data to generate dictionary indexes, the dictionary expansion problem caused by the continuous addition of new data patterns in the traditional differential LZW algorithm is avoided. By using frequency domain transformation and quantization entropy coding processing for high-frequency change segments, data details and instantaneous fluctuations can be captured more effectively, ensuring compression effects in a variety of data scenarios and enhancing adaptability to dynamic data. In addition, the lossless encoding of seed data and detailed segmentation flag information enable the original monitoring data to be accurately restored during decoding, ensuring data integrity and system stability.
[0050] Regarding the details of compressing and encoding the seed data, in some examples of the embodiments of the present application, lossless differential Huffman coding is performed on the historical measured monitoring seed data to generate corresponding seed data compression coding. Specifically, by first differentially encoding the continuous historical monitoring data, adjacent sampled data can be converted into a set of differential values, which can often significantly reduce the data variance and entropy, while ensuring that the original data is fully restored by saving the first data point, thereby achieving lossless coding. Then, when encoding the differential sequence, Huffman coding can utilize the probability of occurrence of each symbol in the data to generate the optimal prefix code and ensure accurate restoration of the data at the decoding end. The differential coding and Huffman coding algorithms have low computational complexity, are suitable for rapid implementation on resource-constrained edge devices, and can be synchronized with the decoding end in real time without the need for complex synchronization mechanisms or large amounts of memory overhead.
[0051] Regarding the packaging details of the measured monitoring compressed data, in some examples of the embodiments of the present application, the segment boundary flag, dictionary index, frequency-domain compression encoding, and seed data compression encoding are used as independent data modules. An appropriate identifier and length field are appended to each data module in the binary stream encapsulation format and encapsulated into a corresponding single overall data packet, thereby obtaining the measured monitoring compressed data for the measured monitoring parameter sequence. Specifically, after each module (segment boundary flag, dictionary index, frequency-domain compression encoding, seed data compression encoding) undergoes its respective encoding process, its respective binary stream is obtained. To package the above module data in an orderly manner, a unified data container format is designed, which can include a data packet header, a module index table, a module data body, and a check code.
[0052] Since the module index table and encoding type marker are already included in the data packet, the decoding end can directly read the metadata of the corresponding module, and then call the corresponding decoder to perform lossless decoding on the module data. Finally, the segment boundary, low-variance segment index, frequency-domain data, and seed data are reconstructed in sequence. After each part of the data is encapsulated in an independent module manner, the decoding end can directly call the corresponding decoding method according to the identifier of each module, without additional complex global synchronization, and can ensure the accuracy of data parsing.
[0053] Through the embodiments of the present application, the unified encapsulation ensures that all key data (segment flag, dictionary index, frequency-domain encoding, seed encoding) are organized in a self-describing manner, and each module information can be accurately and unambiguously identified during decoding, realizing lossless data reconstruction. In addition, each key data is encapsulated in a modular manner, and an identifier and length field are appended to each module in the binary stream encapsulation format (such as the TLV (Tag-Length-Value) structure) and packaged and combined, finally forming a self-describing overall data packet, realizing the efficient and lossless compression transmission of the measured monitoring data of the edge device.
[0054] Figure 2 The operation flowchart of an example of updating the prediction model in the edge network according to the embodiments of the present application is shown.
[0055] Specifically, a second prediction model is deployed at the decoding end of the data center, and the second prediction model has the same model parameters as the first prediction model in the edge device encoding end. To ensure that the first prediction model at the edge device end maintains a high prediction accuracy for the change trend of the monitoring data during long-term operation, the system introduces a model update mechanism of central training and edge synchronization. The model is trained and optimized at the data center end (decoding end), and a parameter update instruction is sent to the edge end to achieve remote iterative synchronization of the model.
[0056] As Figure 2 shown, the operations performed by the edge device encoding end include:
[0057] In step S210, a model parameter update request is received from the data center decoding end. The model parameter update request includes model update parameters obtained by optimizing the second prediction model through sample training.
[0058] In some embodiments, a second prediction model with the same structure as the edge device encoding end is deployed at the data center decoding end. This model uses the compressed data stream received from the edge device (including historical seed data and residual information) for reverse recovery, and further conducts centralized training and optimization based on the recovered data samples. For example, methods such as incremental training, online learning, or periodic batch training are used to adjust the model parameters. Once the second prediction model is optimized, the system constructs a model parameter update request and sends it to the edge device end through a control channel, a security channel, or a model distribution interface.
[0059] Thus, the centralization of the model training task and the edge offloading of the training load are achieved, solving the problems of limited resources and weak training capabilities at the edge end. At the same time, the consistency and controllability of the model performance are ensured through centralized optimization, providing a sustainable model update mechanism for the edge end.
[0060] In step S220, the first prediction model is updated according to the model update parameters, so that the model parameters of the first prediction model and the second prediction model can be updated synchronously.
[0061] In some embodiments, after receiving the model parameter update request, the edge device first verifies the update content included in the request, such as the model version or signature. After verification, the system replaces the parameters of the current first prediction model with the updated parameters, or performs fine-tuning of the parameters according to the differential update strategy, such as cold update or hot update mechanisms. Since the decoding end depends on the same prediction model parameters as the encoding end for data reconstruction, the model synchronization update mechanism can ensure the consistency of encoding prediction and decoding recovery, avoiding problems such as decoding failure and data misrestoration caused by model drift.
[0062] Through the embodiments of this application, the training burden of the prediction model is transferred to the data center, and the edge device only needs to perform model inference tasks, significantly reducing the occupation of local computing and storage resources and meeting the lightweight operation requirements. In addition, the second prediction model is continuously trained based on the latest monitoring data, and has stronger generalization ability and trend capture ability. By synchronously updating the first prediction model, the adaptive ability to new models is achieved through a periodic model update mechanism, and the prediction residual accuracy can be continuously optimized, thereby further reducing the residual entropy value, improving the compression ratio and the reconstruction accuracy.
[0063] In some examples of the embodiments of the present application, the prediction model may adopt an adaptive multi-scale attention error prediction model, which fully considers aspects such as the utilization of historical measured monitoring seed data in the historical time window, multi-scale extraction of data features, dynamic attention mechanism, and error feedback correction, to ensure the final output of the predicted monitoring parameter sequence.
[0064] Figure 3 FIG. shows a schematic structural connection diagram of an example of the adaptive multi-scale attention error prediction model according to the embodiments of the present application.
[0065] As Figure 3 shown, the prediction model 300 includes a cascaded preprocessing input layer 310, a multi-scale convolution feature extraction module 320, an attention mechanism module 330, an error feedback recursive fusion module 340, and a prediction output layer 350.
[0066] The preprocessing input layer 310 is used to preprocess the historical measured monitoring seed data to obtain the corresponding normalized monitoring input data.
[0067] In some embodiments, the historical measured monitoring seed data may be a time series data window corresponding to a fixed length. By normalizing the input data and filtering out noise, it is ensured that the input values for subsequent modules are stable and meet the data distribution requirements during model training.
[0068] Specifically, the input data is normalized according to a predetermined formula, such as using min-max normalization or z-score standardization, so that the data distribution is maintained within the range of [0, 1] to reduce fluctuations during numerical calculations. In addition, simple moving average filtering or wavelet denoising methods can also be used to remove noise from the input data to improve data quality.
[0069] The multi-scale convolution feature extraction module 320 is used to extract multi-scale monitoring features corresponding to the normalized monitoring input data. Here, the multi-scale convolution feature extraction module includes multiple convolutional kernels, and each convolutional kernel has a uniquely corresponding time scale.
[0070] Here, the multi-scale convolution feature extraction module 320 extracts local features at different time scales, captures short-term fluctuations, trend changes, and periodic patterns in the monitoring data, and uses multiple convolutional kernels within the module to perform convolution operations on the preprocessed data in parallel to obtain feature maps corresponding to the scales. For example, short kernels correspond to short time scales and large kernels correspond to long time scales. Finally, the feature maps output from each branch are spliced or fused to form a feature vector containing multi-scale information.
[0071] The attention mechanism module 330 adopts a lightweight attention network SE module, which is used to calculate the importance weights of the features at each scale in the multi-scale monitoring features, and perform weighted fusion on the features at each scale to obtain the corresponding attention fusion features.
[0072] Here, through the lightweight attention network SE (Squeeze and Excitation) module, the contribution weights of the features at each scale are dynamically calculated, and the concatenated multi-scale features are adaptively weighted in the channel direction, so that the information most critical for prediction is focused on.
[0073] In the compression stage, global average pooling is performed on the multi-scale feature map to generate a one-dimensional vector, which reflects the overall importance of each channel. In the excitation stage, two fully connected layers (ReLU activation can be used in the middle and Sigmoid activation is used in the last layer) are used to map this one-dimensional vector into channel weights. Furthermore, the channel weights output by the SE module are multiplied element-wise with the corresponding channels of the multi-scale features to obtain the weighted fusion features, thereby highlighting the temporal patterns that have a greater impact on the prediction results.
[0074] The error feedback recursive fusion module 340 adopts a lightweight recursive unit and is used to recursively fuse the attention fusion features with the error information generated in the previous moment's prediction, realizing dynamic error feedback correction to determine the comprehensive feature representation after dynamic correction and temporal fusion.
[0075] Here, the error feedback recursive fusion module fuses short-term and long-term temporal dependencies and introduces an additional error feedback path. Specifically, inside the recursive unit (such as GRU, LSTM or BiRNN, etc.), not only the hidden state is passed, but also the previous prediction error (i.e., the difference between the actual value and the previous prediction value) is input through an additional feedback path to achieve error compensation, and the deviation generated in the previous prediction is corrected through error feedback, thereby improving the subsequent prediction accuracy.
[0076] The prediction output layer 350 is used to map the comprehensive feature representation to the monitoring parameter space, and generate prediction values that match the number of monitoring parameters in the current sampling time period through a fully connected layer transformation, so as to obtain the corresponding predicted monitoring parameter sequence.
[0077] In some embodiments, through one or more fully connected layers, the feature dimension is gradually reduced and mapped to the prediction parameter space, and the fused high-dimensional features are converted into a predicted monitoring parameter sequence. The dimension of the finally output vector corresponds to the number of monitoring parameters in the current sampling time period, that is, a predicted monitoring parameter sequence is generated.
[0078] Through the embodiments of the present application, the edge device can utilize historical seed data and real-time collected data to achieve high-precision monitoring parameter prediction through an adaptive multi-scale attention error prediction model, and output a predicted monitoring parameter sequence. When processing the difference residual between the prediction and the actual value, since the model prediction error has been dynamically corrected, the entropy of the residual is significantly reduced, thereby providing a more ideal data basis for subsequent low-variance segment clustering and high-variance segment frequency domain compression, greatly improving the lossless compression performance and data reconstruction accuracy of the entire hybrid compression algorithm.
[0079] It should be noted that in the edge device environment, the timing and dynamic changes of monitoring data are crucial. The recursive unit undertakes the important task of capturing historical information and correcting prediction errors. By optimizing the recursive unit and making better use of the prediction error information at the previous moment for correction, the cumulative prediction error can be significantly reduced, and the overall prediction accuracy can be improved.
[0080] In some examples of the embodiments of the present application, the lightweight recursive unit adopts an error feedback gated recursive unit, which adds an error gate on the basis of the standard gated recursive unit to achieve the dynamic introduction of error feedback. By fusing the multi-scale attention fusion feature at the current moment and the error information generated in the previous moment prediction, the dynamic gating mechanism is used to selectively introduce error feedback, realize the automatic correction of the prediction state, and improve the long-term dependence modeling to ensure more accurate prediction of the monitoring parameters.
[0081] Specifically, by introducing an error gate on the basis of the traditional lightweight GRU, while maintaining the computational efficiency, the sensitivity and feedback utilization of the prediction error are enhanced, making the state update process more flexible and more robust, and realizing an error feedback gated recursive unit, which is used to perform the following operations:
[0082] Perform a linear transformation and a non-linear mapping on the error information generated in the previous moment prediction to generate an error gate, which is used to measure the contribution of the error to the current state update.
[0083] , Equation (1)
[0084] In the formula, represents the error gate vector calculated at time and is used to measure the contribution degree of the error information at the previous moment to the current state update; represents the prediction error vector at the previous moment , which is defined as the difference between the actual monitoring value and the previous prediction value; and respectively represent the weight matrix and the bias vector of the error gate, represents the Sigmoid activation function, whose output range is between 0 and 1, thus converting the error information into a regulatory factor.
[0085] In Equation (1), a special weight matrix and bias are used for linear mapping, and then a normalized error gate vector is obtained through the Sigmoid function, ensuring the scale adjustment of the error information, enabling the model to automatically determine how the error should affect the state update. By performing linear transformation and non-linear mapping on the prediction error at the previous moment, the generated error gate is numerically limited between 0 and 1, introducing a dynamic regulatory factor for the subsequent state update, so that the model can selectively introduce the error information into the state update process. Thus, the system can dynamically adjust the feedback strength according to the real-time error level, avoiding the influence of too large or too small errors on the model stability.
[0086] Introducing the error information with error gating in the calculation of the candidate hidden state enables the prediction error to dynamically correct the hidden state.
[0087] , Equation (2)
[0088] , Equation (3)
[0089] In the formula, represents the candidate hidden state vector calculated at time , represents the input feature vector corresponding to time , represents the hidden state vector at the previous time ; represents the reset gate vector at the corresponding time in the standard gated recurrent unit; represents the reset gate input weight matrix, used to map to the reset gate space; represents the reset gate hidden state weight matrix, used to map to the reset gate space; represents the reset gate bias vector; is the weight matrix for the mapping to the candidate hidden state, is the weight matrix contributed by the hidden state controlled by the reset gate, is the weight matrix for mapping the error information modulated by the error gate to the candidate hidden state space; is the bias vector of the candidate hidden state; represents element-wise multiplication, used to modulate the vectors according to the corresponding elements; is the hyperbolic tangent activation function, to ensure that the candidate state output is within .
[0090] It should be noted that in the traditional GRU, the candidate hidden state only depends on the current input and the historical state. By adding the error information gated by the error gate to the candidate hidden state formula, the dynamic introduction of the prediction error at the previous moment is realized. Specifically, in the above equations (2) and (3), the error feedback term enters the candidate state calculation after being multiplied element by element with the corresponding elements of the error gate. In this way, both the original form of the error information is retained, and it is ensured that it is introduced only when necessary, preventing unstable updates caused by excessive errors. At the same time, the candidate state uses the tanh activation function to control the numerical range, ensuring that the output is within the expected range, which is convenient for subsequent updates. Thus, the instant correction of the prediction error is realized, making the state update not only reflect the characteristics of the current input, but also consider the prediction deviation, thereby improving the prediction accuracy and alleviating the error accumulation.
[0091] The update gate is used to complete the update of the final hidden state.
[0092] , Equation (4)
[0093] , Equation (5)
[0094] In the formula, represents the moment The updated final hidden state vector; The update gate vector at the corresponding moment in the standard gated recurrent unit, which is used to control the information fusion ratio between the hidden state at the previous moment and the candidate hidden state; represents the input weight matrix of the update gate, which is used to map to the update gate space; represents the hidden state weight matrix of the update gate, which is used to map to the update gate space; represents the bias vector of the update gate.
[0095] Here, the update of the hidden state still follows the GRU standard formula, and the update gate is used to determine the weight distribution between the historical state and the candidate state. By using the update gate to smoothly fuse the new and old states, it is ensured that there is a proper transition between the candidate state and the historical state after introducing the error correction. Thus, through the standard update gate formula, the corrected candidate state and the historical state are fused in proportion, ensuring the smoothness and continuity of the model state change, and effectively maintaining the long-term dependence and the stable transmission of the information flow.
[0096] Through the embodiments of the present application, a dedicated error gating mechanism is introduced, enabling the prediction error at the previous moment to be dynamically adjusted according to the current input and historical state, thereby achieving real-time correction in candidate state generation, effectively reducing the accumulation of prediction bias, enhancing the self-adaptability of state update, and improving the overall prediction accuracy.
[0097] Regarding the standard deviation threshold in step S120, in some examples of the embodiments of the present application, an adaptive standard deviation threshold can be adopted, which is dynamically updated according to the calculation of quantiles. Specifically, based on a circular buffer, the standard deviations of a preset number of adjacent historical sliding windows are dynamically stored. Furthermore, the standard deviations stored in the circular buffer are sorted, and the standard deviation threshold is dynamically updated according to the sorting result and a preset percentile parameter.
[0098] Here, to improve the adaptability of the compression method under different monitoring data fluctuation states, the standard deviation threshold no longer adopts a static fixed value, but an adaptive threshold adjustment mechanism based on quantile calculation is introduced. Specifically, the standard deviation values corresponding to the most recent several (for example, N = 64) sliding windows are stored in a circular buffer. As the calculation results of new windows are continuously added, the buffer always stores the standard deviation values of the most recent N sliding windows, and the earlier data is gradually overwritten by the new data, ensuring that the buffer always maintains the historical state of data volatility within the most recent period of time. Define a preset percentile parameter (for example ), after sorting the stored data (for example, sorting from small to large) and then calculating the quantile, the interference of local outliers can be eliminated, the risk of misclassification caused by occasional noise or extreme values can be reduced, and the quantile of the standard deviation data saved in the historical buffer is calculated, and at this time the current threshold will be dynamically equal to "the standard deviation value of the first 75% windows in the current buffer".
[0099] Through the embodiments of the present application, since the threshold is calculated based on the real-time updated historical standard deviation distribution, it can automatically adapt to the changes in the fluctuation characteristics of the monitoring data. When the data fluctuation intensifies, the sorted quantile value will increase accordingly, thereby automatically increasing the threshold; conversely, when the data is relatively stable, the threshold will decrease. Thus, it is ensured that the classification (low variance segment or high variance segment) of each sliding window always matches the true statistical characteristics of the current data.
[0100] Figure 4 FIG. shows an operation schematic diagram of an example of encoding high variance segment residual data according to an embodiment of the present application.
[0101] As Figure 4 As shown, in step S410, the high-variance segment is divided into multiple data blocks according to a preset block size, and integer discrete cosine transform is performed on each data block to obtain corresponding frequency-domain coefficients.
[0102] Specifically, for the residual data sequence of the high-variance segment, this segment is divided into several data blocks according to the preset block size parameter. Subsequently, the integer discrete cosine transform (Integer DCT) is applied to the residual values in each data block to convert the original time-domain data into a frequency-domain coefficient representation. Thus, through the discrete cosine transform, the irregularly distributed high-frequency fluctuations in the time domain are mapped into concentrated energy coefficients in the frequency domain, effectively achieving signal energy aggregation and facilitating subsequent quantization and compression.
[0103] In step S420, reversible quantization processing is performed on each data block according to a preset quantization parameter to obtain corresponding quantized integer coefficients.
[0104] For the frequency-domain coefficients after DCT transformation, the system performs integer quantization on them according to the preset quantization parameter (such as quantization step size or hierarchical quantization table). The quantization operation divides the frequency-domain coefficients by the step size and rounds to obtain quantized integer coefficients, and records the quantization parameter for restoration (to be multiplied back for decoding). Thus, through quantization processing, the numerical dynamic range of the frequency-domain data is further compressed, preparing the data basis for entropy coding and helping to suppress invalid information. In addition, the reversible quantization processing ensures that the data can be losslessly restored to an approximate accuracy at the decoding end, taking into account both the compression ratio and the restoration accuracy.
[0105] In step S430, for each quantized integer coefficient, the occurrence times of all symbols in the quantized integer coefficient are counted to generate a corresponding histogram, and the corresponding uniformity index R is calculated.
[0106] The specific calculation formula is:
[0107] , Equation (6)
[0108] , Equation (7)
[0109] In the formula, represents the uniformity index, represents the frequency of the symbol with the most occurrences in the histogram, represents the sum of the occurrence times of all symbols in the histogram, represents the total number of symbol categories in the quantized integer coefficient histogram, represents the average occurrence times of symbols in the histogram.
[0110] In step S440, if R ≤ α, it is determined that the corresponding histogram distribution is uniform, and the corresponding histogram is losslessly compressed based on Huffman coding to generate the corresponding first histogram code, where α represents the uniformity threshold; if R > α, it is determined that the corresponding histogram distribution is sharp, and the corresponding histogram is losslessly compressed based on arithmetic coding to generate the corresponding second histogram code.
[0111] Here, the optimal coding method is dynamically selected according to the symbol distribution in the histogram. When the histogram distribution is relatively uniform, Huffman coding is selected for lossless compression, while when the histogram shows a highly sharp probability distribution, arithmetic coding is selected for lossless compression, enabling the encoder to adapt to the characteristics of different block data.
[0112] It should be noted that Huffman coding is suitable for uniform distribution cases. When the histogram distribution is relatively uniform, the probability gaps between symbols are not large, and the code lengths constructed by Huffman coding are relatively close, and it can efficiently use integer bits for coding. A shorter code length is assigned to the symbol with the highest probability to construct an optimal prefix code, whose average code length is close to the theoretical entropy value. At this time, the integer constraint of the prefix code has little impact because the probability values and code length distributions of each symbol are relatively balanced.
[0113] Arithmetic coding, on the other hand, maps the probability product of the entire message to an interval by continuously subdividing the coding interval, can achieve a coding efficiency close to the theoretical entropy, and can use non-integer bits for coding. When the histogram distribution shows sharp characteristics, the occurrence probabilities of some symbols are much higher than those of other symbols. Huffman coding is still limited by the integer bit length in this case and may not be able to fully utilize the potential compression ratio brought by the higher symbol probabilities; while arithmetic coding can reflect the symbol probabilities with finer granularity, achieve more refined bit allocation, and thus obtain a lower average code length.
[0114] Therefore, corresponding probability models are constructed for the above two discrete distributions respectively, and combined with a preset cumulative frequency table or directly construct a Huffman tree for coding, making full use of the statistical laws of discrete data to achieve adaptive entropy coding. On the basis of making full use of the data probability distribution, the coding redundancy is greatly reduced, the compression ratio is significantly improved, and at the same time, it is ensured that the original data can be perfectly restored at the decoding end.
[0115] In step S450, Huffman coding type metadata is attached to each first histogram code, and arithmetic coding type metadata is attached to each second histogram code, and the frequency domain compression code is obtained by combination.
[0116] Specifically, coding type metadata (such as a flag bit or a type field) is appended before each coding block, respectively indicating whether Huffman coding or arithmetic coding is used for that block. Then, all the encoded histogram data blocks are combined and packed to form a complete frequency-domain compression code, which is used to combine with other compressed segments in the final packing stage. Thus, the coding type metadata provides the necessary structural information for decoding. The decoding end only needs to read the coding type tag of each data block from the data packet to determine which decoding method is used for that data block, ensuring that the decoding end can correctly restore each data block.
[0117] Through the embodiments of the present application, the histogram uniformity index is used to distinguish between uniform and sharp distributions, and Huffman coding and arithmetic coding are respectively used for lossless compression. The adaptive selection of the coding method can cope with the changes in the distribution of different data blocks, significantly improving the compression ratio and the system flexibility, enabling efficient, stable, and low-latency lossless compression in real-time monitoring and data transmission of edge devices.
[0118] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a combination of a series of actions. However, those skilled in the art should know that the present application is not limited by the described action sequence, because according to the present application, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present application. In the above embodiments, each embodiment is described with emphasis. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0119] Figure 5 The structural block diagram of an example of a data compression system based on a hybrid compression algorithm according to an embodiment of the present application is shown.
[0120] As Figure 5 shown, the data compression system 500 based on the hybrid compression algorithm includes an acquisition unit 510, a residual sequence segmentation unit 520, a low-variance segment compression unit 530, a high-variance segment compression unit 540, and an encoding and packing combination unit 550.
[0121] The acquisition unit 510 is used to acquire the measured monitoring parameter sequence corresponding to the current sampling time period and the historical measured monitoring seed data corresponding to the historical time window, and call the first prediction model to determine the predicted monitoring parameter sequence, and calculate the residual sequence between the measured monitoring parameter sequence and the predicted monitoring parameter sequence; the predicted monitoring parameter sequence is the parameter sequence obtained by the first prediction model predicting the monitoring parameters within the current sampling time period according to the historical measured monitoring seed data.
[0122] The residual sequence segmentation unit 520 is used to calculate the standard deviation of the residual sequence in each sliding window, divide the residual sequence into a low-variance segment and a high-variance segment according to a standard deviation threshold, and generate corresponding segmentation boundary flags.
[0123] The low-variance segment compression unit 530 is used to perform clustering processing on the residual values of the low-variance segment to generate corresponding multiple cluster feature clusters, and generate corresponding dictionary indexes according to each of the cluster feature clusters.
[0124] The high-variance segment compression unit 540 is used to perform frequency-domain transformation on the residual values of the high-variance segment to generate a corresponding frequency-domain coefficient distribution, and generate a corresponding frequency-domain compression code through quantization processing and entropy coding.
[0125] The encoding and packing combination unit 550 is used to perform lossless encoding on the historical measured monitoring seed data to generate a corresponding seed data compression code, and pack and combine the segmentation boundary flag, the dictionary index, the frequency-domain compression code, and the seed data compression code, so as to obtain the measured monitoring compression data for the measured monitoring parameter sequence.
[0126] In some embodiments, the embodiment of the present application provides a non-volatile computer-readable storage medium, in which one or more programs including execution instructions are stored, and the execution instructions can be read and executed by an electronic device (including but not limited to a computer, a server, or a network device, etc.) to be used to execute the steps of any one of the above data compression methods based on a hybrid compression algorithm of the present application.
[0127] In some embodiments, the embodiment of the present application further provides a computer program product, the computer program product includes a computer program stored on a non-volatile computer-readable storage medium, the computer program includes program instructions, and when the program instructions are executed by a computer, the computer is enabled to execute the steps of any one of the above data compression methods based on a hybrid compression algorithm.
[0128] In some embodiments, the embodiment of the present application further provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, so that the at least one processor can execute the steps of the data compression method based on a hybrid compression algorithm.
[0129] Figure 6 It is a schematic hardware structure diagram of an electronic device for executing the data compression method based on a hybrid compression algorithm provided by another embodiment of the present application. As Figure 6 shown, the device includes:
[0130] One or more processors 610 and a memory 620, Figure 6 Taking one processor 610 as an example.
[0131] The device for executing the data compression method based on the hybrid compression algorithm may further include: an input device 630 and an output device 640.
[0132] The processor 610, the memory 620, the input device 630, and the output device 640 may be connected through a bus or other means, Figure 6 Taking connection through a bus as an example.
[0133] The memory 620, as a non-volatile computer-readable storage medium, can be used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the data compression method based on the hybrid compression algorithm in the embodiments of the present application. The processor 610 executes various functional applications and data processing of the server by running the non-volatile software programs, instructions, and modules stored in the memory 620, that is, to implement the data compression method based on the hybrid compression algorithm in the above method embodiments.
[0134] The memory 620 may include a program storage area and a data storage area. Among them, the program storage area can store an operating system and application programs required for at least one function; the data storage area can store data created according to the use of the electronic device, etc. In addition, the memory 620 may include a high-speed random access memory, and may also include non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other non-volatile solid-state storage devices. In some embodiments, the memory 620 may optionally include a memory remotely set relative to the processor 610, and these remote memories can be connected to the electronic device through a network. Examples of the above network include but are not limited to the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof.
[0135] The input device 630 can receive input digital or character information, and generate signals related to the user settings and function control of the electronic device. The output device 640 may include a display device such as a display screen.
[0136] The one or more modules are stored in the memory 620, and when executed by the one or more processors 610, execute the data compression method based on the hybrid compression algorithm in any of the above method embodiments.
[0137] The above product can execute the method provided by the embodiments of the present application, and has the corresponding functional modules and beneficial effects for executing the method. For technical details not described in detail in this embodiment, reference can be made to the method provided by the embodiments of the present application.
[0138] The electronic devices in the embodiments of this application exist in various forms, including but not limited to:
[0139] (1) Mobile communication devices: These devices are characterized by having mobile communication functions and mainly aim to provide voice and data communication. Such terminals include: smart phones, multimedia phones, functional phones, and low-end phones, etc.
[0140] (2) Ultra-mobile personal computer devices: These devices belong to the category of personal computers, have computing and processing functions, and generally also have the characteristic of mobile Internet access. Such terminals include: PDA, MID, and UMPC devices, etc.
[0141] (3) Portable entertainment devices: These devices can display and play multimedia content. Such devices include: audio and video players, handheld game consoles, e-books, and smart toys and portable in-vehicle navigation devices.
[0142] (4) Other on-board electronic devices with data interaction functions, such as in-vehicle device installed on a vehicle.
[0143] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0144] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a general hardware platform, and of course, it can also be implemented by hardware. Based on such an understanding, the essence of the above technical solutions or the part that contributes to the related technologies can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to enable a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.
[0145] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, rather than to limit it; although this application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of this application.
Claims
1. A data compression method based on a hybrid compression algorithm, applied to the encoding end of an edge device. The method includes: Obtain the measured monitoring parameter sequence corresponding to the current sampling time period and the historical measured monitoring seed data corresponding to the historical time window, and call the first prediction model to determine the predicted monitoring parameter sequence. Calculate the residual sequence between the measured monitoring parameter sequence and the predicted monitoring parameter sequence; the predicted monitoring parameter sequence is the parameter sequence obtained by the first prediction model predicting the monitoring parameters within the current sampling time period based on the historical measured monitoring seed data; Calculate the standard deviation of the residual sequence in each sliding window, and divide the residual sequence into a low-variance segment and a high-variance segment according to the standard deviation threshold, and generate corresponding segmentation boundary flags; Perform clustering processing on the residual values of the low-variance segment to generate corresponding multiple clustering cluster features, and generate corresponding dictionary indexes according to each of the clustering cluster features; Perform frequency domain transformation on the residual values of the high-variance segment to generate a corresponding frequency domain coefficient distribution, and generate a corresponding frequency domain compression code through quantization processing and entropy coding; Perform lossless coding on the historical measured monitoring seed data to generate a corresponding seed data compression code, and package and combine the segmentation boundary flag, the dictionary index, the frequency domain compression code, and the seed data compression code to obtain the measured monitoring compression data for the measured monitoring parameter sequence.
2. The method according to claim 1, wherein, A second prediction model is deployed at the decoding end of the data center, and the second prediction model has the same model parameters as the first prediction model. The method further includes: Receive a model parameter update request from the decoding end of the data center; the model parameter update request includes model update parameters obtained by optimizing the second prediction model through sample training; Update the first prediction model according to the model update parameters, so that the model parameters of the first prediction model and the second prediction model can be updated synchronously.
3. The method according to claim 1 or 2, wherein The first prediction model adopts an adaptive multi-scale attention error prediction model, which includes a cascaded preprocessing input layer, a multi-scale convolution feature extraction module, an attention mechanism module, an error feedback recursive fusion module, and a prediction output layer; The preprocessing input layer is used to preprocess the historical measured monitoring seed data to obtain corresponding standardized monitoring input data; The multi-scale convolution feature extraction module is used to extract multi-scale monitoring features corresponding to the standardized monitoring input data; among them, the multi-scale convolution feature extraction module includes multiple convolutional kernels, and each convolutional kernel has a uniquely corresponding time scale; The attention mechanism module is used to calculate the importance weights of the features of each scale in the multi-scale monitoring features, and perform weighted fusion on the features of each scale to obtain corresponding attention fusion features; The error feedback recursive fusion module is used to recursively fuse the attention fusion features with the error information generated in the previous moment prediction, realize dynamic error feedback correction, and determine the comprehensive feature representation after dynamic correction and temporal fusion; The prediction output layer is used to map the comprehensive feature representation to the monitoring parameter space, and generate prediction values that match the number of monitoring parameters in the current sampling time period through a fully connected layer, so as to obtain the corresponding predicted monitoring parameter sequence.
4. The method according to claim 3, wherein The error feedback recursive fusion module uses an error feedback gated recurrent unit, which adds an error gate on the basis of a standard gated recurrent unit to realize the dynamic introduction of error feedback; The error feedback gated recurrent unit is used to perform the following operations: Perform linear transformation and non-linear mapping on the error information generated in the previous moment prediction to generate an error gate, which is used to measure the contribution of the error to the current state update: , In the formula, represents the error gate vector calculated at time and is used to measure the contribution degree of the error information at the previous time to the current state update; represents the predicted error vector at the previous time , which is defined as the difference between the actual monitored value and the previous predicted value; represents the weight matrix of the error gate, represents the bias vector of the error gate, represents the Sigmoid activation function, whose output range is between 0 and 1, thus converting the error information into a regulation factor; Introduce the error information after the error gate in the candidate hidden state calculation, so that the prediction error can dynamically correct the hidden state: , , In the formula, represents the candidate hidden state vector calculated at time ; represents the input feature vector corresponding to time ; represents the hidden state vector at the previous time ; represents the reset gate vector at the corresponding time in the standard gated recurrent unit ; represents the reset gate input weight matrix, which is used to map to the reset gate space; represents the reset gate hidden state weight matrix, which is used to map to the reset gate space; represents the reset gate bias vector; is the weight matrix for the mapping from to the candidate hidden state, is the weight matrix contributing to the hidden state controlled by the reset gate, is the weight matrix for mapping the error information modulated by the error gate to the candidate hidden state space; is the bias vector of the candidate hidden state; represents element-wise multiplication, which is used to modulate the vector by corresponding elements; is the hyperbolic tangent activation function to ensure that the candidate state output is within ; Use the update gate to complete the update of the final hidden state: , , In the formula, represents the time The final hidden state vector after update; At the corresponding time in the standard gated recurrent unit The update gate vector, which is used to control the information fusion ratio between the hidden state of the previous time and the candidate hidden state; Represents the update gate input weight matrix, which is used to Map to the update gate space; Represents the update gate hidden state weight matrix, which is used to Map to the update gate space; Represents the update gate bias vector.
5. The method according to claim 1, wherein The standard deviation threshold is dynamically updated according to the quantile calculation, including: Based on a circular buffer, dynamically store the standard deviations of a preset number of adjacent historical sliding windows; Sort the standard deviations stored in the circular buffer, and dynamically update the standard deviation threshold according to the sorting result and a preset percentile parameter.
6. The method according to claim 1, wherein, The process of performing frequency domain transformation on the residual values of the high variance segment to generate the corresponding frequency domain coefficient distribution, and generating the corresponding frequency domain compression coding through quantization processing and entropy coding, includes: Divide the high variance segment into multiple data blocks according to a preset block size, and perform integer discrete cosine transformation on each data block to obtain the corresponding frequency domain coefficients; Perform reversible quantization processing on each of the data blocks according to a preset quantization parameter to obtain the corresponding quantized integer coefficients; For each of the quantized integer coefficients, count the number of occurrences of all symbols in the quantized integer coefficients to generate the corresponding histogram, and calculate the corresponding uniformity index: , , In the formula, represents the uniformity index, represents the frequency of the symbol that appears most frequently in the histogram, represents the sum of the occurrences of all symbols in the histogram, represents the total number of symbol categories in the quantized integer coefficient histogram, represents the average number of occurrences of symbols in the histogram; If , it is determined that the corresponding histogram distribution is uniform, and the corresponding histogram is losslessly compressed based on Huffman coding to generate the corresponding first histogram code, represents the uniformity threshold; If , it is determined that the corresponding histogram distribution is sharp, and the corresponding histogram is losslessly compressed based on arithmetic coding to generate a corresponding second histogram code; Attach Huffman coding type metadata to each of the first histograms and attach arithmetic coding type metadata to each of the second histograms, and combine them to obtain the frequency domain compression coding.
7. The method according to claim 1, wherein The process of performing lossless coding on the historical measured monitoring seed data to generate the corresponding seed data compression coding includes: Perform lossless differential Huffman coding on the historical measured monitoring seed data to generate the corresponding seed data compression coding.
8. The method according to claim 1, wherein Pack and combine the segment boundary flag, the dictionary index, the frequency domain compression coding, and the seed data compression coding to obtain the measured monitoring compressed data for the measured monitoring parameter sequence, including: Take the segment boundary flag, the dictionary index, the frequency domain compression coding, and the seed data compression coding as independent data modules, and use a binary stream encapsulation format to attach the corresponding identification and length fields to each data module and encapsulate them into a corresponding single overall data packet, so as to obtain the measured monitoring compressed data for the measured monitoring parameter sequence.
9. A data compression system based on a hybrid compression algorithm, deployed at the edge device encoding end. The system includes: An acquisition unit is configured to acquire an actual measured monitoring parameter sequence corresponding to a current sampling time period and historical actual measured monitoring seed data corresponding to a historical time window, and call a first prediction model to determine a predicted monitoring parameter sequence, and calculate a residual sequence between the actual measured monitoring parameter sequence and the predicted monitoring parameter sequence; the predicted monitoring parameter sequence is a parameter sequence obtained by the first prediction model predicting the monitoring parameters within the current sampling time period based on the historical actual measured monitoring seed data; A residual sequence segmentation unit is configured to calculate the standard deviation of the residual sequence in each sliding window, divide the residual sequence into a low-variance segment and a high-variance segment according to a standard deviation threshold, and generate corresponding segmentation boundary flags; A low-variance segment compression unit is configured to perform clustering processing on the residual values of the low-variance segment to generate corresponding multiple cluster feature sets, and generate corresponding dictionary indexes respectively according to each of the cluster feature sets; A high-variance segment compression unit is configured to perform frequency-domain transformation on the residual values of the high-variance segment to generate a corresponding frequency-domain coefficient distribution, and generate a corresponding frequency-domain compression code through quantization processing and entropy coding; An encoding and packaging combination unit is configured to perform lossless encoding on the historical actual measured monitoring seed data to generate a corresponding seed data compression code, and package and combine the segmentation boundary flags, the dictionary indexes, the frequency-domain compression code and the seed data compression code, so as to obtain actual measured monitoring compressed data for the actual measured monitoring parameter sequence.
Citation Information
Patent Citations
Data processing method and device, electronic equipment and nonvolatile storage medium
CN118520246A
Time sequence data compression method and device, computer equipment, computer readable storage medium and computer program product
CN118740165A