Dynamic Data Compression and Transmission System and Method Based on Big Data

By introducing hierarchical matching DEW-LZ4 technology based on dynamic entropy weights and machine learning Q-Learning model to dynamically adjust the transmission parameters, the traditional data compression algorithm is solved inefficient in high-entropy data processing and resource waste in the transmission process, and efficient and stable data compression and transmission are achieved.

CN119854377BActive Publication Date: 2025-06-24SHANDONG ZAIQI DATA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510336305.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-06-24
Estimated Expiration
2045-03-21

AI Technical Summary

Technical Problem

In the context of big data, traditional data compression algorithms are inefficient when processing high-entropy data, and bandwidth resources are wasted and delays are increased during transmission, making it difficult to ensure the fast and reliable transmission of data.

Method used

The LZ4 compression algorithm is optimized based on dynamic entropy weights, and the transmission parameters are dynamically adjusted in combination with the machine learning Q-Learning model to ensure efficient transmission under different network conditions.

Benefits of technology

It improves data compression efficiency, adapts to the processing capacity of different types and complex data, reduces the peak CPU occupancy rate, improves the stability and performance of the system, and maintains efficient data transmission in different network environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119854377B_ABST
    Figure CN119854377B_ABST
Patent Text Reader

Abstract

The present invention relates to the technical field of data compression and transmission, and specifically, to a dynamic data compression and transmission system and method based on big data. It includes: a data acquisition and processing unit that obtains and processes raw data from multiple data sources; a data compression execution unit that compresses the raw data through the LZ4 compression algorithm, and introduces the hierarchical matching DEW-LZ4 technology based on dynamic entropy weight in the LZ4 compression algorithm to optimize the compression process; a data transmission unit that transmits the compressed data over the network through the transmission protocol HTTPS, and dynamically adjusts the transmission parameters according to the network state using the machine learning Q-Learning model; a receiving end decompression unit that obtains the transmitted compressed data from the network and decompresses it using the LZ4 decompression algorithm to restore it to the original data format. Through the introduction of the hierarchical matching DEW-LZ4 technology based on dynamic entropy weight, the system of the present invention can intelligently select the optimal compression strategy according to different characteristics (Shannon entropy value) of the data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data compression and transmission. Specifically, it relates to a dynamic data compression and transmission system and method based on big data. Background Art

[0002] With the development of technologies such as the Internet of Things, social media, and mobile devices, the speed and scale of data generation have increased sharply. How to efficiently store, process, and transmit these massive amounts of data has become an important challenge. The data is not only large in quantity but also has a wide range of sources and various formats, including various types such as sensor data, log files, database records, and streaming data. Traditional data processing methods are difficult to effectively handle this multi-source heterogeneous data environment. In the context of big data, transmitting a large amount of data from one place to another requires a large amount of bandwidth resources and may lead to an increase in latency and packet loss rate. Especially in the case of unstable network conditions or limited bandwidth, how to ensure that data can be transmitted quickly and reliably is an urgent problem to be solved. To reduce the amount of data transmitted, data is usually compressed. However, traditional compression algorithms may perform poorly on certain types of high-entropy data (such as log data with strong randomness), resulting in low compression efficiency or a long time-consuming compression / decompression process. Therefore, a dynamic data compression and transmission system and method based on big data are provided. Summary of the Invention

[0003] The purpose of the present invention is to provide a dynamic data compression and transmission system and method based on big data to solve the problems of explosive growth of data volume, difficult integration of heterogeneous data sources, high network transmission efficiency and cost, and difficulty in achieving a balance between data compression and decompression speed and effect proposed in the above background art.

[0004] To achieve the above object, on the one hand, the present invention aims to provide a dynamic data compression and transmission system based on big data, including:

[0005] A data acquisition and processing unit, which acquires and processes raw data from multiple data sources;

[0006] A data compression execution unit, which compresses the raw data through the LZ4 compression algorithm and introduces the hierarchical matching DEW-LZ4 technology based on dynamic entropy weight in the LZ4 compression algorithm to optimize the compression process;

[0007] A data transmission unit, which transmits the compressed data through the transmission protocol HTTPS for network transmission and dynamically adjusts the transmission parameters according to the network state by using the machine learning Q-Learning model;

[0008] The receiving - end decompression unit obtains the transmitted compressed data from the network and uses the LZ4 decompression algorithm to decompress it and restore it to the original data format.

[0009] As a further improvement of this technical solution, the data acquisition and processing unit includes a data acquisition module and a data processing module;

[0010] Among them, the data acquisition module supports real - time acquisition of multi - source heterogeneous data and performs preliminary classification and priority marking on the original multi - source heterogeneous data;

[0011] The data processing module analyzes the data characteristics based on the preliminary classification and marked data, generates metadata tags, and divides the data into data blocks of a fixed size m.

[0012] As a further improvement of this technical solution, the data compression execution unit compresses the original data through the LZ4 compression algorithm and introduces the hierarchical matching based on dynamic entropy weight DEW - LZ4 technology in the LZ4 compression algorithm to optimize the compression process, including the following steps:

[0013] S1.1. Use the hierarchical matching based on dynamic entropy weight DEW - LZ4 technology to find the repeatedly occurring data sequences in the data block;

[0014] S1.2. For the matched part, LZ4 uses pointers to represent the position and length information of the matched sequence relative to the current position;

[0015] S1.3. If the length of the found matched sequence is greater than the threshold n, replace the original data with a dictionary pointer; for the part where no match is found, copy the original data;

[0016] S1.4. Concatenate the compressed data block and the matching information through LZ4 to form the compressed output data.

[0017] As a further improvement of this technical solution, in S1.1, using the hierarchical matching based on dynamic entropy weight DEW - LZ4 technology to find the repeatedly occurring data sequences in the data block includes the following steps:

[0018] S1.11. Calculate the Shannon entropy of the current data block , and based on the metadata tags, establish an entropy - value - matching weight model;

[0019] S1.12. Calculate the entropy weight of the data block according to the entropy - value - matching weight model ;

[0020] S1.13. If the entropy weight is greater than the preset threshold A, use a convolutional neural network to predict the matching hotspots; if the entropy weight If it is greater than the preset threshold B and less than or equal to the preset threshold A, a three-level hash index is constructed for matching; if the entropy weight is less than or equal to the preset threshold B, the traditional LZ4 algorithm is continued to be used for matching;

[0021] S1.14. In all cases, when the change rate of the Shannon entropy value exceeds the set threshold , the update mechanism of the double-buffer dictionary structure is triggered.

[0022] As a further improvement of this technical solution, in S1.11, the entropy value-matching weight model is:

[0023] ;

[0024] wherein, represents the adjustment factor; represents the reference Shannon entropy value.

[0025] As a further improvement of this technical solution, the data transmission unit transmits the compressed data over the network through the transmission protocol HTTPS and dynamically adjusts the transmission parameters according to the network status, including the following steps:

[0026] S2.1. Receive the already compressed data block from the data compression execution unit;

[0027] S2.2. Establish a secure HTTPS connection according to the information of the target server;

[0028] S2.3. Monitor the current network condition in real time and execute the adaptive strategy to adjust the data shard size and compression mode according to the real-time network condition;

[0029] S2.4. Control the sending rate by using the sliding window mechanism, and through ACK confirmation and timeout retransmission, when there is network fluctuation, use selective acknowledgment to only retransmit the lost data segment;

[0030] S2.5. Use the built-in security features of the TLS protocol to encrypt and perform integrity verification on the transmitted data;

[0031] S2.6. Train the machine learning Q-Learning model based on historical data to predict the change trend of the network status and adjust the transmission parameters in advance.

[0032] As a further improvement of this technical solution, in S2.6, training the machine learning Q-Learning model based on historical data to predict the change trend of the network status and adjust the transmission parameters in advance includes the following steps:

[0033] S2.61. Collect the historical data of the network status and record the associated transmission parameters;

[0034] S2.62. Discretize the network state into a multi-dimensional vector;

[0035] S2.63. Define the state of the network and define adjustable transmission parameters as actions;

[0036] S2.64. Design a reward function to evaluate the quality of each action;

[0037] S2.65. Create a state-action matrix, assign initial values to unknown state-action pairs, and initialize the Q value;

[0038] S2.66. Adopt the ε-greedy strategy to balance the relationship between exploring new strategies and exploiting existing knowledge, and update the Q value according to the Q-value update formula. Introduce an algorithm for dynamically adjusting the learning rate based on entropy weight in the Q-value update formula for optimization, so that the learning rate adapts to the entropy weight of the system changes, and for the problems of balancing compression efficiency and computing resources and network environment changes, add an additional reward term to the Q-value update formula for further optimization;

[0039] S2.67. Dynamically adjust the transmission parameters according to the action suggestions output by the machine learning Q-Learning model.

[0040] As a further improvement of this technical solution, in S2.66, the Q-value update formula is:

[0041] ;

[0042] Where, represents the current state; represents the next state; represents the action taken; represents the immediate reward after taking the action; represents the learning rate; represents the discount factor; represents the state under which the action is taken and the expected cumulative reward that can be obtained; represents the time step;

[0043] Introduce an algorithm for dynamically adjusting the learning rate based on entropy weight in the Q-value update formula for optimization, so that the learning rate adapts to the entropy weight of the system changes:

[0044] Where, the algorithm for dynamically adjusting the learning rate based on entropy weight is:

[0045] ;

[0046] Where, represents the learning rate dynamically adjusted based on entropy weight; represents the initial learning rate; represents the average value of the entropy weight; represents the standard deviation measuring the change of the entropy weight;

[0047] Introducing the algorithm for dynamically adjusting the learning rate based on entropy weight into the Q-value update formula is as follows:

[0048] ;

[0049] Aiming at the problems of balancing compression efficiency and computing resources and network environment changes, an additional reward term is added to the Q-value update formula for further optimization:

[0050] Among them, the additional reward term includes the compression efficiency reward and the entropy stability reward :

[0051] ;

[0052] Among them, represents the adjustment parameter;

[0053] Considering that a high compression rate leads to a high CPU occupancy rate, a dynamic cooperation factor is introduced into the compression efficiency reward to quantify the conflict degree between the compression rate and the CPU occupancy rate, and the compression efficiency reward is adjusted through a second-order interaction term:

[0054] ;

[0055] Among them, represents the weight coefficient of the compression rate; represents the weight coefficient of the CPU occupancy rate; represents the weight coefficient of the dictionary hit rate; represents the cooperation strength coefficient; represents the standard deviation of bandwidth fluctuation.

[0056] As a further improvement of this technical solution, the receiving end decompression unit includes a decompression and verification module and a data recombination module;

[0057] Among them, the decompression and verification module decompresses according to the metadata by invoking the LZ4 decompression algorithm and verifies the data integrity;

[0058] The data recombination module combines the block data and restores the original structure.

[0059] On the other hand, the present invention provides a dynamic data compression and transmission method based on big data, which is used for the dynamic data compression and transmission system based on big data described in any one of the above, and includes the following steps:

[0060] S3.1. Obtain and process raw data from multiple data sources;

[0061] S3.2. Compress the raw data through the LZ4 compression algorithm, and introduce the hierarchical matching DEW-LZ4 technology based on dynamic entropy weight into the LZ4 compression algorithm to optimize the compression process;

[0062] S3.3. Transmit the compressed data over the network through the HTTPS transport protocol, and dynamically adjust the transmission parameters according to the network status using the machine learning Q-Learning model;

[0063] S3.4. Obtain the transmitted compressed data from the network and decompress it using the LZ4 decompression algorithm to restore it to the original data format.

[0064] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0065] 1. In the dynamic data compression and transmission system and method based on big data, by introducing the hierarchical matching DEW-LZ4 technology based on dynamic entropy weight, the system can intelligently select the optimal compression strategy according to different characteristics of the data (such as Shannon entropy value). For high-entropy data, a convolutional neural network is used to predict the hot spots to accelerate compression; for medium-entropy data, a three-level hash index is used to reduce invalid searches and improve the compression ratio; for low-entropy data, the traditional LZ4 algorithm is still used for efficient compression. This method not only improves the overall compression efficiency but also enhances the adaptability to different types and complexities of data, especially suitable for processing large-scale and heterogeneous data sets. In addition, the application of the double-buffer dictionary structure further optimizes the memory usage, reduces the peak CPU occupancy rate, and improves the stability and performance of the system.

[0066] 2. In the dynamic data compression and transmission system and method based on big data, the HTTPS protocol is used to ensure security during data transmission, and the machine learning Q-Learning model is used to dynamically adjust the transmission parameters according to the real-time network status, such as the data shard size, compression mode, etc., so as to maintain efficient data transmission under different network conditions. For example, when the bandwidth is high, the data shard size is increased to improve the transmission efficiency; when the bandwidth is low or the latency is high, the shard size is reduced to ensure the stability of the transmission. At the same time, the sliding window mechanism is used to control the sending rate, combined with the ACK confirmation and timeout retransmission mechanisms to ensure the reliability of data transmission. This adaptive strategy greatly reduces the latency and packet loss rate, providing a more stable quality of service for users, especially suitable for scenarios with large network fluctuations. Description of the Drawings

[0067] Figure 1 It is the overall flow block diagram of the present invention;

[0068] The meanings of each label in the figure are as follows:

[0069] 1. Data acquisition and processing unit; 11. Data acquisition module; 12. Data processing module; 2. Data compression execution unit; 3. Data transmission unit; 4. Receiver decompression unit; 41. Decompression and verification module; 42. Data recombination module. Specific implementation manners

[0070] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0071] Embodiment 1: Please refer to Figure 1 As shown, a dynamic data compression and transmission system based on big data is provided, including.

[0072] The data acquisition and processing unit 1 acquires raw data from multiple data sources (data sources include sensors, devices, user terminals, etc.) and processes it;

[0073] In this embodiment, the data acquisition and processing unit 1 includes a data acquisition module 11 and a data processing module 12;

[0074] Among them, the data acquisition module 11 supports real-time acquisition of multi-source heterogeneous data (sensors, logs, databases, streaming data), and performs preliminary classification and priority marking (including data volume, timeliness, sensitivity) on the original multi-source heterogeneous data;

[0075] The data processing module 12 analyzes data characteristics (including entropy value, repetition rate, temporal correlation) based on the preliminarily classified and marked data, generates metadata tags, and divides the data into blocks of a fixed size m for subsequent compression.

[0076] The data compression execution unit 2 compresses the raw data through the LZ4 compression algorithm, and introduces the hierarchical matching DEW-LZ4 technology based on dynamic entropy weight in the LZ4 compression algorithm to optimize the compression process;

[0077] In this embodiment, traditional LZ4 relies on a fixed hash table and a sliding window to detect duplicate sequences, while DEW-LZ4 dynamically adjusts the matching strategy through entropy weights, solving the problem of low matching efficiency of traditional methods in heterogeneous data (such as highly random logs); for high-entropy data (such as sensor streams), CNN prediction of hot spots can improve the compression speed; for medium-entropy data (such as database records), a three-level hash index reduces invalid searches and improves the compression ratio; the dynamic entropy weight model can identify different characteristics of text, binary, and time-series data and automatically switch compression strategies. For example, for highly random log data (high entropy), the optimized compression ratio is higher than that of traditional LZ4; the double-buffer dictionary reduces memory occupancy and avoids the overhead of frequently rebuilding the hash table; the dictionary update is triggered by the entropy change rate, reducing the peak CPU occupancy rate; based on the extreme compression of traditional LZ4, DEW-LZ4 breaks through the dependence of fixed algorithms on data characteristics through information entropy theory + machine learning, matches appropriate algorithms for data with different complexities, and makes an intelligent trade-off between resource consumption and compression efficiency;

[0078] The data compression execution unit 2 compresses the original data through the LZ4 compression algorithm and introduces the hierarchical matching DEW-LZ4 technology based on dynamic entropy weights in the LZ4 compression algorithm to optimize the compression process, including the following steps:

[0079] S1.1. Use the hierarchical matching DEW-LZ4 technology based on dynamic entropy weights to find duplicate data sequences in the data block. The repeated sequences are recorded as "dictionary" entries. The hierarchical matching DEW-LZ4 based on dynamic entropy weights is a technology designed to optimize the performance of the LZ4 compression algorithm. It improves data compression efficiency, especially when dealing with large-scale, heterogeneous data sets, by introducing a dynamic entropy-aware mechanism and a hierarchical matching architecture;

[0080] Among them, using the hierarchical matching DEW-LZ4 technology based on dynamic entropy weights to find duplicate data sequences in the data block includes the following steps:

[0081] S1.11. Calculate the Shannon entropy of the current data block (The Shannon entropy is the probability-weighted average of the information amounts of each event in the discrete event set), and based on the metadata label, establish an entropy value - matching weight model. The entropy value - matching weight model is a mechanism that calculates the entropy weight according to the Shannon entropy of the data block and then guides the selection of the matching strategy in the data compression process;

[0082] The entropy value - matching weight model is:

[0083] ;

[0084] Among them, represents the adjustment factor (recommended value 0.5); represents the reference Shannon entropy value (obtained through historical data analysis);

[0085] S1.12. Calculate the entropy weight of the data block according to the entropy value - matching weight model ;

[0086] S1.13. If the entropy weight is greater than the preset threshold A (the preset threshold A in this embodiment is 0.7), then use a convolutional neural network to predict the matching hotspots; if the entropy weight is greater than the preset threshold B (the preset threshold B in this embodiment is 0.3) and less than or equal to the preset threshold A, then construct a three - level hash index for matching; if the entropy weight is less than or equal to the preset threshold B, then continue to use the traditional LZ4 algorithm for matching (the traditional LZ4 algorithm uses the sliding window technique to find and replace repeated sequences in the data with short references for compression purposes);

[0087] In this embodiment, the steps of using a convolutional neural network to predict the matching hotspots are as follows: The input of the CNN is the entropy weight features of the current data block (such as Shannon entropy E, repetition rate in metadata tags, temporal correlation, etc.). Local pattern features are extracted through the convolutional layer. The model architecture includes multiple convolutional layers + pooling layers. The output is the probability distribution map of potential matching hotspots in the data block. High - entropy regions (such as log data with strong randomness) may identify periodic or local repeated patterns through the CNN. When the entropy weight is greater than the preset threshold A (0.7), the data block has high complexity and the traditional hash matching efficiency is low. The CNN guides the hash table to match preferentially in these regions by predicting the hotspots (such as high - frequency repeated byte sequences), reducing ineffective searches;

[0088] The steps of constructing a three - level hash index for matching are as follows: Design a three - level hash structure: The first - level hash: Use the hash value of the 4 - byte sequence in the data block as the key and map it to the initial bucket (similar to the hash table of LZ4); The second - level hash: Construct a chained hash for the conflicting key values and disperse the conflicts through a secondary hash function; The third - level hash: On the basis of the first two levels, construct a prefix tree (Trie) index for long matching sequences (> 8 bytes) to accelerate long - sequence matching. When the entropy weight is between the preset threshold B (0.3) and A (0.7), the data block has a medium repetition rate. The three - level hash index gradually narrows the search range layer by layer: The first - level hash quickly locates the candidate positions; The second - level hash resolves conflicts; The third - level hash confirms the longest repeated sequence through prefix matching, significantly improving the matching accuracy. Compared with the single - layer hash of the traditional LZ4, the three - level index reduces the hash conflict rate by about 60%, especially suitable for heterogeneous data (such as mixed text and binary streams);

[0089] S1.14. In all cases, when the change rate of the Shannon entropy value exceeds the set threshold ( When the entropy value change rate exceeds the standard ( ), or when the matching hit rate continuously decreases, the update mechanism of the dual - cache dictionary structure is triggered. The specific composition of the dual - cache dictionary structure includes the cooperation mechanism between the main dictionary and the backup dictionary. The triggering conditions of the update mechanism are: the entropy value change rate exceeds the standard (

[0090]

[0091]

[0092]

[0093] The data transmission unit 3 transmits the compressed data over the network through the transmission protocol HTTPS and dynamically adjusts the transmission parameters according to the network status using the machine learning Q - Learning model;

[0094] In this embodiment, the HTTPS protocol provides end - to - end encryption, ensuring the security and integrity of the data during transmission. This is particularly important for the transmission of sensitive information and can effectively prevent data from being eavesdropped or tampered with; by compressing the data before transmission, the amount of data actually transmitted over the network is reduced, which can not only speed up the transmission speed but also reduce the bandwidth usage cost; the system can monitor the current network conditions (such as bandwidth, latency, packet loss rate, etc.) in real - time and dynamically adjust the transmission parameters (such as data shard size and compression mode) based on this information. This adaptive strategy enables effective data transmission even under poor network conditions while optimizing resource consumption;

[0095] The data transmission unit 3 transmits the compressed data over the network through the transmission protocol HTTPS and dynamically adjusts the transmission parameters according to the network status, including the following steps:

[0096] S2.1. Receive the already compressed data block from the data compression execution unit 2;

[0097] S2.2. Establish a secure HTTPS connection according to the information of the target server (such as URL, port, etc.), which includes the handshake process to ensure that the communication between the client and the server is encrypted;

[0098] S2.3. Monitor the current network status in real time, including but not limited to key indicators such as bandwidth, latency, packet loss rate, etc., and execute an adaptive strategy to adjust the data shard size and compression mode according to the real-time network status, so as to ensure efficient and reliable data transmission in an environment with limited or fluctuating bandwidth, while reducing latency and resource consumption;

[0099] The adaptive strategy is as follows: increase the data shard size when the bandwidth is high (in this embodiment, from 1MB to 2MB), and reduce the shard size when the bandwidth is low or the latency is high (such as 512KB); dynamically adjust the number of parallel connections according to the network congestion level (the default number of connections in HTTP / 1.1 is 6, and HTTP / 2 enables multiplexing); in extreme network conditions (in this embodiment, the bandwidth < 1Mbps), degrade to use the fast compression mode (from LZ4-HC to LZ4-fast);

[0100] S2.4. Use the sliding window mechanism to control the sending rate, and ensure the reliability of data transmission through ACK confirmation and timeout retransmission (RTO is dynamically calculated) (using the sliding window mechanism to control the sending rate, through ACK confirmation and timeout retransmission is a network data transmission control method, which ensures reliable data transmission and optimizes transmission efficiency by adjusting the sending window size, using acknowledgment reply (ACK) and setting timeout retransmission). When the network fluctuates, use selective acknowledgment (SACK) to only retransmit the lost data segments instead of the entire data packet. SACK is a technology that optimizes the data retransmission process and improves network transmission efficiency;

[0101] S2.5. Use the built-in security features of the TLS protocol to encrypt (AES-256) and perform integrity verification (HMAC) on the transmitted data. This part of the operation is automatically processed by the TLS layer without additional configuration;

[0102] S2.6. Train a machine learning Q-Learning model based on historical data to predict the changing trend of the network status and adjust the transmission parameters in advance;

[0103] Among them, the machine learning Q-Learning model is a reinforcement learning algorithm that makes decisions by learning the action value function. It enables the agent to select the optimal action in a given state to maximize the long-term reward. By learning and predicting the changing trends of network states, the system can dynamically adjust transmission parameters (such as shard size, compression mode, etc.) to adapt to different network conditions, thus maintaining efficient data transmission in various network environments. This model can react in advance to upcoming network state changes. For example, adjusting the transmission strategy before network congestion helps reduce latency and packet loss rate, providing a more stable quality of service. Intelligent adjustment of data transmission configuration according to real-time network conditions avoids unnecessary resource waste. For example, increasing the data shard size at high bandwidth and reducing the shard size in case of low bandwidth or high latency makes resource usage more reasonable.

[0104] Training the machine learning Q-Learning model based on historical data to predict the changing trends of network states and adjusting transmission parameters in advance includes the following steps:

[0105] S2.61. Collect historical data of network states: including indicators such as bandwidth, latency, packet loss rate, transmission success rate, etc., and record associated transmission parameters: including compression ratio, data block size, transmission protocol type (HTTPS / MQTT), etc. (refer to the transmission unit design in the historical conversation);

[0106] S2.62. Discretize the network state into a multi-dimensional vector;

[0107] S2.63. Define the state of the network, including the current bandwidth, latency, etc., and define adjustable transmission parameters as actions, such as increasing / decreasing the shard size, etc.;

[0108] S2.64. Design a reward function to evaluate the quality of each action. For example, a positive reward can be obtained for improving transmission efficiency, and a negative reward otherwise;

[0109] S2.65. Create a state-action matrix, assign initial values (such as all zeros or random small values) to unknown state-action pairs, and initialize the Q value (the Q value represents the expected long-term return (or cumulative reward) of taking a certain action in a given state);

[0110] S2.66. Adopt the ε-greedy strategy to balance the relationship between exploring new strategies and exploiting existing knowledge (the ε-greedy strategy is a strategy used in reinforcement learning to balance exploration and exploitation. It randomly selects actions with a probability of ε for exploration and selects the currently known best action with a probability of 1 - ε for exploitation), and update the Q value according to the Q value update formula. Introduce an algorithm for dynamically adjusting the learning rate based on entropy weight to optimize the Q value update formula, so that the learning rate Entropy weight of the adaptive system and in view of the problem of balancing compression efficiency and computing resources and the change of network environment, an additional reward term is added to the Q-value update formula for further optimization;

[0111] Furthermore, the Q-value update formula is:

[0112] ;

[0113] wherein, represents the current state; represents the next state; represents the action taken; represents the immediate reward after taking the action; represents the learning rate; represents the discount factor; represents the state when taking the action the expected cumulative reward that can be obtained; represents the time step;

[0114] Entropy weight reflects the degree of uncertainty of the system: when the system is in a high-entropy state (such as multi-objective conflicts and drastic environmental dynamic changes), increases, indicating that a more cautious update strategy is required; conversely, in a low-entropy state (clear goals and stable environment), decreases, allowing a more aggressive parameter update; in a complex network environment, a fixed learning rate may not be able to adapt to different state changes. In a high-entropy environment, the network state changes greatly, and a higher learning rate is required in the learning process to accelerate learning and adapt to these changes. While in a low-entropy environment, the system is relatively stable, and an overly high learning rate may lead to overfitting or unnecessary fluctuations, reducing the system performance; introducing a dynamically adjusted learning rate based on entropy weight aims to enable Q-Learning to automatically adjust the learning speed according to the complexity of the current network state, thereby improving the adaptability and effectiveness of the algorithm. Especially in network transmission optimization, the fluctuations and uncertainties of the network state are relatively large, and the dynamic learning rate can effectively improve the transmission performance and stability of the network;

[0115] Introduce an algorithm for dynamically adjusting the learning rate based on entropy weight into the Q-value update formula for optimization, so that the learning rate adapts to the change of the entropy weight of the system as follows:

[0116] wherein, the algorithm for dynamically adjusting the learning rate based on entropy weight is:

[0117] ;

[0118] wherein, represents the learning rate dynamically adjusted based on entropy weight; represents the initial learning rate; represents the average value of the entropy weight; represents the standard deviation measuring the change of the entropy weight;

[0119] Introducing the algorithm for dynamically adjusting the learning rate based on entropy weight into the Q-value update formula is as follows:

[0120] ;

[0121] In traditional Q-Learning, the reward function usually only focuses on a single goal, such as the success rate of network transmission, latency, bandwidth, etc. However, in real-world network transmission optimization, multiple goals are usually involved; a single reward function may not be able to take these goals into account simultaneously, easily leading to over-optimization of the model in some aspects while ignoring other important factors. For example, optimizing the transmission speed may increase the CPU occupancy rate, or optimizing the compression efficiency may affect the latency. Therefore, introducing additional reward terms can enable the model to perform balanced optimization on multiple goals and avoid the emergence of local optimal solutions; increasing the compression ratio helps reduce the amount of transmitted data, thereby improving the transmission efficiency; a certain amount of computing resources is usually consumed during the compression process. If the compression algorithm is efficient and occupies less CPU resources, it is more conducive to the stability and response speed of the system; the efficiency of the compression algorithm is usually related to the dictionary hit rate. The higher the dictionary hit rate, the better the compression effect of duplicate data, and thus the amount of data transmission is reduced; by adding a compression efficiency reward, Q-Learning can not only focus on traditional network transmission parameters (such as bandwidth, latency, etc.) when updating the Q-value, but also optimize the performance of the compression algorithm; in network transmission, the stability and predictability of the system are very important, especially in high-load or highly variable network environments. Entropy measures the uncertainty or chaos of the system. A higher entropy value usually means the system is unstable or highly variable, which may lead to fluctuations in network performance. The purpose of introducing an entropy stability reward is to encourage the model to maintain the stability of the system as much as possible and avoid excessive fluctuations;

[0122] Regarding the issues of balancing compression efficiency and computing resources and network environment changes, additional reward terms are added to the Q-value update formula for further optimization:

[0123] Among them, the additional reward terms include a compression efficiency reward and an entropy stability reward :

[0124] ;

[0125] Among them, represents a tuning parameter used to measure the entropy stability reward , and the degree of influence on the overall Q-value update;

[0126] Considering that high compression ratio leads to high CPU occupancy rate, a dynamic cooperation factor is introduced in the compression efficiency reward to quantify the conflict degree between the compression ratio and the CPU occupancy rate (i.e., , where the compression ratio represents the compression ratio quantified by the introduced dynamic cooperation factor, and the CPU occupancy rate represents the introduced CPU occupancy rate), and the compression efficiency reward is adjusted through a second-order interaction term (quantifying the conflict between objectives through the real-time ratio of the compression ratio to the CPU occupancy rate, dynamically adjusting the cooperation effect intensity in combination with bandwidth fluctuations; using the hyperbolic tangent function to avoid the influence of extreme values, enabling the model to automatically suppress the conflict reward term when the bandwidth changes suddenly, and the dynamic cooperation factor quantifying the dynamic relationship between objectives through the non-linear interaction term, enhancing the multi-objective balance ability, and adjusting through the second-order interaction term and the real-time negative correlation ( coefficient), solving the problem that traditional linear superposition cannot quantify the objective conflict):

[0127] ;

[0128] Among them, represents the weight coefficient of the compression ratio; represents the weight coefficient of the CPU occupancy rate; represents the weight coefficient of the dictionary hit rate (the probability of successfully finding a duplicate data sequence in the dictionary); represents the cooperation intensity coefficient, which is dynamically adjusted by statistically analyzing the negative correlation between the compression ratio and the CPU occupancy rate in historical data through a sliding window (the stronger the conflict, the smaller); represents the standard deviation of bandwidth fluctuations;

[0129] In this embodiment, the value basis of the weight coefficient is as follows: in a typical network environment (bandwidth 1 Mbps to 100 Mbps, end-to-end delay 10 ms to 500 ms) and multi-type data (text, binary, log) scenarios, the system performance indicators are balanced through parameter tuning. The optimization objective aims to maximize the compression ratio and transmission efficiency while strictly controlling the CPU occupancy rate. Based on the multi-objective optimization theory (Pareto optimality), this solution coordinates the conflicts among the compression ratio, the CPU occupancy rate, and the dictionary hit rate to ensure the optimal compromise among the three; the cooperation intensity coefficient is adaptively adjusted according to the negative correlation between the compression ratio and the CPU occupancy rate in historical data (sliding window statistics);

[0130] The value range of the weight coefficient is: (weight coefficient of the compression ratio): 0.5 to 0.7 (scenario with high compression ratio priority); (weight coefficient of the CPU occupancy rate): 0.2 to 0.4 (scenario with low resource consumption priority); (Weight coefficient of dictionary hit rate): 0.1 - 0.3 (for high repetition rate data scenarios); (Collaboration strength coefficient): 0.05 - 0.2 (dynamically scaled according to the standard deviation of bandwidth fluctuation);

[0131] This formula is a customized improvement for the data compression and transmission scenario under the Q - learning framework. It integrates a multi - objective reward design for compression efficiency and is based on entropy - aware dynamic learning rate and correction terms; traditional Q - learning uses a fixed learning rate. However, in this formula an entropy - dependent dynamic adjustment is introduced, and the newly added compression reward term is used to quantify compression efficiency (including compression ratio and processing speed); combined with the entropy change rate in information theory ( ), it is used to dynamically adjust the compression strategy.

[0132] S2.67. Dynamically adjust the transmission parameters according to the action suggestions output by the machine learning model Q - Learning, including changing the shard size, adjusting the number of concurrent connections, etc.

[0133] The decompression unit 4 at the receiving end obtains the transmitted compressed data from the network and decompresses it using the LZ4 decompression algorithm to restore it to the original data format.

[0134] In this embodiment, the decompression unit 4 at the receiving end includes a decompression and verification module 41 and a data recombination module 42.

[0135] Among them, the decompression and verification module 41 decompresses according to the metadata by calling the LZ4 decompression algorithm, verifies the data integrity. The principle of calling the LZ4 decompression algorithm to decompress is to reverse the matching and pointer operations used in the compression process to restore the compressed data to the original data format.

[0136] The data recombination module 42 merges the chunked data and restores the original structure (the time - series data in this embodiment is sorted by timestamp).

[0137] Embodiment 2: The difference between Embodiment 2 and Embodiment 1 of the present invention is that this embodiment introduces the data compression and transmission method used in the dynamic data compression and transmission system based on big data.

[0138] The dynamic data compression and transmission method based on big data, which is used for the dynamic data compression and transmission system in any one of the above, includes the following steps:

[0139] S3.1. Obtain the original data from multiple data sources and process it.

[0140] S3.2. Compress the original data through the LZ4 compression algorithm, and introduce the hierarchical matching DEW-LZ4 technology based on dynamic entropy weight into the LZ4 compression algorithm to optimize the compression process;

[0141] S3.3. Transmit the compressed data over the network through the HTTPS transport protocol, and dynamically adjust the transmission parameters according to the network status using the machine learning Q-Learning model;

[0142] S3.4. Obtain the transmitted compressed data from the network and decompress it using the LZ4 decompression algorithm to restore it to the original data format.

[0143] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed.

Claims

1. A dynamic data compression and transmission system based on big data, characterized in that: include: A data acquisition and processing unit (1), wherein the data acquisition and processing unit (1) acquires raw data from a variety of data sources and processes the data; A data compression execution unit (2), wherein the data compression execution unit (2) compresses the original data using an LZ4 compression algorithm, and introduces a layered matching DEW-LZ4 technology based on dynamic entropy weights into the LZ4 compression algorithm to find data sequences that appear repeatedly in the data block and optimize the compression process; The method of using the hierarchical matching DEW-LZ4 technology based on dynamic entropy weight to find repeated data sequences in the data block includes the following steps: S1.

11. Calculate the Shannon entropy of the current data block , and based on metadata tags, establish an entropy value-matching weight model; S1.

12. Calculate the entropy weight of the data block based on the entropy value-matching weight model ; S1.13, if the entropy weight If the entropy weight is greater than the preset threshold A, a convolutional neural network is used to predict matching hotspots. If the entropy weight is greater than the preset threshold B and less than or equal to the preset threshold A, a three-level hash index is constructed for matching; If it is less than or equal to the preset threshold B, the traditional LZ4 algorithm is continued to be used for matching; S1.

14. In all cases, when the Shannon entropy value change rate exceeds the set threshold When , the update mechanism of the double buffer dictionary structure is triggered; A data transmission unit (3), wherein the data transmission unit (3) transmits the compressed data over a network via a transmission protocol HTTPS, and dynamically adjusts transmission parameters using a machine learning Q-Learning model according to a network state; A receiving end decompression unit (4), wherein the receiving end decompression unit (4) obtains the transmitted compressed data from the network, and decompresses the data using an LZ4 decompression algorithm to restore the data to an original data format.

2. The big data based dynamic data compression transmission system according to claim 1, characterized in that: The data acquisition and processing unit (1) comprises a data acquisition module (11) and a data processing module (12); The data acquisition module (11) supports real-time acquisition of multi-source heterogeneous data, and performs preliminary classification and priority marking on the original multi-source heterogeneous data; The data processing module (12) analyzes data features based on the preliminary classified and labeled data, generates metadata tags, and divides the data into blocks of fixed size m.

3. The big data based dynamic data compression transmission system according to claim 2, characterized in that: The data compression execution unit (2) compresses the original data using the LZ4 compression algorithm, and introduces a layered matching DEW-LZ4 technology based on dynamic entropy weight into the LZ4 compression algorithm to optimize the compression process, including the following steps: S1.

1. Use the DEW-LZ4 technology based on dynamic entropy weight to find repeated data sequences in the data block; S1.

2. For the matching part, LZ4 uses a pointer to indicate the position and length information of the matching sequence relative to the current position; S1.3, if the length of the matching sequence found is greater than the threshold n, the original data is replaced by the dictionary pointer; for the part where no match is found, the original data is copied; S1.

4. Concatenate the compressed data blocks and matching information through LZ4 to form compressed output data.

4. The big data based dynamic data compression transmission system according to claim 1, characterized in that: In S1.11, the entropy value-matching weight model is: ; in, represents the regulating factor; Represents the benchmark Shannon entropy value.

5. The big data based dynamic data compression transmission system according to claim 4, characterized in that: The data transmission unit (3) transmits the compressed data over the network via the transmission protocol HTTPS and dynamically adjusts the transmission parameters according to the network status, including the following steps: S2.1, receiving a compressed data block from the data compression execution unit (2); S2.

2. Establish a secure HTTPS connection based on the target server information; S2.3, monitor the current network status in real time, and execute adaptive strategies to adjust the data segment size and compression mode according to the real-time network status; S2.4, using a sliding window mechanism to control the sending rate, through ACK confirmation and timeout retransmission, when the network fluctuates, using selective confirmation to retransmit only the lost data segments; S2.

5. Use the built-in security features of the TLS protocol to encrypt and verify the integrity of the transmitted data; S2.

6. Train the machine learning Q-Learning model based on historical data to predict the trend of network status changes and adjust the transmission parameters in advance.

6. The big data based dynamic data compression transmission system according to claim 5, characterized in that: In S2.6, training a machine learning Q-Learning model based on historical data, predicting the trend of network status changes, and adjusting transmission parameters in advance include the following steps: S2.

61. Collect historical data of network status and record associated transmission parameters; S2.62, discretize the network state into a multi-dimensional vector; S2.63, define the state of the network and define adjustable transmission parameters as actions; S2.

64. Design a reward function to evaluate the quality of each action; S2.65, create a state-action matrix, assign initial values ​​to unknown state-action pairs, and initialize Q values; S2.66, use the ε-greedy strategy to balance the relationship between exploring new strategies and using existing knowledge, and update the Q value according to the Q value update formula. In the Q value update formula, an algorithm for dynamically adjusting the learning rate based on entropy weight is introduced to optimize the learning rate. Entropy weight of the adaptive system In order to balance compression efficiency and computing resources and network environment changes, we added additional reward items to the Q value update formula for further optimization. S2.

67. Dynamically adjust transmission parameters based on the action recommendations output by the machine learning Q-Learning model.

7. The big data based dynamic data compression transmission system according to claim 6, characterized in that: In S2.66, the Q value update formula is: ; in, Indicates the current state; Indicates the latter state; Indicates the action taken; Indicates the immediate reward after taking an action; represents the learning rate; represents the discount factor; Indicates status Take action The expected cumulative rewards that can be obtained; represents the time step; The algorithm for dynamically adjusting the learning rate based on entropy weight is introduced into the Q value update formula for optimization, so that the learning rate Entropy weight of the adaptive system Changes: Among them, the algorithm for dynamically adjusting the learning rate based on entropy weight is: ; in, Represents the learning rate dynamically adjusted based on the entropy weight; represents the initial learning rate; represents the average value of entropy weight; represents the standard deviation of the change in entropy weight; The algorithm for dynamically adjusting the learning rate based on entropy weight is introduced into the Q value update formula as follows: ; In order to balance compression efficiency, computing resources and network environment changes, additional reward items are added to the Q value update formula for further optimization: Among them, additional rewards include compression efficiency rewards and entropy stability bonus : ; in, represents the adjustment parameter; Considering that high compression rate will cause high CPU usage, reward compression efficiency A dynamic synergy factor is introduced to quantify the conflict between compression rate and CPU usage, and the compression efficiency reward is adjusted through the second-order interaction term: ; in, The weight coefficient representing the compression ratio; Indicates the weight coefficient of CPU usage; The weight coefficient representing the dictionary hit rate; represents the synergy strength coefficient; Indicates the standard deviation of bandwidth fluctuation.

8. The big data based dynamic data compression transmission system according to claim 7, characterized in that: The receiving end decompression unit (4) comprises a decompression and verification module (41) and a data reassembly module (42); The decompression and verification module (41) calls the LZ4 decompression algorithm to decompress data according to the metadata and verifies the data integrity; The data reorganization module (42) merges the block data and restores the original structure.

9. A method for dynamic data compression and transmission based on big data, used in a dynamic data compression and transmission system based on big data as claimed in any one of claims 1 to 8, characterized in that: The steps include: S3.

1. Obtain and process raw data from various data sources; S3.2, compressing the original data by LZ4 compression algorithm, and introducing the layered matching DEW-LZ4 technology based on dynamic entropy weight into the LZ4 compression algorithm to optimize the compression process; S3.3, the compressed data is transmitted over the network through the transmission protocol HTTPS, and the transmission parameters are dynamically adjusted according to the network status using the machine learning Q-Learning model; S3.4, obtain the transmitted compressed data from the network, and decompress it using the LZ4 decompression algorithm to restore it to the original data format.

Citation Information

Patent Citations

  • Efficient data compression storage algorithm based on adaptive entropy optimization model

    CN119543954A

  • KR20200037700A