Self-adaptive data compression method, system and equipment
By calculating the Shannon entropy value and repetition rate of the data, adaptively selecting the optimal data compression algorithm, solving the problems of low efficiency and poor adaptability in complex data processing, and achieving efficient data compression effect.
Patent Information
- Application Number
- CN202510854517.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-25
- Publication Date
- 2025-07-22
- Estimated Expiration
- 2045-06-25
AI Technical Summary
Traditional data compression algorithms have low compression efficiency and poor adaptability when processing complex data, making it difficult to meet the needs of different types of data.
By calculating the Shannon entropy value and repetition rate of the data to be compressed, the data is divided into the first and second halfs, and the adaptive compression algorithm is used to determine the optimal data compression algorithm for the model to select, including standard LZW, reset LZW, block LZW, sliding window LZW and threshold-based dictionary reset LZW, for adaptive compression processing.
It improves data compression efficiency and universality, and is suitable for many types of data compression, especially in the field of large-scale data storage and transmission.
Smart Images

Figure CN120357909A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of data compression and information processing, and particularly relates to an adaptive data compression method, system and device. Background Art
[0002] With the rapid development of information technology, the amount of data is increasing continuously, and the demand for data compression algorithms is increasing day by day. Traditional data compression algorithms, such as Huffman coding and arithmetic coding, often face problems such as low compression efficiency and poor adaptability to different data types when dealing with a large amount of complex data. Especially when dealing with complex data with a high entropy value or strong repeatability, the existing compression methods often cannot achieve an ideal compression ratio, resulting in poor compression effects.
[0003] As an emerging compression technology, the adaptive compression method can dynamically adjust the compression method according to different characteristics of data, improving the compression efficiency and decompression speed. However, there are still certain limitations in the method selection and optimization of traditional adaptive compression methods, making it difficult to fully meet the needs of different types of data. Therefore, there is an urgent need for a data compression method based on an adaptive method that can flexibly select a compression method according to different characteristics of data, thereby improving the compression ratio while reducing the computational complexity. Summary of the Invention
[0004] The purpose of this application is to provide an adaptive data compression method, system and device, improving the compression efficiency of data to be compressed and the universality of the compression method.
[0005] To achieve the above purpose, this application provides the following solutions: In the first aspect, this application provides an adaptive data compression method, and the adaptive data compression method includes: Calculating the Shannon entropy value and the repetition rate of the data to be compressed; the repetition rate of the data to be compressed is the occurrence probability of the highest-frequency character in the data to be compressed; Dividing the data to be compressed into the first half and the second half equally according to the data length, and calculating the Shannon entropy value of the first half of the data to be compressed and the Shannon entropy value of the second half of the data to be compressed; Based on the Shannon entropy value of the data to be compressed, the repetition rate of the data to be compressed, the Shannon entropy value of the first half of the data to be compressed, the Shannon entropy value of the second half of the data to be compressed and the adaptive compression algorithm determination model, determining the target data compression algorithm; the adaptive compression algorithm determination model is obtained by training a feedforward fully-connected neural network model with a training set, and the training set includes: the Shannon entropy value of the sample data, the repetition rate of the sample data, the Shannon entropy value of the first half of the sample data, the Shannon entropy value of the second half of the sample data and the categories of multiple data compression algorithms; Using the target data compression algorithm to perform compression processing on the data to be compressed to obtain an optimal compressed codeword sequence.
[0006] In one embodiment, calculating the Shannon entropy value and the repetition rate of the data to be compressed specifically includes: Calculating the occurrence times of each character in the data to be compressed; Based on the occurrence times of each character and the data length of the data to be compressed, determining the occurrence probability of the highest-frequency character in the data to be compressed, and taking the occurrence probability of the highest-frequency character as the repetition rate of the data to be compressed; Based on the occurrence probabilities of each character in the data to be compressed, calculating the Shannon entropy value of the data to be compressed.
[0007] In one embodiment, based on the Shannon entropy value of the data to be compressed, the repetition rate of the data to be compressed, the Shannon entropy value of the first half of the data to be compressed, the Shannon entropy value of the second half of the data to be compressed, and the adaptive compression algorithm determination model, determining the target data compression algorithm specifically includes: Inputting the Shannon entropy value of the data to be compressed, the repetition rate of the data to be compressed, the Shannon entropy value of the first half of the data to be compressed, and the Shannon entropy value of the second half of the data to be compressed into the adaptive compression algorithm determination model, and outputting the optimal category; Determining the data compression algorithms corresponding to each optimal category as the target data compression algorithm.
[0008] In one embodiment, the optimal category is one or more, and one optimal category corresponds to one target data compression algorithm.
[0009] In one embodiment, using the target data compression algorithm to perform compression processing on the data to be compressed to obtain the optimal compressed codeword sequence specifically includes: When the target data compression algorithm is one, using the target data compression algorithm to perform compression processing on the data to be compressed to obtain the optimal compressed codeword sequence; When the target data compression algorithm is multiple, using multiple target data compression algorithms to perform compression processing on the data to be compressed respectively to obtain the compressed codeword sequences corresponding to the multiple target data compression algorithms; Determining the compression rate of the compressed codeword sequences corresponding to each target data compression algorithm; Determining the compressed codeword sequence corresponding to the target data compression algorithm with the highest compression rate as the optimal compressed codeword sequence.
[0010] In one embodiment, the training process of the adaptive compression algorithm determination model specifically includes: Constructing a training set; Constructing a feedforward fully connected neural network model; Taking the Shannon entropy value of the sample data, the repetition rate of the sample data, the Shannon entropy value of the first half of the sample data, and the Shannon entropy value of the second half of the sample data as inputs, and taking the data compression algorithm as the output, the feedforward fully-connected neural network model is iteratively trained until the number of iterations reaches the maximum value or the loss function reaches the minimum value, and the iterative training is stopped to obtain the adaptive compression algorithm determination model.
[0011] In one embodiment, the data compression algorithm includes: standard LZW compression algorithm, reset LZW compression algorithm, block LZW compression algorithm, sliding window LZW compression algorithm, and threshold-based dictionary reset LZW compression algorithm.
[0012] In one embodiment, the adaptive data compression method further includes: Using the corresponding data decompression algorithm to decompress the optimal compressed codeword sequence to obtain the decompressed data; Comparing whether the decompressed data is consistent with the data to be compressed; If so, taking the optimal compressed codeword sequence as the final compression result of the data to be compressed; If not, updating the model parameters of the adaptive compression algorithm determination model, and using the updated adaptive compression algorithm determination model to determine the target data compression algorithm.
[0013] In a second aspect, the present application provides an adaptive data compression system, which is used to implement the adaptive data compression method. The adaptive data compression system includes: The first calculation unit is used to calculate the Shannon entropy value and repetition rate of the data to be compressed; the repetition rate of the data to be compressed is the occurrence probability of the highest-frequency character in the data to be compressed; The second calculation unit is used to equally divide the data to be compressed into the first half and the second half according to the data length, and calculate the Shannon entropy value of the first half of the data to be compressed and the Shannon entropy value of the second half of the data to be compressed; The target data compression algorithm determination unit is used to determine the target data compression algorithm based on the Shannon entropy value of the data to be compressed, the repetition rate of the data to be compressed, the Shannon entropy value of the first half of the data to be compressed, the Shannon entropy value of the second half of the data to be compressed, and the adaptive compression algorithm determination model; the adaptive compression algorithm determination model is obtained by training the feedforward fully-connected neural network model with a training set, and the training set includes: the Shannon entropy value of the sample data, the repetition rate of the sample data, the Shannon entropy value of the first half of the sample data, the Shannon entropy value of the second half of the sample data, and the categories of various data compression algorithms; The optimal compressed codeword sequence determination unit is used to compress the data to be compressed by using the target data compression algorithm to obtain the optimal compressed codeword sequence.
[0014] In a third aspect, the present application provides a computer device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor executes the computer program to implement the adaptive data compression method described in any one of the above.
[0015] According to the specific embodiments provided by the present application, the present application has the following technical effects: The present application discloses an adaptive data compression method, system, and device. First, calculate the Shannon entropy value and repetition rate of the data to be compressed; secondly, divide the data to be compressed into the first half and the second half according to the data length, and calculate the Shannon entropy value of the first half of the data to be compressed and the Shannon entropy value of the second half of the data to be compressed; furthermore, based on the Shannon entropy value of the data to be compressed, the repetition rate of the data to be compressed, the Shannon entropy value of the first half of the data to be compressed, the Shannon entropy value of the second half of the data to be compressed, and the adaptive compression algorithm determination model, determine the target data compression algorithm; finally, use the target data compression algorithm to compress the data to be compressed to obtain the optimal compressed codeword sequence. By establishing an adaptive compression algorithm determination model, the present application can adaptively select the optimal data compression method that matches the data to be compressed, improving the data compression efficiency; this method is applicable to the compression of various types of data, improving the universality of the data compression method. Description of the Drawings
[0016] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0017] Figure 1 Schematic flowchart of the adaptive data compression method provided by an embodiment of the present application; Figure 2 Schematic diagram of the training process provided by an embodiment of the present application; Figure 3 Schematic diagram of the structure of a computer device provided by an embodiment of the present application. Detailed Embodiments
[0018] The following will clearly and completely describe the technical solutions in the embodiments of the present application in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present application.
[0019] To make the above objects, features, and advantages of the present application more apparent and understandable, the present application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0020] In an exemplary embodiment, as Figure 1 shown, an adaptive data compression method is provided, including the following steps. Wherein: Step S1, calculate the Shannon entropy value and the repetition rate of the data to be compressed; the repetition rate of the data to be compressed is the occurrence probability of the highest-frequency character in the data to be compressed.
[0021] As an optional implementation manner, step S1 specifically includes: Step S11, calculate the occurrence times of each character in the data to be compressed.
[0022] Step S12, based on the occurrence times of each character and the data length of the data to be compressed, determine the occurrence probability of the highest-frequency character in the data to be compressed, and use the occurrence probability of the highest-frequency character as the repetition rate of the data to be compressed. The occurrence probability of each character is calculated using the following formula: p i = n i / N (1) Wherein, p i represents the occurrence probability of the i th character; n i represents the occurrence times of the i th character; N represents the data length.
[0023] Step S13, based on the occurrence probabilities of each character in the data to be compressed, calculate the Shannon entropy value of the data to be compressed. The calculation formula of the Shannon entropy value of the data to be compressed is as follows: H = -∑ p i log2( p i ) (2) Wherein, H is the Shannon entropy value of the data to be compressed.
[0024] Specifically, first, the entire segment of data to be compressed is used as input. In order to count the occurrence times of each character in the data to be compressed, an empty frequency dictionary is first established, and then each character in the data to be compressed is scanned in turn; every time a character is encountered, the count of that character in the dictionary is incremented by one. After the character statistics are completed, the total data length N (i.e., the total number of characters) and the occurrence times ni of each character are obtained.
[0025] Next, calculate the occurrence probability of each character p i, and based on this, the Shannon entropy value H is calculated. At the same time, the occurrence probabilities of all characters are compared, and the maximum probability value (i.e., the occurrence probability of the highest-frequency character) is taken as the repetition rate of the data to be compressed. R max . Among them, the Shannon entropy value is used to evaluate the randomness of the data to be compressed, and the repetition rate is used to evaluate the repeatability of the data to be compressed.
[0026] Step S2: Divide the data to be compressed into the first half and the second half according to the data length, and calculate the Shannon entropy value of the first half of the data to be compressed and the Shannon entropy value of the second half of the data to be compressed.
[0027] The calculation formulas for the Shannon entropy value of the first half of the data to be compressed and the Shannon entropy value of the second half of the data to be compressed are shown in Equation (2).
[0028] Step S3: Determine the target data compression algorithm based on the Shannon entropy value of the data to be compressed, the repetition rate of the data to be compressed, the Shannon entropy value of the first half of the data to be compressed, the Shannon entropy value of the second half of the data to be compressed, and the adaptive compression algorithm determination model; the adaptive compression algorithm determination model is obtained by training a feedforward fully connected neural network model using a training set, and the training set includes: the Shannon entropy value of the sample data, the repetition rate of the sample data, the Shannon entropy value of the first half of the sample data, the Shannon entropy value of the second half of the sample data, and the categories of various data compression algorithms.
[0029] As an optional implementation manner, in step S3, the data compression algorithms include: the standard LZW (short for Lempel-Ziv-Welch) compression algorithm, the reset LZW compression algorithm, the block LZW compression algorithm, the sliding window LZW compression algorithm, and the dictionary reset LZW compression algorithm based on a threshold.
[0030] As an optional implementation manner, in step S3, the training process of the adaptive compression algorithm determination model specifically includes: Step S300: Construct a training set.
[0031] Step S301: Construct a feedforward fully connected neural network model.
[0032] Step S302: Use the Shannon entropy value of the sample data, the repetition rate of the sample data, the Shannon entropy value of the first half of the sample data, and the Shannon entropy value of the second half of the sample data as inputs, and use the data compression algorithm as the output to perform iterative training on the feedforward fully connected neural network model until the number of iterations reaches the maximum value or the loss function reaches the minimum value, and then stop the iterative training to obtain the adaptive compression algorithm determination model. The schematic diagram of the training process is as Figure 2 shown.
[0033] Among them, when determining the model for the training of the adaptive compression algorithm, the Shannon entropy value of the data, the repetition rate, the Shannon entropy value of the first half of the data, and the Shannon entropy value of the second half of the data are used as input features. Additionally, the entropy difference (i.e., the difference between the Shannon entropy value of the first half of the sample data and the Shannon entropy value of the second half of the sample data) can be added to accelerate the convergence of the model determined by the adaptive compression algorithm.
[0034] Specifically, during model training, the optimal data compression algorithm is adaptively selected through a preset rule. The preset rule is as follows: 1) If the overall Shannon entropy value of the data to be compressed is higher than the preset Shannon entropy threshold and the repetition rate (i.e., the occurrence probability of the most frequent character) is lower than the preset probability threshold, indicating that the data is relatively random and lacks obvious repeatability, then the standard LZW compression algorithm is selected.
[0035] 2) If the data length of the data to be compressed is less than the length threshold, and the overall Shannon entropy value of the data to be compressed is lower than the preset Shannon entropy threshold or the repetition rate is higher than the probability threshold, then the reset LZW compression algorithm is selected to efficiently utilize the repeat pattern in short data and avoid the overhead of complex methods.
[0036] 3) If the data length of the data to be compressed is greater than the length threshold and the overall Shannon entropy value is lower than the preset Shannon entropy threshold (indicating that the data is long and generally has a high repeatability or redundancy), then the dictionary reset LZW compression algorithm based on the threshold is selected to prevent the dictionary from growing excessively and improve the overall compression efficiency.
[0037] 4) If the structural strength of the data to be compressed exceeds the preset strength threshold or the repetition rate is higher than the probability threshold, then the block LZW compression algorithm is selected.
[0038] Specifically, if the structural strength of the data to be compressed exceeds the preset strength threshold, it is considered that the data to be compressed has strong structure; otherwise, it is considered that the data to be compressed has weak structure. The key to judging the "structural" strength of the data to be compressed lies in identifying whether there are reusable stable patterns between and within the windows. That is, the data to be compressed is divided into multiple windows of a fixed length. If the average intersection-over-union ratio of the shared substrings between adjacent windows is not lower than the preset intersection-over-union ratio threshold, or high-frequency repeated frame header / frame tail bytes are found at fixed offsets in each window, the block LZW compression algorithm is selected.
[0039] In actual operation, first, the data to be compressed is sliced into a series of consecutive windows with a fixed length. Then, several substrings of appropriate length are extracted within each window, and the set of these substrings is used to measure the similarity between adjacent windows: if more than about half of the substrings appear repeatedly in adjacent windows, it indicates that the content of each window is highly templated and belongs to strongly structured data; at the same time, the byte distribution is also statistically analyzed at a fixed offset position in each window. If the same frame header or frame tail bytes appear at the same offset in a large number of windows, and the number of such "template bits" reaches more than eight or accounts for more than ten percent of the window bytes, it also indicates that the data to be compressed has an obvious protocol frame structure. For the data to be compressed that meets any of the above situations, the block LZW compression algorithm can be used to independently maintain a dictionary for each window, thus avoiding the dilution of the compression effect by differences between different blocks and making full use of the structural information to improve the compression effect. When both detection results are not significant, it can be regarded as weakly structured.
[0040] 5) If the data to be compressed has strong local dependencies or temporal variability, the sliding window LZW compression algorithm is selected.
[0041] Specifically, if the conditional entropy of the characters in the data to be compressed within a short distance is lower than the preset conditional entropy threshold, and the Kullback-Leibler (KL) divergence between adjacent windows is lower than the preset smoothing threshold, the sliding window LZW compression algorithm is selected to dynamically capture and utilize the local repetitive patterns in the data.
[0042] Among them, "local dependence or temporal variability" means that the same data stream can be predicted with high precision within a very short byte interval, and this predictable statistical pattern drifts slowly over time. To quantify this phenomenon, first divide the entire data to be compressed into consecutive windows according to a fixed number of bytes (such as 256 bytes, 512 bytes, or 1 KB), and then count the joint occurrence frequencies of adjacent or near-neighbor bytes (such as 1 to 3 byte intervals) within each window, so as to calculate the average conditional entropy within the window. When this conditional entropy is significantly lower than the preset conditional entropy threshold (such as 3.5 bits, much lower than the maximum uncertainty of 8 bits for a single byte), it indicates that the "short-distance" predictability is very strong, that is, the local dependence is significant. Then, calculate the Kullback-Leibler divergence of the overall character distributions of two adjacent windows and perform smooth tracking. If this divergence is continuously lower than the preset smoothing threshold (such as 0.10 to 0.15) for a long time, it indicates that the statistical distribution evolves very smoothly over time, showing "pattern slippage" rather than mutation. Only when both low conditional entropy and low KL divergence are observed simultaneously is it determined that the data has strong local dependence and temporal smooth variation, and the sliding window LZW compression algorithm is selected because the dictionary established in the previous window still has a high hit rate in the next window, and the gradually sliding out old entries can make room for the gradually changing new patterns. On the contrary, if the conditional entropy does not drop below the threshold or the KL divergence frequently exceeds the limit, it means that either the local predictability is insufficient or the pattern change is too drastic, and the sliding window LZW compression algorithm is difficult to bring advantages. Other data compression methods can be selected based on global indicators such as entropy value, repetition rate, and data length.
[0043] All thresholds are initially determined through cross-validation on typical data sets and written back to the feature library together with the compression effect during operation for the neural network to adjust the model parameters online, so as to ensure that the judgment criteria can be automatically updated for new types of data, thereby realizing public, sufficient, and implementable local dependence detection and algorithm selection.
[0044] During training, set the corresponding compression rate threshold, store the compression results and compression rates of the compression algorithms whose compression rates of the output target data compression algorithm are better than the compression rate threshold in the compression feature library, and combine the compression feature library to process the subsequent data using pattern recognition technology to obtain an adaptive compression algorithm determination model, improving the efficiency of data compression and the accuracy of decompression. Then, feedback the compression rate of each data compression algorithm to the compression feature library through a feedback mechanism, so as to improve the selection of subsequent data compression algorithms, enabling the adaptive compression algorithm determination model to dynamically adjust the optimal data compression algorithm according to the latest data characteristics, optimize the compression rate and decompression performance, and enhance the adaptability and compression efficiency of the adaptive compression algorithm determination model.
[0045] As an optional implementation manner, step S3 specifically includes: Step S31: Input the Shannon entropy value of the data to be compressed, the repetition rate of the data to be compressed, the Shannon entropy value of the first half of the data to be compressed, and the Shannon entropy value of the second half of the data to be compressed into the adaptive compression algorithm determination model, and output the optimal category.
[0046] Step S32: Determine the data compression algorithm corresponding to each optimal category as the target data compression algorithm.
[0047] As an optional implementation manner, in step S31, the optimal category is one or more, and one optimal category corresponds to one target data compression algorithm.
[0048] Step S4: Use the target data compression algorithm to perform compression processing on the data to be compressed, and obtain the optimal compressed codeword sequence.
[0049] As an optional implementation manner, step S4 specifically includes: Step S41: When the target data compression algorithm is one, use the target data compression algorithm to perform compression processing on the data to be compressed, and obtain the optimal compressed codeword sequence.
[0050] Step S42: When the target data compression algorithm is multiple, use the multiple target data compression algorithms to perform compression processing on the data to be compressed respectively, and obtain the compressed codeword sequences corresponding to the multiple target data compression algorithms.
[0051] Step S43: Determine the compression ratio of the compressed codeword sequences corresponding to each target data compression algorithm.
[0052] Step S44: Determine the compressed codeword sequence corresponding to the target data compression algorithm with the highest compression ratio as the optimal compressed codeword sequence.
[0053] Specifically, when the data to be compressed meets two or more preset rules at the same time, it indicates that the data to be compressed has complex or composite structural characteristics. First, perform compression processing on the data to be compressed according to each rule respectively and calculate the compression ratio. Finally, select the compressed codeword sequence corresponding to the data compression algorithm with the highest compression ratio as the optimal compressed codeword sequence.
[0054] As an optional implementation manner, in step S4, the compression processing includes: by matching characters one by one, updating the dictionary and generating corresponding codewords. When the dictionary reaches the maximum capacity, perform reset or block operations according to different methods to ensure that the dictionary remains effective during the compression process and avoid dictionary overflow.
[0055] As an optional implementation manner, the adaptive data compression method further includes: Step S5: Use the corresponding data decompression algorithm to decompress the optimal compressed codeword sequence to obtain the decompressed data.
[0056] Step S6, compare whether the decompressed data is the same as the data to be compressed.
[0057] Step S7, if so, use the optimal compression codeword sequence as the final compression result of the data to be compressed.
[0058] Step S8, if not, update the model parameters of the adaptive compression algorithm determination model, and use the updated adaptive compression algorithm determination model to determine the target data compression algorithm.
[0059] Specifically, use the corresponding LZW decompression algorithm to restore the original data (i.e., the data to be compressed), ensure the data integrity during the decompression process, and avoid the situation that the decompressed data is inconsistent with the original data.
[0060] Beneficial effects: By establishing an adaptive compression algorithm determination model to perform real-time analysis on the characteristics of the data to be compressed and adaptively select the optimal compression algorithm, it not only effectively improves the compression efficiency, but also can adapt to various different types of data, with good universality and compression performance. Compared with traditional compression technologies, it has obvious advantages in processing complex data, especially in the field of large-scale data storage and transmission, and has broad application prospects.
[0061] Based on the same inventive concept, the embodiments of the present application also provide an adaptive data compression system for implementing the above-mentioned adaptive data compression method. The implementation solutions provided by this system to solve problems are similar to the implementation solutions described in the above method. Therefore, the specific limitations in one or more embodiments of the following adaptive data compression system can refer to the limitations on the adaptive data compression method in the above text, and will not be repeated here.
[0062] In an exemplary embodiment, an adaptive data compression system is provided, including: A first calculation unit, configured to calculate the Shannon entropy value and the repetition rate of the data to be compressed; the repetition rate of the data to be compressed is the occurrence probability of the highest-frequency character in the data to be compressed.
[0063] A second calculation unit, configured to equally divide the data to be compressed into a first half and a second half according to the data length, and calculate the Shannon entropy value of the first half of the data to be compressed and the Shannon entropy value of the second half of the data to be compressed.
[0064] A target data compression algorithm determination unit, configured to determine a target data compression algorithm based on the Shannon entropy value of the data to be compressed, the repetition rate of the data to be compressed, the Shannon entropy value of the first half of the data to be compressed, the Shannon entropy value of the second half of the data to be compressed, and an adaptive compression algorithm determination model; the adaptive compression algorithm determination model is obtained by training a feedforward fully-connected neural network model with a training set, and the training set includes: the Shannon entropy value of the sample data, the repetition rate of the sample data, the Shannon entropy value of the first half of the sample data, the Shannon entropy value of the second half of the sample data, and the categories of multiple data compression algorithms.
[0065] An optimal compression codeword sequence determination unit, configured to perform compression processing on the data to be compressed by using the target data compression algorithm to obtain an optimal compression codeword sequence.
[0066] In an exemplary embodiment, a computer device is provided, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and the processor executes the computer program to implement an adaptive data compression method.
[0067] In an exemplary embodiment, a computer device is provided. The computer device may be a server or a terminal, and its internal structure diagram may be as Figure 3 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. The computer program, when executed by the processor, implements an adaptive data compression method.
[0068] Those skilled in the art can understand that Figure 3 the structure shown in is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.
[0069] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0070] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc.
[0071] The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.
[0072] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered to be within the scope described in this specification.
[0073] In this text, specific examples are used to elaborate on the principles and implementation manners of the present application. The description of the above embodiments is only used to help understand the method, system and its core idea of the present application; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. An adaptive data compression method, characterized in that, The described adaptive data compression method includes: Calculating the Shannon entropy value and the repetition rate of the data to be compressed; the repetition rate of the data to be compressed is the occurrence probability of the highest-frequency character in the data to be compressed; Dividing the data to be compressed into the first half and the second half equally according to the data length, and calculating the Shannon entropy value of the first half of the data to be compressed and the Shannon entropy value of the second half of the data to be compressed; Based on the Shannon entropy value of the data to be compressed, the repetition rate of the data to be compressed, the Shannon entropy value of the first half of the data to be compressed, the Shannon entropy value of the second half of the data to be compressed, and an adaptive compression algorithm determination model, determining the target data compression algorithm; the adaptive compression algorithm determination model is obtained by training a feedforward fully-connected neural network model with a training set, and the training set includes: the Shannon entropy value of the sample data, the repetition rate of the sample data, the Shannon entropy value of the first half of the sample data, the Shannon entropy value of the second half of the sample data, and the categories of multiple data compression algorithms; Using the target data compression algorithm to perform compression processing on the data to be compressed to obtain the optimal compressed codeword sequence.
2. The adaptive data compression method according to claim 1, wherein Calculating the Shannon entropy value and the repetition rate of the data to be compressed specifically includes: Calculating the occurrence times of each character in the data to be compressed; Based on the occurrence times of each character and the data length of the data to be compressed, determining the occurrence probability of the highest-frequency character in the data to be compressed, and taking the occurrence probability of the highest-frequency character as the repetition rate of the data to be compressed; Based on the occurrence probabilities of each character in the data to be compressed, calculating the Shannon entropy value of the data to be compressed.
3. The adaptive data compression method according to claim 1, wherein Based on the Shannon entropy value of the data to be compressed, the repetition rate of the data to be compressed, the Shannon entropy value of the first half of the data to be compressed, the Shannon entropy value of the second half of the data to be compressed, and an adaptive compression algorithm determination model, determining the target data compression algorithm specifically includes: Inputting the Shannon entropy value of the data to be compressed, the repetition rate of the data to be compressed, the Shannon entropy value of the first half of the data to be compressed, and the Shannon entropy value of the second half of the data to be compressed into the adaptive compression algorithm determination model, and outputting the optimal category; Determining the data compression algorithm corresponding to each optimal category as the target data compression algorithm.
4. The adaptive data compression method according to claim 3, wherein The optimal category is one or more, and one optimal category corresponds to one target data compression algorithm.
5. The adaptive data compression method according to claim 4, wherein Using the target data compression algorithm to perform compression processing on the data to be compressed to obtain the optimal compressed codeword sequence specifically includes: When the target data compression algorithm is one, using the target data compression algorithm to perform compression processing on the data to be compressed to obtain the optimal compressed codeword sequence; When the target data compression algorithm is multiple, using multiple target data compression algorithms to perform compression processing on the data to be compressed respectively to obtain the compressed codeword sequences corresponding to the multiple target data compression algorithms; Determining the compression ratio of the compressed codeword sequences corresponding to each target data compression algorithm; Determining the compressed codeword sequence corresponding to the target data compression algorithm with the highest compression ratio as the optimal compressed codeword sequence.
6. The adaptive data compression method according to claim 1, wherein The training process of the adaptive compression algorithm determination model specifically includes: Constructing a training set; Constructing a feedforward fully-connected neural network model; Taking the Shannon entropy value of the sample data, the repetition rate of the sample data, the Shannon entropy value of the first half of the sample data, and the Shannon entropy value of the second half of the sample data as inputs, and taking the data compression algorithm as the output, iteratively train the feedforward fully connected neural network model until the number of iterations reaches the maximum value or the loss function reaches the minimum value, then stop the iterative training to obtain the adaptive compression algorithm determination model.
7. The adaptive data compression method according to claim 1, characterized in that, The data compression algorithms include: standard LZW compression algorithm, reset LZW compression algorithm, block LZW compression algorithm, sliding window LZW compression algorithm, and dictionary reset LZW compression algorithm based on threshold.
8. The adaptive data compression method according to claim 1, wherein The adaptive data compression method further includes: Using the corresponding data decompression algorithm to decompress the optimal compressed codeword sequence to obtain the decompressed data; Comparing whether the decompressed data is consistent with the data to be compressed; If so, taking the optimal compressed codeword sequence as the final compression result of the data to be compressed; If not, updating the model parameters of the adaptive compression algorithm determination model, and using the updated adaptive compression algorithm determination model to determine the target data compression algorithm.
9. An adaptive data compression system, characterized in that, The adaptive data compression system is used to implement the adaptive data compression method according to any one of claims 1-8. The adaptive data compression system includes: The first calculation unit is used to calculate the Shannon entropy value and repetition rate of the data to be compressed; the repetition rate of the data to be compressed is the occurrence probability of the highest frequency character in the data to be compressed; The second calculation unit is used to equally divide the data to be compressed into the first half and the second half according to the data length, and calculate the Shannon entropy value of the first half of the data to be compressed and the Shannon entropy value of the second half of the data to be compressed; The target data compression algorithm determination unit is used to determine the target data compression algorithm based on the Shannon entropy value of the data to be compressed, the repetition rate of the data to be compressed, the Shannon entropy value of the first half of the data to be compressed, the Shannon entropy value of the second half of the data to be compressed, and the adaptive compression algorithm determination model; the adaptive compression algorithm determination model is obtained by training the feedforward fully connected neural network model with a training set, and the training set includes: the Shannon entropy value of the sample data, the repetition rate of the sample data, the Shannon entropy value of the first half of the sample data, the Shannon entropy value of the second half of the sample data, and the categories of various data compression algorithms; The optimal compressed codeword sequence determination unit is used to use the target data compression algorithm to compress the data to be compressed to obtain the optimal compressed codeword sequence.
10. A computer device, comprising: A memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the processor executes the computer program to implement the adaptive data compression method according to any one of claims 1-8.
Citation Information
Patent Citations
Machine learning based image compression setting reflecting user preferences
CN114080615A
Data compression method of ERP (Enterprise Resource Planning) management system
CN115269940A
Log data processing method and device and computer readable storage medium
CN117093557A
Database data compression method and storage device
CN118202339A
Adaptive data compression method, system, device and product for database
CN118838879A