An Adaptive Optimization Method for LZ Series Compression Algorithms

Through the adaptive optimization method of LZ77 sliding window dictionary and auxiliary memory dictionary, the problem of low compression efficiency of LZ algorithm is solved. Through partition compression and adaptive update, the data compression efficiency and the utilization efficiency of auxiliary memory dictionary are improved.

CN115361026BActive Publication Date: 2025-07-04ZHENGZHOU UNIVERSITY OF AERONAUTICS
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211021912.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-08-24
Publication Date
2025-07-04
Estimated Expiration
2042-08-24

AI Technical Summary

Technical Problem

During the compression process, the compression efficiency of the LZ algorithm decreases due to the distance between repeated data exceeding the length of the sliding window. Increasing the dictionary length of the sliding window will increase the search time, resulting in low compression efficiency.

Method used

Adaptive optimization methods of LZ77 sliding window dictionary and auxiliary memory dictionary are adopted. By dividing the data to be compressed into multiple partitions, the initial partition is compressed and label values ​​are established using the LZ77 sliding window dictionary, high-frequency long statements are selected to build auxiliary memory dictionary, and adaptive updates and attenuation management are performed during the compression process.

Benefits of technology

Improve compression efficiency, reduce the amount of data in the auxiliary memory dictionary, enhance adaptability and compression value, and optimize the compression process.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115361026B_ABST
    Figure CN115361026B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of data compression technology. By obtaining the data to be compressed, dividing the data to be compressed into multiple partitions according to the length of the LZ77 sliding window, compressing the statements in the initial partition using the LZ77 sliding window dictionary, establishing a tag value for the compressed statements and incrementing the tag value of the statement by one each time compression is performed, excluding the statements with shorter lengths among the statements with the same data in the initially compressed partition and retaining the statements with longer lengths, and updating the tags of the retained statements, screening out the statements that meet the criteria to obtain an auxiliary memory dictionary, performing parallel auxiliary compression on each of the other partitions using the LZ77 sliding window dictionary, updating the auxiliary memory dictionary each time a statement is compressed, obtaining the attenuation value of each statement in the auxiliary memory dictionary, and deleting the statements with attenuation values less than the attenuation threshold, the method realizes the adaptive update of the auxiliary memory dictionary and improves the compression efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data compression, and particularly relates to an adaptive optimization method for LZ series compression algorithms. Background Art

[0002] Nowadays, with the rapid development of technology and the increasing popularity of the Internet, people have various ways to obtain information. Whether the information is obtained from the Internet, mobile devices, terminal devices, etc., data transmission and storage are required. To improve the data transmission performance, data compression is often required before data transmission, and lossy / lossless compression algorithms are used to reduce the total amount of data to be transmitted.

[0003] The LZ algorithm is the most commonly used lossless compression algorithm. The LZ algorithm builds a dictionary with a sliding window, including a data area to be compressed and a buffer data area, and retrieves and matches the data in the buffer data area in the buffer data area for matching compression.

[0004] However, in the LZ sliding window dictionary of the LZ algorithm, during the process of compressing data, the interval distance between many repeated data often exceeds the length of a sliding window, resulting in the inability to match and compress the repeated data, reducing the compression efficiency. The conventional solution is to increase a fixed length value for the LZ sliding window dictionary, so that the dictionary established by the sliding window can contain more data. However, increasing the length will increase the retrieval time of the dictionary, resulting in low compression efficiency. Summary of the Invention

[0005] The present invention provides an adaptive optimization method for LZ series compression algorithms to solve the problem of low compression efficiency of the LZ algorithm, and adopts the following technical solutions:

[0006] Obtain the data to be compressed;

[0007] First, obtain the data to be compressed with the length of an LZ77 sliding window dictionary, then increase the length of the LZ77 sliding window dictionary by one each time, and obtain the data repeatability within each length according to the probability of each data appearing within each length. If the data repeatability within the increased length is less than the data repeatability within the length before the increase, stop increasing the length, and use the data to be compressed within the length before the increase as the initial partition;

[0008] Use the LZ77 sliding window dictionary to compress the data in the initial partition, and regard each compressed data as a statement. When each statement is compressed for the first time, establish a tag value for the statement and initialize it. When compressing the same statement as this statement, increase the tag value of this statement by one until the initial partition is compressed;

[0009] Obtain statements with the same data in the initial partition, retain the statement with the longest length among the statements with the same data, exclude the remaining statements, and use the sum of the tag values of the excluded statements and the retained statement as the tag value of the retained statement;

[0010] Judge whether each retained statement in the initial partition meets the entry criteria of the auxiliary memory dictionary according to the tag value and length of the retained statement, and initialize the auxiliary memory dictionary with the retained statements that meet the entry criteria;

[0011] Obtain each partition other than the initial partition, and retrieve and match the data in each other partition in the auxiliary memory dictionary and the LZ77 sliding window dictionary;

[0012] If it can be matched only in the LZ77 sliding window dictionary, use the LZ77 sliding window dictionary for compression. If it can be matched only in the auxiliary memory dictionary, compress it with the auxiliary memory dictionary. If it can be matched in both the auxiliary memory dictionary and the LZ77 sliding window dictionary, compress it with the LZ77 sliding window dictionary;

[0013] Regardless of whether it is compressed with the LZ77 sliding window dictionary or the auxiliary memory dictionary, for each compressed statement, retrieve the statement with the same data as this statement in the auxiliary memory dictionary, and replace the statement with the shortest length among this statement and the statement with the same data with the statement with the longest length for adaptive update.

[0014] The method for performing adaptive update is as follows:

[0015] Whether it is compressing a statement using the LZ77 sliding window or using the auxiliary memory dictionary;

[0016] Retrieve this statement in the auxiliary memory dictionary. If it can be retrieved, increment the tag value of this statement by one;

[0017] If this statement cannot be retrieved, retrieve the statement with the same data as this statement;

[0018] If the statement with the same data as this statement cannot be retrieved, establish and initialize a tag value for this statement;

[0019] If the statement with the same data as this statement is retrieved, compare the length of this statement and the length of the statement with the same data as this statement;

[0020] If the length of this statement is greater than the length of the statement with the same data as this statement, replace the statement with the same data as this statement with this statement, and this statement inherits the tag value of the statement with the same data as this statement and increments it by one;

[0021] If the length of this statement is less than the length of the statement with the same data, no replacement is performed, and only the label value of the statement with the same data is incremented by one.

[0022] The adaptive update further includes that when the length of the statement stored in the auxiliary memory dictionary is greater than or equal to the LZ77 sliding window, the decay value of each statement stored in the auxiliary memory dictionary is calculated according to the length, label value, and time interval between the last compression time and the current time of each statement, and the statements with decay values less than the decay value threshold are deleted.

[0023] The method for calculating the decay value of each statement stored in the auxiliary memory dictionary according to the length, label value, and time interval between the last compression time and the current time of each statement is as follows:

[0024]

[0025] In the formula, G i is the decay value of the i-th statement, e is the natural constant, and E i is the time interval between the last compression time and the current time of the i-th statement in the auxiliary memory dictionary, and m i is the length of the i-th statement, and F i is the label value of the i-th statement.

[0026] The method for obtaining the decay threshold is as follows:

[0027] Obtain the maximum decay value and the minimum decay value of all statements stored in the auxiliary memory dictionary;

[0028] Obtain the difference between the maximum decay value and the maximum decay value, and the value obtained by dividing the difference by the adjustment parameter is the decay threshold, and the adjustment parameter is set by oneself.

[0029] The statements with the same data refer to that if one statement contains data that can cover another statement among two statements, then the two statements are statements with the same data.

[0030] The method for obtaining each partition other than the initial partition is the same as the method for obtaining the initial partition.

[0031] The method for determining whether the retained statement meets the entry criteria of the auxiliary memory dictionary according to the label value and length of each retained statement in the initial partition is as follows:

[0032] Obtain the product C1 of the length of each retained statement in the initial partition and the label value of this statement;

[0033] Obtain the product C2 of the length of each retained statement in the initial partition and the average value of the label value of this statement;

[0034] If the difference between C1 and C2 is greater than 0, then the retention statement meets the entry criteria for the auxiliary memory dictionary;

[0035] If the difference between C1 and C2 is greater than 0 and less than or equal to 0, then the retention statement does not meet the entry criteria for the auxiliary memory dictionary.

[0036] The beneficial effects of the present invention are:

[0037] (1) Using the length of the LZ77 sliding window dictionary to divide the data to be compressed into multiple partitions, compressing the statements in the initial partition with the LZ77 sliding window dictionary, setting a tag value for the statement, and excluding and retaining statements with the same data; the method not only ensures the accuracy of the statements for constructing the auxiliary memory dictionary, but also reduces the amount of data for constructing the auxiliary memory;

[0038] (2) Judging whether the statement meets the entry criteria for the auxiliary memory dictionary according to the tag value and length of each statement, and obtaining the auxiliary memory dictionary according to the statements that meet the entry criteria for the auxiliary memory dictionary; the method screens out high-frequency and long-length statements to construct the auxiliary memory dictionary. This method does not increase the sliding window length of the LZ77 dictionary, but by establishing an auxiliary memory dictionary, extracts long-length statements, that is, improves the compression value of the statements in the auxiliary memory dictionary;

[0039] (3) Using the LZ77 sliding window dictionary to perform parallel auxiliary compression on each other partition. For each compressed statement, retrieve whether there is a related statement in the auxiliary memory dictionary, and update the auxiliary memory dictionary according to the retrieval result; this method adaptively replaces and updates high-frequency statements in the auxiliary memory dictionary during the compression process, improves the adaptability of the auxiliary memory dictionary, and improves the compression efficiency;

[0040] (4) When the total length of the statements stored in the auxiliary memory dictionary is greater than or equal to the length of the LZ77 sliding window dictionary, obtain the attenuation value of each statement according to the length, tag value of each statement, and the length of the statements between the last compression and the current moment, and delete the statements with an attenuation value less than the attenuation threshold; this method uses the attenuation function to delete the statements in the auxiliary memory dictionary, reduces the bloat of the auxiliary memory dictionary, is a further optimization of the compression method, and further improves the compression efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0042] Figure 1 It is a schematic flow diagram of an adaptive optimization method for the LZ series compression algorithm of the present invention. Specific implementation manners

[0043] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without making creative efforts shall fall within the protection scope of the present invention.

[0044] An embodiment of an adaptive optimization method for the LZ series compression algorithm of the present invention is as Figure 1 shown and includes:

[0045] Step 1: Obtain the data to be compressed; first obtain the data to be compressed with the length of an LZ77 sliding window dictionary, then increase the length of the LZ77 sliding window dictionary by one each time, and obtain the data repeatability within each length according to the probability of each data appearance within each length. If the data repeatability within the increased length is less than that within the length before the increase, stop increasing the length, and use the data to be compressed within the length before the increase as the initial partition.

[0046] The purpose of this step is to utilize the repeatability between the data to be compressed and perform interval partitioning on the data to be compressed based on the length of the LZ77 window dictionary.

[0047] Among them, the data to be compressed obtained in the present invention is character data.

[0048] The scenario targeted by the present invention is: when using the LZ77 algorithm for data compression, a sliding window is used as an automatic dictionary to compress the data. However, during the compression process, many repeated information cannot be compressed because the interval distance exceeds a sliding window, resulting in a reduction in compression efficiency. Therefore, the present invention realizes the purpose of high-efficiency compression without increasing the length of the dynamic window dictionary in the LZ77 algorithm by establishing an auxiliary memory dictionary for parallel auxiliary compression on the basis of using the LZ77 algorithm to compress information with a dynamic window dictionary.

[0049] Among them, the method for obtaining the initial partition is:

[0050] First obtain the data to be compressed with the length of an LZ77 sliding window dictionary, then increase the length of the LZ77 sliding window dictionary by one each time, and obtain the data repeatability within each length according to the probability of each data appearance within each length. If the data repeatability within the increased length is less than that within the length before the increase, stop increasing the length, and use the data to be compressed within the length before the increase as the initial partition.

[0051] The specific method is as follows:

[0052] (1) First, starting from the first data of the data to be compressed, obtain the data to be compressed within the length range of one LZ77 sliding window dictionary, and calculate the data repeatability within this length:

[0053]

[0054] In the formula, represents the repeatability of the data within the length range of one LZ77 sliding window dictionary when the data is used as the first partition (initial partition). The subscript 1 represents the data to be compressed within the length range of one LZ77 sliding window dictionary, the superscript 1 represents the initial partition, e is the natural constant, a is the a-th data within the length range of one LZ77 sliding window dictionary, representing an independent (non-repeating) data, l represents the length of one LZ77 sliding window dictionary, that is, the total number of data in this partition, P a represents the probability of each independent data appearing in this interval, P a is the probability of data a appearing in this partition, log2P a is the logarithmic function;

[0055] The purpose of this formula: Quantify the data repetition rate under this partition through the probability of each independent data appearing. Although the information compression in the LZ77 compression process is not completely targeted at data information compression, the repeated appearance of long sentences is based on data repetition. For example, if you want the sentence AB to repeat, the first basic condition is the repetition of data A. If A appears only once in the entire data to be compressed, then any length of sentence starting with A cannot appear. Therefore, predict long sentences through the probability of data. The larger P a is repeated, the larger the value, the greater the probability of sentence repetition regardless of length.

[0056] (2) Increase the length of one LZ77 sliding window dictionary. At this time, obtain the data to be compressed within the length range of two LZ77 sliding window dictionaries, and calculate the data repeatability within this length. The method is the same as (1):

[0057] Take the data to be compressed within the length range of two LZ77 sliding window dictionaries as the initial partition, and calculate the data repeatability of this partition

[0058]

[0059] In the formula, The meaning is to use the data to be compressed within the length range of 2 LZ77 sliding window dictionaries as the first partition (initial partition);

[0060] (3) When using the data to be compressed within the length range of 1 LZ77 sliding window dictionary as the initial partition, the data repeatability of this partition and when using the data to be compressed within the length range of 2 LZ77 sliding window dictionaries as the initial partition, the data repeatability of this partition are compared:

[0061] If it indicates that the effect of initial partitioning with the dictionary length of 2 LZ77 sliding windows is better than that with the dictionary length of 1 LZ77 sliding window;

[0062] (4) Continuously increase the length range and calculate the data repeatability within each length. When the data repeatability within the increased length is less than that within the previous length, stop increasing the length, and use the data to be compressed within the previous length as the initial partition. If when increasing to the (n + 1)-th LZ77 sliding window dictionary length, the data repeatability within the (n + 1)-th LZ77 sliding window dictionary length is less than or equal to the data repeatability within the n-th LZ77 sliding window dictionary length it indicates that the repetition rate of the data to be compressed within the first n sliding window dictionary lengths is the highest. Therefore, select the data to be compressed within the first n LZ77 sliding window dictionary lengths as the initial data area;

[0063] (5) The method for obtaining each partition other than the initial partition is the same as that for obtaining the initial partition: that is, for the data to be compressed other than the initial partition, through the methods of (1) to (4), each partition is obtained.

[0064] So far, all the divided data partitions are obtained, and the probability of any length of statements repeating in all data partitions is the highest.

[0065] This step uses the characteristics (repeatability) of the information to be compressed and the length of the sliding window in the LZ77 algorithm to partition the information to be compressed, and then establishes an initial auxiliary memory dictionary based on the compression effect of the first partition.

[0066] It should be noted that the establishment and update process of the auxiliary memory dictionary described in the present invention is completed by using high-frequency compressed information based on the LZ77 algorithm sliding window dictionary for information compression, and the memory auxiliary dictionary needs to be adaptively updated. Therefore, this step is required to perform quantization partitioning on the overall data to be compressed to achieve the purpose of maximizing the compression efficiency and minimizing the resource utilization of the auxiliary memory dictionary.

[0067] Step 2: Use the LZ77 sliding window dictionary to compress the data in the initial partition. Take the data compressed each time as a statement. When each statement is compressed for the first time, establish a tag value for this statement and initialize it. When compressing a statement identical to this statement, increment the tag value of this statement by one until the compression of the initial partition is completed; Obtain the statements with the same data in the initial partition, retain the statement with the longest length among the statements with the same data, exclude the remaining statements, and use the sum of the tag values of the excluded statements and the retained statement as the tag value of the retained statement;

[0068] The purpose of this step is to compress the statements in the initial partition using the LZ77 sliding window dictionary, establish a tag value for each statement, and count the compression times of this statement.

[0069] Among them, statements with the same data refer to that if in two statements, the data contained in one statement can cover the other statement, then these two statements are statements with the same data. For example, in ABC and BC, ABC can cover BC because ABC itself contains BC.

[0070] Among them, the method of using the LZ77 sliding window dictionary to compress the data in the initial partition is as follows:

[0071] Take the data compressed each time as a statement. When each statement is compressed for the first time, establish a tag value for this statement and initialize it. When compressing a statement identical to this statement, increment the tag value of this statement by one until the compression of the initial partition is completed;

[0072] Since the information area is partitioned, there is a repetition rate of the statement lengths within each information area itself, which means that the occurrence times of repeated statements in two different information areas are not many. Therefore, use the first information area to initialize the auxiliary memory dictionary. Specifically:

[0073] First, when the LZ77 sliding window dictionary slides in R 1 During the compression process, R 1 is the initial partition. For each independent statement compressed, establish a tag F i , F i is the tag of the i-th statement, and the initial value of the tag is 1. Then, each time this statement is compressed, increment the corresponding statement tag by one until R 1 is compressed completely;

[0074] Then, exclude the already compressed data recorded by the tag:

[0075] Obtain the statements with the same data in the initial partition, retain the statement with the longest length among the statements with the same data, exclude the remaining statements, and use the sum of the tag values of the excluded statements and the tag value of the retained statement as the tag value of the retained statement. The specific method is as follows:

[0076] Exclude the already compressed data recorded by the tags. Statements with the same data are statements with the same data. Use the long statement to exclude the short statement. The exclusion method is that for the same data, that is, statements with the same data, they can be retrieved during the compression process. For example, for statement ABC and statement ABCD, their common feature is ABC. When retrieving statements in the dictionary, identify the statements with the same data and exclude them. At the same time, retain the long statement with the same data and add the tag value corresponding to the excluded short statement to the tag value corresponding to the long statement. For example, for one statement ABC, the corresponding tag value F1 = 15, and the long statement with the same data is ABCD, the corresponding tag value F2 = 19. Then exclude the short statement ABC, retain the long statement ABCD, and reset the tag value of the long statement to F2 = 15 + 19 = 34. The principle is that the long statement ABCD can completely compress the short statement ABC, but the short statement ABC cannot completely compress the long statement ABCD;

[0077] Finally, obtain the statements with established tag values in the already compressed initial partition according to this step.

[0078] Step 3: Determine whether each retained statement in the initial partition meets the entry criteria of the auxiliary memory dictionary according to the tag value and length of the retained statement, and initialize the auxiliary memory dictionary with the retained statements that meet the entry criteria;

[0079] This step is to screen the already compressed statements with established tag values, screen out the high-frequency statements among them for initializing the entry of the auxiliary memory dictionary, and construct the auxiliary memory dictionary.

[0080] Among them, the method for determining whether each retained statement meets the entry criteria of the auxiliary memory dictionary according to the tag value and length of the statement is as follows:

[0081] (1) Obtain the product C1 of the length of each retained statement in the initial partition and the tag value of the statement:

[0082] C1 = m i ×F i

[0083] In the formula, m i is the length of the i-th statement, and F i is the tag value of the i-th statement;

[0084] (2) Obtain the product C2 of the length of each statement in the initial partition and the average value of the statement label values:

[0085]

[0086] In the formula, m i is the length of the i-th statement, F i is the label value of the i-th statement, I1 is the total number of statements in the current data area, i is the i-th statement, is the average value of the statement label values;

[0087] (3) If the difference between C1 and C2 is greater than 0, then this statement meets the entry criteria for the auxiliary memory dictionary;

[0088] (4) If the difference between C1 and C2 is greater than 0 and less than or equal to 0, then this statement does not meet the entry criteria for the auxiliary memory dictionary.

[0089] Specifically as follows:

[0090] Obtain the product C1 of the length of each retained statement and the label value of each retained statement; obtain the product C2 of the length of each statement and the average value of the statement label values; subtract the product of the length of each statement and the average value of the statement label values from the product of the length of each statement and the label value of each statement to obtain E i :

[0091] E i = C1 - C2

[0092] In the formula, E i means whether the i-th statement meets the entry criteria for the auxiliary memory dictionary. If the difference value E i is greater than 0, then this statement meets the entry criteria for the auxiliary memory dictionary. If the difference value is less than or equal to 0, then this statement does not meet the entry criteria for the auxiliary memory dictionary.

[0093] Formula meaning: During the data compression process, for statements with longer lengths compared to those with shorter lengths, the improvement in compression efficiency is more obvious. Specifically, short statements cannot fully compress long statements, but long statements can fully compress short statements; and the label value indicates the number of times the statement is compressed. The larger the label value, the more times it is compressed. Therefore, the present invention uses the length of the statement as the weight and the label value of the statement as the basis to quantify the screening and entry criteria for this statement in this information area. Using this criteria and the average criteria of the entire marked data as the difference value to perform the screening for whether to be entered, the larger the standard value, the higher the possibility of being entered. When it is greater than the average value of all marked data in the entire interval, the present invention considers it to be commonly used (high frequency) and having compression value (long statement length), and it can be used as a dictionary statement in the auxiliary memory dictionary.

[0094] Using the above method for R 1 That is, screening all the tag value statements in the first data area, the first data area R in the initialized auxiliary memory dictionary can be obtained 1 There are I1' compressed statements in R, where I1' is the statement reserved for the first data area

[0095] Furthermore, an auxiliary memory dictionary is obtained according to the statements that meet the entry criteria of the auxiliary memory dictionary. The statements that meet the entry criteria of the auxiliary memory dictionary are sequentially entered into the auxiliary memory dictionary to obtain the auxiliary memory dictionary

[0096] Step 4: Obtain each partition other than the initial partition, and retrieve and match the data in each other partition in the auxiliary memory dictionary and the LZ77 sliding window dictionary. If it can be matched only in the LZ77 sliding window dictionary, then use the LZ77 sliding window dictionary for compression. If it can be matched only in the auxiliary memory dictionary, then use the auxiliary memory dictionary for compression. If it can be matched in both the auxiliary memory dictionary and the LZ77 sliding window dictionary, then use the LZ77 sliding window dictionary for compression; Regardless of whether it is compressed by the LZ77 sliding window dictionary or the auxiliary memory dictionary, for each compressed statement, retrieve the statement with the same data as this statement in the auxiliary memory dictionary, and replace the statement with the shortest length with the statement with the longest length among this statement and the statement with the same data, for adaptive update

[0097] The purpose of this step is to perform parallel auxiliary compression on the statements in each other partition using the auxiliary memory dictionary and the LZ77 dictionary, and perform adaptive update on the auxiliary memory dictionary

[0098] Among them, the method for adaptive update is

[0099] Whether it is using LZ77 sliding window compression or using the auxiliary memory dictionary to compress a statement

[0100] Retrieve this statement in the auxiliary memory dictionary. If it can be retrieved, then increment the tag value of this statement by one

[0101] If this statement cannot be retrieved, then retrieve the statement with the same data as this statement

[0102] If the statement with the same data as this statement cannot be retrieved, then establish a tag value for this statement and initialize it

[0103] If the statement with the same data as this statement is retrieved, then compare the length of this statement and the length of the statement with the same data as this statement

[0104] If the length of this statement is greater than the length of the statement with the same data, replace the statement with the same data with this statement. This statement inherits the tag value of the statement with the same data and increments it by one.

[0105] If the length of this statement is less than the length of the statement with the same data, no replacement is made. Only the tag value of the statement with the same data is incremented by one.

[0106] Specifically, taking the second partition as an example:

[0107] Because the auxiliary memory dictionary has been initialized and established using R 1 Now starting from R 2 That is, starting from the second partition, replace and update the statements in the auxiliary memory dictionary according to the actual compression effect. The specific logic is that whether it is the sliding window dictionary in the LZ77 dictionary or the auxiliary memory dictionary, after compressing all the information after R 2 starts, it is necessary to retrieve in the auxiliary memory dictionary and update according to the retrieval result. When the data is the same and the length is greater than the statement in the auxiliary memory dictionary, replace it. If there is no statement exclusive to R 2 in the auxiliary memory dictionary, establish the tag value and input it in the same way as in step two. The implementation process is as follows:

[0108] First, use the LZ77 sliding window dictionary and the auxiliary memory dictionary to parallelly retrieve and compress all the data after R 2 in units of each statement. Taking R 2 as an example, retrieve and compress the information in R 2 using the LZ77 sliding window dictionary and the auxiliary memory dictionary.

[0109] Then perform pre - processing before inputting the statements in R 2 . The processing method is that for each statement compression in R 2 , first retrieve in the auxiliary memory dictionary whether there is a corresponding characteristic statement. If it exists, determine whether its statement length is greater than the length of the statement with the same data in the auxiliary dictionary. If it is greater than the length of the statement with the same data in the auxiliary dictionary, replace the statement with the same data in the auxiliary dictionary with this statement and inherit its tag value and increment it by 1. If it is not greater, no replacement is made, and only the tag value is incremented by one, and no statement replacement is performed.

[0110] Finally, when performing a statement compression in R 2 and no corresponding characteristic statement is retrieved in the auxiliary dictionary, use the method in step two to establish the statement tag value belonging to R 2 and initialize it. The adaptive replacement and update method of this step is common to all partitions.

[0111] Among them, the adaptive update further includes that when the length of the statement stored in the auxiliary memory dictionary is greater than or equal to the LZ77 sliding window, the decay value of each statement stored in the auxiliary memory dictionary is calculated according to the length of each statement, the tag value, and the time interval between the last compression time and the current time, and the statements with decay values less than the decay value threshold are deleted;

[0112] Among them, the method for calculating the decay value of each statement stored in the auxiliary memory dictionary according to the length of each statement, the tag value, and the time interval between the last compression time and the current time is as follows:

[0113]

[0114] In the formula, G i is the decay value of the i-th statement, e is the natural constant, and E i is the time interval between the last compression time and the current time of the i-th statement in the auxiliary memory dictionary. In this embodiment, the time is characterized by the data length of the interval between the last compression of this statement and the current compression time, and m i is the length of the i-th statement, and F i is the tag value of the i-th statement.

[0115] The purpose of this formula is to use the uncompressed duration of the statement in the auxiliary memory dictionary (quantified by the information length of the interval between the last compression of this statement and the calculation of the decay function), the statement length, and the tag value as parameters to set the memory decay function to discard the statements in the auxiliary memory dictionary in a certain period to achieve the effect of reducing the bloatedness of the auxiliary memory dictionary and improving the compression efficiency.

[0116] It should be noted that using the decay value to decay the statements in the auxiliary memory dictionary has the beneficial effect of giving each statement in the auxiliary memory dictionary a certain "existence tolerance space", that is, the initial decay amplitude is not large, and as the parameters change, the decay degree becomes larger and larger. The actual physical meaning is that as the uncompressed duration increases, the length of the i-th statement does not change (no compression means no update), and the tag value of the length of the i-th statement does not change (no compression means no change), and the decay process of its decay function becomes faster and faster until the decay ends.

[0117] Among them, the method for obtaining the decay threshold is: obtaining the maximum decay value and the minimum decay value of all statements stored in the auxiliary memory dictionary; obtaining the difference between the maximum decay value and the maximum decay value, and dividing the obtained difference by the adjustment parameter to obtain the decay threshold, and the adjustment parameter is set by itself.

[0118] The specific formula is:

[0119] Set a threshold K, and discard the corresponding statements whose decay function values are less than the threshold K. A calculation method for setting a threshold is as follows:

[0120]

[0121] In the formula, max{G i} is the maximum decay value, min{G i} is the minimum decay value, γ is an adjustment parameter, which can be adjusted according to the implementation requirements of the implementer. In this embodiment, γ = 0.5.

[0122] So far, the memory decay function is set.

[0123] It should be noted that during the process of replacing and updating the auxiliary memory dictionary, due to the continuous increase of the information area, it is easy to cause the auxiliary memory dictionary to be bloated (too many statements), so that the retrieval time is too long when using the auxiliary memory dictionary for auxiliary compression. Therefore, the present invention uses the uncompressed duration of the statements in the auxiliary memory dictionary (quantified by the information length between the last compression and the calculation of the decay function), the statement length, and the tag value as parameters to set the memory decay function to discard the statements in the auxiliary memory dictionary at a certain period to achieve the effect of reducing the bloat of the auxiliary memory dictionary. The certain period is when using the auxiliary memory dictionary to assist the LZ77 dynamic window dictionary for the R r th information area compression, when the total length l' of the statements stored in the auxiliary memory dictionary is greater than or equal to the LZ77 dynamic window dictionary l, the decay function is used for decay, and some statements (statements reaching the threshold) stored in the auxiliary memory dictionary are discarded;

[0124] The present invention first partitions the information to be compressed by using the characteristics (repeatability) of the information to be compressed and the length L of the sliding window in the LZ77 algorithm, then establishes an initial auxiliary memory dictionary through the compression effect of the first partition, and then enters, replaces, and discards the high-frequency statements in the initial memory dictionary through the compression effect of each partition. During this process, the auxiliary memory dictionary is simultaneously used to assist the LZ77 sliding window dictionary for information compression.

[0125] It should be noted that for the conventional lz77 algorithm during the process of information compression, because the dictionary used for compression is a dynamic dictionary, during the compression process, some information is the same, but due to the long interval distance between the same information, exceeding the length of the dynamic compression dictionary, the same information cannot be compressed, which has a greater impact on the compression efficiency of the information.

[0126] Through Steps 1 to 4, the present invention has completely established the auxiliary memory dictionary. Now, the auxiliary memory dictionary is used to assist the LZ77 sliding window dictionary to compress the information to be compressed. The specific method is parallel auxiliary compression, that is, when compressing, the information to be compressed is simultaneously retrieved and matched for the compression length in the copy memory dictionary and the LZ77 sliding window dictionary. If it can be retrieved simultaneously, the LZ77 sliding window dictionary is used for compression. If it can only be retrieved in the auxiliary memory dictionary, the auxiliary memory dictionary is used for compression, and the already compressed data is transmitted and stored.

[0127] Based on the lz77 algorithm for information compression, the present invention establishes an auxiliary adaptive and automatically updated auxiliary memory dictionary through the frequently compressed data of the already compressed information, the dictionary length of the LZ77 algorithm, and the characteristics of the data to be compressed. Then, through the auxiliary dynamic dictionary to assist the dynamic window dictionary, parallel auxiliary compression of the data is carried out without increasing the dictionary length to improve the compression efficiency.

[0128] Further, the present embodiment is illustrated by way of example:

[0129] (1) Use the LZ77 sliding window dictionary to compress the initial partition data: The character data to be compressed in the initial partition is: ABABCDABCDBCE. Set the length of the LZ77 sliding window dictionary to 8 bits 00000000, and use 0 to represent the empty position where there is no data. It includes 3 bits in the data area to be compressed and 5 bits in the buffer data area;

[0130] Basic compression rule: When the statement in the data area to be compressed is not retrieved and matched in the buffer data area, the unmatched symbol is encoded into a symbol mark, and this symbol mark only contains the symbol itself without a compression process.

[0131] When the statement in the data area to be compressed is retrieved and matched in the buffer data area, the matched statement is compressed into (the offset in the sliding window, the match length, the next data to be compressed after the match ends).

[0132] Initial state:

[0133] Since the 3 bits 000 in the data area to be compressed and the 5 bits 00000 in the buffer data area are both empty, the LZ77 sliding window dictionary is slid 3 bits to the right starting from the character data to be compressed. At this time, the data area to be compressed contains the data ABA as the initial state;

[0134] The compression process is as follows:

[0135] a. The first character A in the data area to be compressed is not retrieved and matched in the buffer data area. A is not compressed, and A is output. At this time, the buffer data area is 0000A, and the data area to be compressed is BAB;

[0136] b. The first character B in the data area to be compressed is not found in the buffered data area. Output B without compression. At this time, the buffered data area is 000AB and the data area to be compressed is ABC;

[0137] c. The first character A in the data area to be compressed is found in the buffered data area. Then continue to search for the characters AB in the buffered data area. AB is found, then continue to search for ABC. ABC is not found, so output (3, 2, C), only compress AB. AB is compressed for the first time, establish a tag value and initialize it to 1. At this time, the buffered data area is 0ABAB and the data area to be compressed is CDA;

[0138] d. The first character C in the data area to be compressed is not found in the buffered data area, so output C. At this time, the buffered data area is ABABC and the data area to be compressed is DAB;

[0139] e. The first character D in the data area to be compressed is not found in the buffered data area, so output D. At this time, the buffered data area is BABCD and the data area to be compressed is ABC;

[0140] f. The first character A in the data area to be compressed is found in the buffered data area. Continue to search for AB, which is found. Then continue to search for ABC, which is also found. So compress ABC and output (1, 3, D). ABC is compressed for the first time, establish a tag value and initialize it to 1. At this time, the buffered data area is CDABC and the data area to be compressed is DBC;

[0141] g. The first character D in the data area to be compressed is found in the buffered data area. Continue to search for DB, which is not found. So only compress D and output (1, 1, B). D is compressed for the first time, establish a tag value and initialize it to 1. At this time, the buffered data area is DABCD and the data area to be compressed is BCE;

[0142] h. The first character B in the data area to be compressed is found in the buffered data area. Continue to search for BC, which is found. Then continue to search for BCE, which is not found. So only compress BC and output (2, 2, E). BC is compressed for the first time, establish a tag value and initialize it to 1. At this time, the buffered data area is BCDBC and the data area to be compressed is E;

[0143] i. The character E in the data area to be compressed is not found in the buffered data area, so output E. At this time, the buffered data area is CDBCE and the data area to be compressed is empty, and the compression is completed;

[0144] The compressed data obtained at this time is:

[0145] AB(3, 2, C)CD(1, 3, D)(1, 1, B)(2, 2, E)E

[0146] (2) Obtain the compressed statements and their tag values obtained from the compression process:

[0147] AB = 1, ABC = 1, D = 1, BC = 1. Exclude the short statements with the same data, retain the long statements and update the tag values:

[0148] AB, BC, and ABC are statements with the same data. Exclude AB and BC and retain ABC. And modify the tag value of ABC to the sum of its own tag value and the tag values of AB and BC. ABC = 3; D has no statements with the same data, so it is directly retained;

[0149] (2) Initialize the auxiliary memory dictionary:

[0150] Select the statements that meet the input criteria according to the length and tag values of the retained statements in the initial partition. Assume that ABC meets the input criteria, then initialize the auxiliary memory dictionary using ABC. Assume the auxiliary memory dictionary is also 8 bits, including a 3-bit compressed data area and a 5-bit data buffer area. At this time, the buffer area of the initialized auxiliary memory dictionary is 00ABC, and the buffer area of the LZ77 sliding window dictionary finally obtained in step (1) is CDBCE;

[0151] (3) Perform parallel auxiliary compression using the LZ77 dictionary and the auxiliary memory dictionary:

[0152] a. If the data in a certain partition is ADBAB and the data area to be compressed is ADB at this time, first search for A in the LZ77 dictionary and the auxiliary memory dictionary. Only the auxiliary memory dictionary can be searched and matched, so compress A using the auxiliary memory dictionary;

[0153] b. After A is compressed, first search in the auxiliary memory dictionary to see if there is a statement with the same data as statement A. The statement ABC with the same data is found. Then judge whether its statement length is greater than the length of the statement with the same data in the auxiliary dictionary. It is found that it is less than the length of the statement with the same data in the auxiliary dictionary, so no replacement is performed, and only the tag value of ABC in the auxiliary dictionary is incremented by one. At this time, the value is 3, and no statement replacement is performed;

[0154] c. After the auxiliary memory dictionary is compressed, its data buffer area is 0ABCA, and the data area to be compressed is DBA;

[0155] d. Search for D in the LZ77 dictionary and the auxiliary memory dictionary. Only the LZ77 sliding window dictionary can be searched and matched. D is retrieved, and DB is also retrieved. At this time, compress DB. At this time, the data in its buffer area is BCEDB, and the data area to be compressed is BAB;

[0156] e. After the LZ77 dictionary is compressed, first search in the auxiliary memory dictionary to see if there is a statement in the statement DB with the same data. If no statement with the same data is found, a tag value is established for DB and initialized to 1.

[0157] f. Search for B in the LZ77 dictionary and the auxiliary memory dictionary. If it can be found in both, then use the LZ77 dictionary to compress B. The data buffer is CEDBB and the data area to be compressed is A.

[0158] g. After the LZ77 dictionary is compressed, first search in the auxiliary memory dictionary to see if there is a statement in the statement B with the same data. If a statement with the same data ABC is found, then compare the statement lengths. If the length of ABC is larger, no replacement is made and the tag value of ABC is incremented. At this time, the value is 4.

[0159] h. Search for A in the LZ77 dictionary and the auxiliary memory dictionary. If it can only be found in the auxiliary memory dictionary, then use the auxiliary memory dictionary to compress it. After compression, the data buffer of the auxiliary memory dictionary is ABCAA and the data area to be compressed is empty. The compression is complete.

[0160] i. After the compression of the auxiliary memory dictionary is complete, first search in the auxiliary memory dictionary to see if there is a statement in the statement A with the same data. If a statement with the same data ABC is found, then determine whether its statement length is greater than the length of the statement with the same data in the auxiliary dictionary. If it is found to be smaller than the length of the statement with the same data in the auxiliary dictionary, no replacement is made and only the tag value of ABC in the auxiliary dictionary is incremented. At this time, the tag value of ABC is 5.

[0161] Perform parallel auxiliary compression on the data of each other partition according to the method in (3).

[0162] The above is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.

Claims

1. An adaptive optimization method for the LZ series compression algorithm, characterized in that, Including: Obtain the data to be compressed; First, obtain the data to be compressed with the length of the LZ77 sliding window dictionary. Then, each time increase the length of the LZ77 sliding window dictionary by one, and obtain the data repeatability within each length according to the probability of each data occurrence within each length. When the data repeatability within the increased length is less than that within the length before the increase, stop increasing the length, and use the data to be compressed within the length before the increase as the initial partition; Use the LZ77 sliding window dictionary to compress the data in the initial partition. Take the data compressed each time as a statement. When each statement is compressed for the first time, establish a tag value for the statement and initialize it. When compressing a statement identical to this statement, increase the tag value of this statement by one until the compression of the initial partition is completed; Obtain the statements in the initial partition with the same data. Retain the statement with the longest length among the statements with the same data, exclude the remaining statements, and use the sum of the tag values of the excluded statements and the retained statement as the tag value of the retained statement; Judge whether each retained statement in the initial partition meets the entry criteria of the auxiliary memory dictionary according to the tag value and length of each retained statement, and initialize the auxiliary memory dictionary with the retained statements that meet the entry criteria; Obtain each partition other than the initial partition, and retrieve and match the data in each other partition in the auxiliary memory dictionary and the LZ77 sliding window dictionary; If it can be matched only in the LZ77 sliding window dictionary, use the LZ77 sliding window dictionary for compression. If it can be matched only in the auxiliary memory dictionary, use the auxiliary memory dictionary for compression. If it can be matched in both the auxiliary memory dictionary and the LZ77 sliding window dictionary, use the LZ77 sliding window dictionary for compression; Regardless of whether it is compressed by the LZ77 sliding window dictionary or the auxiliary memory dictionary, for each statement compressed, retrieve in the auxiliary memory dictionary the statement with the same data as this statement, and replace the statement with the shortest length with the statement with the longest length among this statement and the statement with the same data for adaptive update; The said adaptive update further includes that when the length of the statement stored in the auxiliary memory dictionary is greater than or equal to the LZ77 sliding window, calculate the attenuation value of each statement stored in the auxiliary memory dictionary according to the length, tag value, and the time interval between the last compression time and the current time, and delete the statements with attenuation values less than the attenuation value threshold; 2. The adaptive optimization method for the LZ series compression algorithm according to claim 1, wherein The method for the said adaptive update is: Whether compressing a statement using the LZ77 sliding window or using the auxiliary memory dictionary; Retrieve this statement in the auxiliary memory dictionary. If it can be retrieved, increase the tag value of this statement by one; If this statement cannot be retrieved, retrieve the statement with the same data as this statement; If the statement with the same data as this statement cannot be retrieved, establish a tag value for this statement and initialize it; If the statement with the same data as this statement is retrieved, compare the length of this statement with the length of the statement with the same data as this statement; If the length of this statement is greater than the length of the statement with the same data, then replace the statement with the same data with this statement. This statement inherits the tag value of the statement with the same data and increments it by one; If the length of this statement is less than the length of the statement with the same data as this statement, no replacement is performed, and only the tag value of the statement with the same data is incremented by one.

3. An adaptive optimization method for the LZ series compression algorithm according to claim 1, characterized in that, The method for calculating the decay value of each statement stored in the auxiliary memory dictionary according to the length, tag value, and time interval between the last compression time and the current time of each statement is as follows: Wherein, is the attenuation value of the i-th statement, is the natural constant, is the time interval between the last compression time and the current time of the i-th statement in the auxiliary memory dictionary, is the length of the i-th statement, is the label value of the i-th statement.

4. A self - adaptive optimization method for the LZ series compression algorithm according to claim 1, characterized in that The method for obtaining the decay value threshold is as follows: Obtain the maximum decay value and the minimum decay value of all statements in the statements stored in the auxiliary memory dictionary; Obtain the difference between the maximum decay value and the maximum decay value, and divide the difference by the adjustment parameter. The value obtained is the decay value threshold, and the adjustment parameter is set by oneself.

5. The adaptive optimization method for the LZ series compression algorithm according to claim 1, wherein, The statements with the same data refer to that if one statement can overwrite another statement among two statements, then these two statements are statements with the same data.

6. A self-adaptive optimization method for the LZ series compression algorithm according to claim 1, characterized in that The method for obtaining each partition other than the initial partition is the same as the method for obtaining the initial partition.

7. A self-adaptive optimization method for the LZ series compression algorithm according to claim 1, characterized in that The method for determining whether the retained statement meets the entry criteria of the auxiliary memory dictionary according to the tag value and length of each retained statement in the initial partition is as follows: Obtain the product C1 of the length of each retained statement in the initial partition and the tag value of this statement; Obtain the product C2 of the length of each statement in the initial partition and the average value of the statement tag values; If the difference between C1 and C2 is greater than 0, then this statement meets the entry criteria of the auxiliary memory dictionary; If the difference between C1 and C2 is greater than 0 and less than or equal to 0, then this statement does not meet the entry criteria of the auxiliary memory dictionary.

Citation Information

Patent Citations

  • Lossless data compression method based on LZ77, error code repair method, encoder and decoder

    CN108880556A

  • Device and method for quickly implementing LZ77 compression based on FPGA

    CN109672449A