An improved differential compression algorithm for unstructured data

Through the differential compression improvement algorithm for unstructured data, the network burden problems caused by the inconsistency between the cost model and the actual storage mode and the quadratic matching in the prior art are solved, and more efficient data storage and transmission are achieved.

CN118300615BActive Publication Date: 2025-06-06XIAN TECH UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410483985.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-04-22
Publication Date
2025-06-06
Estimated Expiration
2044-04-22

AI Technical Summary

Technical Problem

The cost model of the existing differential compression algorithm does not match the actual storage mode of the computer system, and it is impossible to objectively evaluate its actual application performance. The quadratic matching leads to repeated operation, increasing network burden and storage space occupation.

Method used

A differential compression improvement algorithm for unstructured data is proposed. By setting a buffer to store unmerged commands, running the 1.5-pass algorithm and dividing the incremental cost sequence, step three and step four are alternately executed, and the commands are merged in turn to reduce the cost of the incremental coding sequence.

Benefits of technology

The cost model of this algorithm is closer to the actual storage mode, reduces repeated matching and editing commands, reduces data storage space and network transmission bandwidth, and avoids network congestion and transmission delay.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118300615B_ABST
    Figure CN118300615B_ABST
Patent Text Reader

Abstract

The present invention belongs to the field of electronic information technology, and specifically relates to an improved differential compression algorithm for unstructured data. The following steps are included: step 1, setting a buffer; step 2, reading reference data R and target data T; step 3, running the 1.5-pass algorithm, when the ADD command appears, or the data T has been matched, suspending the operation, and placing the commands of the selected area and the new area into the buffer in sequence; step 4, sequentially merging the commands of the selected area in the buffer with rules R5, R1, R3 or R4, R4 or R3, R2, outputting the new command after completion and suffixing it to the corrected area; step 5, alternately executing steps 3 and 4 until there is no new area, and ending the operation. The present invention proposes a cost model based on the actual storage method of the computer system, which reduces the cost of the incremental coding sequence by merging instructions, and then uses the incremental cost coding sequence to replace the original data for storage and transmission, which can effectively save data storage space and network transmission speed.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the field of electronic information technology, and in particular relates to an improved differential compression algorithm for unstructured data. Background Art

[0002] In recent years, with the rapid development of high-tech technologies such as artificial intelligence, the Internet of Things, and cloud computing, the amount of global data is growing and expanding without limit. These data have penetrated into various aspects of society, economy, scientific research, and are a microcosm of various things in real life. If we can effectively use these data, scientific research may achieve breakthroughs, and technology will become more adaptable, personalized, and robust. It can be seen that data, as a carrier of various information, contains rich knowledge and has great value.

[0003] Before extracting valuable information from these data, the raw data needs to be stored, managed, and analyzed. It can be said that data storage is the basic technology. However, only a small part of the massive data is actually stored, and the vast majority of the data may be discarded, and its potential social and economic value may not be realized. Therefore, how to effectively store and manage this data has become a rather thorny issue. Although the price of storage devices is constantly falling, it is still far behind the rising speed of the amount of data stored and processed. Therefore, the advancement and innovation of storage technology is a key factor in overcoming the data capacity gap.

[0004] Since the sensor devices that collect data have the problem of repeated sampling, which will cause a lot of redundancy between data, differential compression technology is applied to obtain higher storage efficiency. Differential compression is a lossless data compression technology that compresses data by using the statistical correlation between data. Specifically, differential compression takes the reference file R and the target file T as input, identifies the common segments that can be copied and the segments that need to be added in T, and edits them into COPY and ADD commands respectively to obtain an incremental coding sequence Δ. Δ is much smaller than T, so using Δ instead of T for storage and transmission can save storage space and network transmission bandwidth.

[0005] At present, all differential compression algorithms have some common characteristics. They are mainly composed of three steps, namely, selecting hash functions, matching common segments, and constructing incremental coding sequences. When a simple cost model is adopted, that is, the cost of a COPY command is 1, the cost of an ADD command is the length of the added segment, and the construction cost of the incremental coding sequence is the sum of the costs of the COPY and ADD commands it contains, the cost of the incremental coding sequence constructed by the greedy algorithm proposed by Reichenberger is the smallest, but the algorithm has the problem of high time and space complexity. In order to improve the time and space complexity of the greedy algorithm, AJTAI et al. designed the one-pass algorithm and the 1.5-pass algorithm successively. The one-pass algorithm and the 1.5-pass algorithm can complete compression in linear time and constant space, but many common segments are lost in the matching process, which seriously affects the compression performance. Therefore, AJTAI et al. revised these two improved algorithms and named them the revised one-pass algorithm and the revised 1.5-pass algorithm respectively.

[0006] Although the two modified algorithms have improved performance, especially the modified 1.5-pass algorithm, whose compression ratio is close to that of the greedy algorithm and whose complexity is much lower than that of the greedy algorithm, both the modified 1.5-pass algorithm and the modified one-pass algorithm inevitably perform secondary matching, which increases the workload.

[0007] In summary, there are two main problems with existing differential compression algorithms. First, an integer occupies different memory spaces in different computer systems, so the simple cost model that generalizes the cost of COPY to 1 does not conform to the actual storage mode of the computer system, resulting in the inability to objectively evaluate the actual application performance of the differential compression algorithm; second, the secondary matching of the modified 1.5-pass algorithm and the modified one-pass algorithm leads to repeated operation, which increases the burden of the algorithm. At the same time, the storage space occupied by the data is still large, which also affects the transmission speed of the network, causing network congestion and transmission delay. Summary of the invention

[0008] The present invention provides an improved differential compression algorithm for unstructured data, which aims to solve the problems that the cost model in the prior art does not match the actual storage mode of the computer system and cannot objectively evaluate its actual application performance; the secondary matching of the existing algorithm leads to repeated operation, which will cause network congestion and transmission delay.

[0009] In order to achieve the above object, the technical solution of the present invention is as follows: an improved differential compression algorithm for unstructured data, comprising the following steps:

[0010] Step 1: Setting a buffer for storing unmerged commands;

[0011] Step 2: Read reference data R and target data T;

[0012] Step 3: Run the 1.5-pass algorithm. Once an ADD command appears or the target data T has been matched, stop running, divide the obtained incremental cost sequence Δ into the corrected area, the selected area and the new area, and put the commands of the selected area and the new area into the buffer in sequence;

[0013] Step 4: Process the commands in the selected area in the buffer: merge them using rules R5, R1, R3 or R4, R4 or R3, R2 in sequence, and after completion, output the new commands from the buffer and suffix them to the corrected area;

[0014] Step 5: Alternately execute steps 3 and 4 until there are no new regions and obtain an incremental coding sequence Δ with a smaller cost.

[0015] Furthermore, in the above step three, the corrected area only contains the merged commands; the new area refers to the latest generated part, which is composed of an ADD command, or an ADD and a COPY command; the selected area is the command sequence to be merged, and its head and tail ADD commands are borrowed from the corrected area and the new area respectively.

[0016] Furthermore, in the above step 3, the cost model of the incremental coding sequence Δ obtained by running the 1.5-pass algorithm is:

[0017] Δ= <CM 1 ,CM 2 ,……,CM n >,

[0018] Among them: CM is a COPY or ADD command. The format of the COPY command is COPY(lgc,pst), which means copying the fragment of length lgc starting at pst in data R to data T. The format of the ADD command is ADD(lga,sgm), which means adding the string sgm of length lga to data T. The costs of the COPY and ADD commands are as follows:

[0019] cost(COPY(lgc,pst))=2L

[0020] cost(ADD(lga,sgm))=lga+L

[0021] Where: L represents the storage space occupied by an integer in the same computer system;

[0022] The cost of the Δ sequence is the sum of the costs of the commands it contains.

[0023]

[0024] Furthermore, in the above step 4, the 5 merge rules are

[0025] R1:

[0026] In any case, a full-ADD sequence should be combined into a single ADD command;

[0027] R2:

[0028] When satisfied When dual-ADD sequence Δ= <ADD(lga 1 ,sgm 1 ),COPY(lgc 1 ,pst 1 ),…,COPY(lgc n ,pst n ),ADD(lga 2 ,sgm 2 )>can be combined into one command

[0029] R3:

[0030] When satisfied When head-ADD sequence Δ= <ADD(lga 1 ,sgm 1 ),COPY(lgc 1 ,pst 1 ),…,COPY(lgc n ,pst n )> can be combined into one command

[0031] R4:

[0032] When satisfied When the end-ADD sequence Δ= <COPY(lgc 1 ,pst 1 ),…,COPY(lgc n ,pst n ),ADD(lga 1 ,sgm 1 )> can be combined into one command

[0033] R5:

[0034] When satisfied When no-ADD sequence Δ= <COPY(lgc 1 ,pst 1 ),…,COPY(lgc n ,pst n )> can be combined into one command

[0035] Compared with the prior art, the advantages of the present invention are:

[0036] 1. The same type of data occupies different memory sizes in different computer systems. In the cost model proposed in the present invention, the costs of the COPY and ADD commands are determined by their corresponding operand types, and the total cost of the incremental encoding sequence is determined by the included commands. Therefore, the cost model proposed in the present invention is closer to the actual storage mode of the computer system, and the differential compression algorithm based on this model has more realistic and effective performance.

[0037] 2. Steps 3 and 4 of the present invention are performed alternately, step 3 generates commands, step 4 performs regional division and merges in turn using rules R5, R1, R3 or R4, R4 or R3, R2, and the merging follows the principle of from short to long, first merging shorter sequences, and then merging longer sequences. Such a merging order can reduce the number of merging times and improve merging efficiency. By merging instructions to reduce the cost of incremental coding sequences, and by using incremental cost coding sequences instead of original data for storage and transmission, it can avoid repeated matching and editing commands, and can also reduce the cost of generating incremental coding sequences, effectively saving data storage space and network transmission bandwidth, which can save data storage space and network transmission speed, prevent network congestion, and reduce transmission delays. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 It is the overall idea of ​​the present invention;

[0039] Figure 2 This is a combined schematic diagram of step 4 of the present invention;

[0040] Figure 3 This is the curve showing how the cost of generating incremental encoding sequences varies with L. DETAILED DESCRIPTION

[0041] The present invention will be described in detail below with reference to the accompanying drawings and embodiments.

[0042] Embodiment: An improved differential compression algorithm for unstructured data specifically comprises the following steps:

[0043] Step 1: Set a buffer to store unmerged commands;

[0044] Step 2: Read reference data R and target data T;

[0045] Step 3: Run the 1.5-pass algorithm. Once an ADD command appears or the target data T has been fully matched, stop running, divide the obtained incremental cost sequence Δ into the corrected area, the selected area and the new area, and put the instructions of the selected area and the new area into the buffer in sequence:

[0046] Combination Figure 1 It can be seen that the corrected area refers to the command sequence after the merger; the new area is composed of the latest generated ADD and COPY commands, either an ADD command or an ADD and COPY command, which is determined by the 1.5-pass algorithm; the selected area contains the commands to be merged, and its head and tail ADD commands are borrowed from the corrected area and the new area respectively.

[0047] The incremental cost sequence Δ obtained after the 1.5-pass algorithm is run is in the form:

[0048] The incremental encoding sequence Δ consists of COPY commands and ADD commands. The format of the COPY command is COPY(lgc,pst), which means copying a segment of length lgc from position pst in R to T; the format of the ADD command is ADD(lga,sgm), which means adding a segment sgm of length lga to data T. In actual computer systems, the size of the space used to store an integer is fixed, represented by L; the size of the space used to store a segment is determined by the segment length, i.e. lga, so we get a new cost model:

[0049] Given an incremental encoding sequence Δ = <CM 1 ,CM 2 ,……,CM n >, where CM is a COPY or ADD command.

[0050] The costs of COPY, ADD, and Δ are:

[0051] (1) The cost of any COPY command is 2L, expressed as cost(COPY(lgc,pst))=2L;

[0052] (2) The cost of an ADD command is the sum of the length of the corresponding segment and L, expressed as cost(ADD(lga,sgm))=lga+L;

[0053] (3) The cost of the Δ sequence is the cumulative sum of the costs of the commands it contains, that is,

[0054] Step 4: Process the commands in the selected area in the buffer, and merge them using rules R5, R1, R3 or R4, R54 or R3, R2 in sequence. After completion, output the new command from the buffer and suffix it to the corrected area. The merging idea is as follows:

[0055] The Δ sequences are classified according to the position and number of ADD commands, and the following five categories are obtained:

[0056] Given an incremental encoding sequence Δ = <CM 1 ,CM 2 ,……,CM n >, where CM is a COPY or ADD command,

[0057] (1) If CM 1 ,……,CM n If all are ADD, then Δ is a full-ADD sequence;

[0058] (2) If in addition to CM 1 and CM n is ADD, and the others are COPY commands, then Δ is a dual-ADD sequence;

[0059] (3) If only CM 1 is ADD, the rest are CM 2 ,……,CM n All are COPY, then Δ is a head-ADD sequence;

[0060] (4) If only CM n is ADD, the rest are CM 1 ,……,CM n-1 All are COPY, then Δ is an end-ADD sequence;

[0061] (5) If CM 1 ,……,CM n All are COPY, then Δ is a no-ADD sequence;

[0062] Among them, the head-ADD sequence and the end-ADD sequence are collectively referred to as a single-ADD sequence.

[0063] For a full-ADD sequence Δ= <ADD(lga 1 ,sgm 1 ),…,ADD(lga n ,sgm n )>, in any case, can be merged into an ADD(lga 1 +…+lga n ,sgm 1 +…+sgmn ) command, this is because That is, cost(Δ)>cost(ADD(lga 1 +…+lga n ,sgm 1 +…+sgm n )).

[0064] Using the same design idea, the present invention provides 5 merging rules:

[0065] R1: In any case, a full-ADD sequence should be combined into a single ADD command;

[0066] R2: When satisfied When dual-ADD sequence Δ= <ADD(lga 1 ,sgm 1 ),COPY(lgc 1 ,pst 1 ),…,COPY(lgc n ,pst n ),ADD(lga 2 ,sgm 2 )> can be combined into one command

[0067] R3: When satisfied When head-ADD sequence Δ= <ADD(lga 1 ,sgm 1 ),COPY(lgc 1 ,pst 1 ),…,COPY(lgc n ,pst n )> can be combined into one command

[0068] R4: When satisfied When the end-ADD sequence Δ= <COPY(lgc 1 ,pst 1 ),…,COPY(lgc n ,pst n ),ADD(lga 1 ,sgm 1 )> can be combined into one command

[0069] R5: When satisfied When no-ADD sequence Δ= <COPY(lgc 1 ,pst 1),…,COPY(lgc n ,pst n )> can be combined into one command

[0070] It can be seen that from rule R1 to rule R5, the conditions required to be met are becoming more and more stringent. So, how can we efficiently select rules R1 to R5, merge and improve instructions, and thus reduce the cost of the selected area? Figure 2 , the merging process of the present invention can be divided into four stages: Since the replacement condition of rule R5 is the most stringent, in the first stage, R5 is used to merge COPY commands to obtain more ADDs. In order to replace and be more rigorous and reduce the cost of incremental encoding, COPY instructions are merged from few to many, first merging two adjacent COPY instructions, then merging three or four adjacent instructions... until they cannot be merged. During the merging process, adjacent ADD instructions may be generated and need to be merged with R1. After the first stage is completed, the selected area can be regarded as being connected by some head-ADD sequences, so R3 is used for merging, and the merging is also from short to long, which is the second stage. In the third stage, the selected area can also be regarded as consisting of some end-ADD sequences, so rule R4 is used for merging and replacement, and the merging principle is the same as the second stage, also from short to long. The third stage can be swapped with the second stage, that is, the order of R3 and R4 can be interchanged. In this embodiment, it is carried out in the order of R3 and R4, that is, the head-ADD sequence is merged first, and then the end-ADD sequence is merged. After the first three stages are completed, the selected area can only be merged as a whole, so the fourth stage is the application of R2. So far, the merging of the selected area is completed, and the new command is output from the buffer and suffixed to the corrected area.

[0071] Step 5: Alternately execute step 3 and step 4 until all target data T are matched and the incremental coding sequence Δ is merged, that is, there is no new area, and finally an incremental coding sequence Δ with a lower cost is obtained.

[0072] In order to verify the performance of the improved differential compression algorithm proposed in the present invention, a simulation was conducted. The experiment was conducted on a 1.6GHz Intel(R) Core(TM) i5-8265 CPU, 8GB main memory, and Windows operating system. The input data R and T are represented by strings. Data R is randomly generated. T consists of two parts, one is randomly generated, and the other is randomly captured from R. In order to ensure fairness, the lengths of these two parts and the positions in which they are placed in T are also random. Before the experiment, the correlation factor Φ(R,T) was defined to represent the similarity between data T and R, as follows:

[0073] Assume x 1 ,x 2,…,x n is a random fragment captured from data R into T, |x i | is fragment x i The length of the data T and R is defined as:

[0074]

[0075] The experiment compares the performance of greedy, one-pass, 1.5-pass, modified one-pass, modified 1.5-pass and the improved algorithm proposed by the present invention. The length of data R and T is about 5000 bytes. The program is run 1000 times and the average value is calculated. The results are as follows: Figure 3 As shown in the figure, Cor_one-pass and Cor_1.5-pass represent the modified one-pass and modified 1.5-pass algorithms respectively, and CSIM_1.5-pass represents the improved algorithm proposed by the present invention. It can be seen that no matter what value Φ(R,T) takes, the cost of the incremental coding sequence generated by these algorithms increases with the increase of L. This is because the increase of L reduces the number of matched common fragments, so the COPY commands in the sequence in the incremental coding are reduced, resulting in an increase in the total cost. However, in any case, the cost of the incremental coding sequence generated by the CSIM_1.5-pass algorithm is always less than that of other algorithms, that is, the effectiveness of the incremental compression improved algorithm proposed by the present invention has been verified, and it has significant advantages in reducing the storage space and transmission bandwidth of similar files.

[0076] The above description is only the technical method and implementation mode of the present invention, and the protection scope of the present invention is not limited thereto. Any person who makes equivalent replacement or change based on the technical solution and inventive concept of the present invention is covered by the protection scope of the present invention.

Claims

1. An improved differential compression algorithm for unstructured data, characterized in that: The following steps are involved: Step 1: Setting a buffer for storing unmerged commands; Step 2: Read reference data R and target data T; Step 3: Run the 1.5-pass algorithm. Once an ADD command appears or the target data T has been matched, stop running, divide the obtained incremental cost sequence Δ into the corrected area, the selected area and the new area, and put the commands of the selected area and the new area into the buffer in sequence; Step 4: Process the commands in the selected area in the buffer: merge them using rules R5, R1, R3 or R4, R4 or R3, R2 in sequence, and after completion, output the new commands from the buffer and suffix them to the corrected area; Step 5: Alternately execute steps 3 and 4 until there are no new regions and obtain an incremental coding sequence Δ with a smaller cost; In step 3, the cost model of the incremental coding sequence Δ obtained by running the 1.5-pass algorithm is: D= <CM1,CM2,……,CM n >, Among them: CM is a COPY or ADD command. The format of the COPY command is COPY(lgc,pst), which means copying the fragment of length lgc starting at pst in data R to data T. The format of the ADD command is ADD(lga,sgm), which means adding the string sgm of length lga to data T. The costs of the COPY and ADD commands are as follows: cost(COPY(lgc,pst))=2L cost(ADD(lga,sgm))=lga+L Where: L represents the storage space occupied by an integer in the same computer system; The cost of the Δ sequence is the sum of the costs of the commands it contains. In step 4, the 5 merging rules are: R1: In any case, a full-ADD sequence should be combined into a single ADD command; R2: When satisfied When dual-ADD sequence Δ= <ADD(lga1,sgm1),COPY(lgc1,pst1),…,COPY(lgc n ,pst n ),ADD(lga2,sgm2)> can be combined into one command R3: When satisfied When head-ADD sequence Δ= <ADD(lga1,sgm1),COPY(lgc1,pst1),…,COPY(lgc n ,pst n )> can be combined into one command R4: When satisfied When the end-ADD sequence Δ= <COPY(lgc1,pst1),…,COPY(lgc n ,pst n ),ADD(lga1,sgm1)> can be combined into one command R5: When satisfied When no-ADD sequence Δ= <COPY(lgc1,pst1),…,COPY(lgc n ,pst n )> can be combined into one command 2. The improved differential compression algorithm for unstructured data according to claim 1, characterized in that: In step 3, the corrected region only includes the merged commands; The new area refers to the most recently generated part, which is composed of an ADD command, or an ADD and a COPY command; the selected area is the command sequence to be merged, and its head and tail ADD commands are borrowed from the modified area and the new area respectively.

Citation Information

Patent Citations

  • Method for optimizing compressed storage format in data processing process

    CN111858391A

  • Dynamic compression method and dynamic compression system based on improved travel length coding

    CN112615627A