A compression method, device and medium based on LZ77
By using double-byte stepping to perform data matching and replacement matching pairing in LZ77 encoding, the problems of cumbersome steps and high hardware resource consumption under single-byte matching are solved, and more efficient data compression and simplified hardware implementation are achieved.
Patent Information
- Application Number
- CN202210111481.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-29
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-01-29
AI Technical Summary
The existing LZ77 encoding uses a single-byte matching method in data compression, resulting in cumbersome steps, excessive number of repeated matches, and consumes a large amount of hardware resources, which is not conducive to hardware implementation.
The data matching is used in double-byte steps, the matching string is obtained, and the matching pair is replaced, which reduces the number of repeated matches, and uses the hash value to process the data in 4 bytes as the matching unit.
The consumption of hardware resources is reduced through double-byte matching, the matching process is simplified, the efficiency of data compression is improved, and the complexity of hardware implementation is reduced.
Smart Images

Figure CN114567331B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of computer data compression, and in particular to a compression method, device and medium based on LZ77. Background Art
[0002] In current communications, computer file archiving, etc., data compression is often required, among which LZ77 encoding is particularly widely used. LZ77 encoding always contains a sliding window and a read-ahead buffer. The sliding window is a historical buffer, which is used to store information about the first m bytes of the input stream. The data range of a dynamic window can be up to 64K; the read-ahead buffer corresponds to the dynamic window, which is used to store the first n bytes of the input stream. The size of the read-ahead buffer is usually between 0 and 258, and the read-ahead buffer is filled with the next n bytes; then the dynamic window is searched for the most matching data in the read-ahead buffer. If the length of the matching data is greater than the minimum matching length (usually depends on the encoder and the size of the dynamic window, such as a 4K dynamic window, its minimum matching length is 2), then a pair of <length, distance> arrays is output, and this pair of arrays is called a matching pair. The length (length, LE) is the length of the matching data, and the distance (distance, DI) indicates how many bytes in the input stream the matching data can be found; the corresponding matching pair is used to replace the original data, and finally the data is compressed.
[0003] In recent years, when using LZ77 encoding, single-byte full matching is usually used to match data, and then the length of each matching data is obtained and compressed accordingly. However, this single-byte matching method has cumbersome steps, and the number of repeated matches is too many, which consumes a lot of hardware resources and is not conducive to hardware implementation.
[0004] Therefore, technicians in this field are in urgent need of a compression method based on LZ77 to solve the problem that the current single-byte matching method has cumbersome steps, too many repeated matches, consumes a lot of hardware resources, and is not conducive to hardware implementation. Summary of the invention
[0005] The purpose of this application is to provide a compression method, device and medium based on LZ77, so as to solve the problem that the current single-byte matching method has cumbersome steps, too many repeated matches, consumes a lot of hardware resources and is not conducive to hardware implementation.
[0006] In order to solve the above technical problems, the present application provides a compression method based on LZ77, comprising:
[0007] Get the data to be compressed;
[0008] The data to be compressed is matched in double-byte steps to obtain a matching string that matches the subsequent data; wherein the subsequent data is at least two bytes;
[0009] Get the LE value and DI value of the matching string; the LE value is the length of the matching string, and the DI value is the distance between the matching string and the subsequent data that matches it;
[0010] The obtained LE value and DI value form a matching pair, and replace the matching string in the data to be compressed with the matching pair;
[0011] After all the data to be compressed have been matched and the matching character strings have been replaced with matching pairs, compressed data is obtained.
[0012] Preferably, after obtaining the matching character string, the method further includes: re-matching the single bytes before and after the matching character string to obtain a new matching character string.
[0013] Preferably, when the LE value in the matching pair exceeds the number of bytes that the pre-read buffer can accommodate, the method further includes: splitting the matching pair into multiple matching pairs according to the LE value, wherein the LE value of the split matching pairs does not exceed the number of bytes that the pre-read buffer can accommodate.
[0014] Preferably, replacing the matching character string in the to-be-compressed data with a matching pair comprises:
[0015] The original data of the matching string is replaced by the matching pair byte by byte, and the original data is replaced by an empty bubble mark where there is a lack of matching pairs.
[0016] Correspondingly, before obtaining the compressed data, it also includes:
[0017] Remove the empty bubble marks in the data to be compressed.
[0018] Preferably, obtaining the data to be compressed includes: obtaining the data to be compressed using four bytes as a matching unit according to a hash value of the data to be compressed;
[0019] Correspondingly, after obtaining the data to be compressed, it also includes: judging whether the data to be compressed meets the matching rules, wherein the matching rules include: whether the data determined according to the DI value is within the dynamic window, and whether the obtained data to be compressed is the original text.
[0020] Preferably, when the LE value exceeds the number of bytes of data that can be processed by a single clock, the method further includes:
[0021] Replace the next byte of the to-be-compressed data of the matching string with an empty bubble marker;
[0022] Determine whether the previous byte of the matching string matches;
[0023] If it matches, increase the LE value by one;
[0024] If there is no match, the data to be compressed of the next byte of the matching string is restored to the empty bubble mark closest to the byte where the LE value is located.
[0025] Preferably, after obtaining the compressed data, the method further comprises: returning prompt information to prompt the data receiver or operator that the data compression is completed.
[0026] In order to solve the above technical problems, the present application also provides a compression device based on LZ77, comprising:
[0027] An acquisition module, used for acquiring data to be compressed;
[0028] A matching module, used for matching the data to be compressed in double-byte steps to obtain a matching string that matches the subsequent data; wherein the subsequent data is at least two bytes;
[0029] A calculation module is used to obtain the LE value and DI value of the matching string; wherein the LE value is the length of the matching string, and the DI value is the distance between the matching string and the subsequent data matching it;
[0030] A replacement module, used to form a matching pair with the acquired LE value and DI value, and replace the matching string in the data to be compressed with the matching pair;
[0031] The compression module is used to obtain compressed data after all the data to be compressed have been matched and the matching character strings have been replaced with matching pairs.
[0032] Preferably, after obtaining the matched string, the method further comprises: a single-byte retrieval module for re-matching the single bytes before and after the matched string to obtain a new matched string.
[0033] Preferably, when the LE value in the matching pair exceeds the number of bytes that the pre-read buffer can accommodate, it also includes: a splitting module, used to split the matching pair into multiple matching pairs according to the LE value, wherein the LE value of the split matching pair does not exceed the number of bytes that the pre-read buffer can accommodate.
[0034] Preferably, when the LE value exceeds the number of bytes of data that can be processed by a single clock, it also includes: a legacy original text processing module, which is used to replace the data to be compressed of the next byte of the matching string with an empty bubble mark; determine whether the previous byte of the matching string matches; if it matches, increase the LE value by one; if it does not match, restore the data to be compressed of the next byte of the matching string to the empty bubble mark closest to the byte where the LE value is located.
[0035] Preferably, after the compressed data is obtained, the method further comprises: a prompt module for returning prompt information to prompt the data receiver or operator that the data compression is completed.
[0036] In order to solve the above technical problems, the present application also provides a compression device based on LZ77, comprising:
[0037] Memory for storing computer programs;
[0038] A processor is used to implement the steps of the above-mentioned LZ77-based compression method when executing a computer program.
[0039] In order to solve the above technical problems, the present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the compression method based on LZ77 as described above are implemented.
[0040] The present application provides a compression method based on LZ77, which reduces the number of repeated matches by matching data in double-byte steps. At the same time, when obtaining a matching string, since the matching unit obtained by the hash value each time is mostly 4 bytes as a matching unit, when obtaining a matching string longer than 4 bytes, it is necessary to concatenate the matching units to form a longer matching string. Using double-byte steps is simpler than the currently used single-byte concatenation, which further reduces resource consumption.
[0041] The LZ77-based compression device and computer-readable storage medium provided in this application correspond to the above method and have the same effects as above. BRIEF DESCRIPTION OF THE DRAWINGS
[0042] In order to more clearly illustrate the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0043] Figure 1 A flowchart of a compression method based on LZ77 provided by the present invention;
[0044] Figure 2 A schematic diagram of data processing for double-byte fusion provided by the present invention;
[0045] Figure 3 A data processing schematic diagram of a compression method based on LZ77 provided by the present invention;
[0046] Figure 4 A schematic diagram of data processing for single-byte retrieval provided by the present invention;
[0047] Figure 5 A schematic diagram of the processing of the legacy original text data provided by the present invention;
[0048] Figure 6 A structural diagram of a compression device based on LZ77 provided by the present invention;
[0049] Figure 7 A structural diagram of another LZ77-based compression device provided by the present invention. DETAILED DESCRIPTION
[0050] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0051] The core of this application is to provide a compression method, device and medium based on LZ77.
[0052] In order to enable those skilled in the art to better understand the present application, the present application is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0053] In today's communications and computer file storage, LZ77 encoding is often used to achieve data compression. This is done by comparing whether there are two identical strings of data in the data stream that are greater than the minimum matching length. If so, one string can be used to represent another matching string. This string is called a matching string.
[0054] For example, Figure 2 As shown, Figure 2 (a) represents the original input data; Figure 2(b) represents the data after double-byte fusion; each square represents a byte; the letters in the square represent the value of the data, and different letters represent different values of the data; the data flows from left to right; the vertical dotted line is the boundary junction between the dynamic window and the pre-read buffer, the left side of the dotted line is the pre-read buffer, and the right side of the dotted line is the dynamic window; LI indicates that the data in the upper square is the uncompressed original data; DI indicates that the upper square is the DI value; LE indicates that the upper square is the LE value; INV indicates that the upper square is an empty bubble mark. In the original input data stream, there are two identical strings framed by dotted lines, each of which is 4 bytes long. A string can be found by shifting eight bytes to find another string, so a string can be represented by another string through the relationship of <length, distance>. For example, in this example, it is <4, 8>, which means that after 8 bytes, a string of 4 bytes is consistent with the string at this location.
[0055] Therefore, it can be seen from the above that how to find two matching strings is the key to the implementation of LZ77 encoding. How to reduce the number and complexity of matching operations while ensuring that the length of the matched string is as long as possible is a problem that people in this technical field pay attention to. At present, single-byte full matching is usually used to match data. This method requires a large number of matching operations. At the same time, when the length of the matching string is long, it needs to be connected multiple times, which consumes a lot of hardware resources. Therefore, if Figure 1 As shown, please provide a compression method based on LZ77, including:
[0056] S11: Obtain data to be compressed.
[0057] In most cases, the hardware cannot process the entire data stream, so a hash algorithm is usually used to calculate a hash value with the same characteristics for all input data, and the data is stored in a cache based on the hash value to form a continuously updated lookup table. When certain data is needed, it is only necessary to search and read it based on its corresponding hash value. Generally speaking, when performing read and write queries, 4 bytes are used as a unit for data matching, which is called a matching unit.
[0058] S12: performing data matching on the data to be compressed in double-byte steps to obtain a matching character string that matches the subsequent data; wherein the subsequent data is at least two bytes.
[0059] After obtaining the above-mentioned matching units, data matching is performed in double-byte steps. Considering that the length of the matching string may be more than 4 bytes, the complete matching string cannot be obtained through one matching unit, so it is also necessary to connect the matching units that meet the matching conditions in series to finally obtain the complete matching string.
[0060] S13: Obtain the LE value and DI value of the matching string; wherein the LE value is the length of the matching string, and the DI value is the distance between the matching string and subsequent data that matches it.
[0061] After obtaining the matching string, it is easy to obtain the byte length of the matching string and the byte distance between it and the corresponding matching string, such as Figure 2 As shown, the two strings enclosed by the dotted lines are strings that match each other, and the first string is the matching string mentioned above, so it is easy to know that the length of the matching string is 4 bytes and the distance is 8 bytes.
[0062] S14: The obtained LE value and DI value are combined into a matching pair, and the matching character string in the data to be compressed is replaced with the matching pair.
[0063] After obtaining the above LE and DI values, we can get a pair of<DI,LE> This pair of data is called a matching pair. Figure 2 The matching string shown in the figure has a matching pair of <8, 4>, which is represented in hexadecimal in actual encoding as Figure 2 <0x8, 0x4> in .
[0064] It should be noted that the present embodiment provides<DI,LE> The matching pair of the form is only one embodiment, and the matching pair can also be<LE,DI> In the meantime, the DI value is generally 16 bits, which occupies 2 bytes, and the length is generally limited to between 0 and 258 in practical applications, which occupies one byte, so the matching pair is represented in byte form as<DI,DI,LE> , where two DIs represent the 2-byte value of the same DI value. So Figure 2 The matching pair in is <0x0, 0x8, 0x4>.
[0065] S15: After all the data to be compressed have been matched and the matching character strings have been replaced with matching pairs, compressed data are obtained.
[0066] After all the data is matched, all the matching strings and their corresponding matching pairs can be obtained, and all the matching strings are replaced with matching pairs to achieve data compression. Figure 2 In the example, the length of the original input data stream is 16 bytes, and the length of the matching string is 4 bytes. After the matching string is replaced by the matching pair <0x0, 0x8, 0x4>, the length of the data stream becomes 15 bytes, which means that data compression is performed. Although only one byte is compressed in this example, it is easy to know that when the data stream is long, there will be a large number of matching strings and longer matching strings. In this way, the effect of data compression will be very impressive.
[0067] The present application provides a compression method based on LZ77, which reduces the number of repeated matches when performing data matching by using double bytes as the stepping for data matching. When connecting the matching units in series to obtain the final matching series string, using double bytes as the stepping is also simpler than using a single byte, further reducing the consumption of hardware resources and making it easier to implement in hardware.
[0068] In the above embodiment, double bytes are used as the stepping of data matching to achieve the effect of reducing the number of matches and reducing the consumption of hardware resources. However, only by using double bytes as the stepping to perform data matching, the matching string with an odd length is not accurately recognized, so that when the matching string is replaced with a matching pair, there are still matching bytes that are not recognized, and then no replacement is performed, which affects the compression rate of data compression. Therefore, based on the above embodiment, this embodiment also provides a preferred implementation scheme, which, after obtaining the matching string, further includes:
[0069] S16: Rematch the single bytes before and after the matching string to obtain a new matching string.
[0070] Since data matching is performed in double bytes, a match can only be identified when the two sets of data with a length of two bytes are both consistent. If only one byte is consistent, it cannot be identified.
[0071] like Figure 3 As shown in (a), for the 5-byte matching string outlined in the dotted box, only the following can be identified by double-byte step matching: Figure 3 The 4-byte matching string in the dotted box in (b) cannot match the data E, so a preferred solution provided in this embodiment further re-matches the matching string obtained by double-byte step matching, and the matching objects are the first and last bytes of the original matching string, such as Figure 3 In (c), since the previous byte of the original matching string exceeds the range of the pre-read buffer, only the next byte is re-matched, that is, data E. After matching, it is found that data E also meets the matching rules, so the length of the matching string is increased by one, and the matching pair becomes <0x0, 0x8, 0x5>.
[0072] To further illustrate a preferred solution provided in this embodiment, Figure 4 As shown, Figure 4 (a) retrieve the matching original data for a single byte; Figure 4 (b) retrieve the data when the forward match is successful for a single byte; Figure 4 (c) Single byte retrieval of data when backward matching is successful; Figure 4 (d) Single-byte retrieval of data when both forward and backward matching are successful.
[0073] like Figure 4 As shown in (b), if the single-byte forward match is successful, the data of the previous byte is replaced with an empty bubble mark and the LE value is increased by one; Figure 4 As shown in (c), if the single-byte retrieval and backward matching are successful, the data of the next byte is replaced with an empty bubble mark, and the LE value is increased by one; Figure 4 As shown in (d), if the single-byte retrieval matches both forward and backward successfully, the data of the previous byte and the next byte are replaced with an empty bubble mark, and the LE value is increased by two.
[0074] This embodiment provides a preferred solution for single-byte retrieval, so that the compression process is initially matched in double-byte steps, and then single-byte retrieval is used to make up for the insufficient compression rate caused by double-byte steps. At the same time, this embodiment combines the use of double bytes as matching steps in the above embodiments to split the originally complex single-byte full matching into two matching processes of double bytes and single bytes, thereby simplifying the matching compression process, reducing the number of matches, and further reducing the consumption of hardware resources.
[0075] At the same time, in actual applications, the range of the dynamic window is relatively large, which can reach 64K at most, while the range of LE is relatively small, which is only 258 at most. At present, the length of the matching string in the dynamic window is mainly calculated, so when the calculated LE value exceeds the limit, it must be sliced. Therefore, this embodiment also provides a preferred implementation scheme. When the LE value in the matching pair exceeds the number of bytes that can be accommodated by the pre-read buffer, the method further includes:
[0076] S21: Decompose the matching pair into multiple matching pairs according to the LE value, wherein the LE value of the matched pair after decomposition does not exceed the number of bytes that the pre-read buffer can accommodate.
[0077] For example Figure 3 The length of the matching string in is 5 bytes. If the pre-read buffer size is only two bytes, the LE value should not exceed 2. The matching pair should be split into two LE values 2 and one LE value 1. The matching pair with one LE value of 2 is as follows: Figure 3 As shown in (d), it is <0x0, 0x8, 0x2>.
[0078] It is worth noting that in actual applications, some users have the need to convert length coding into Huffman coding. At this time, according to the Huffman coding correspondence principle, the length 3 to 258 is mapped to 0 to 255 coding, and whether to perform coding conversion can be decided according to actual needs.
[0079] This embodiment slices the LE value that exceeds the number of bytes that the pre-read buffer can accommodate to meet the requirements, thereby avoiding adverse effects on compression caused by the LE value exceeding the number of bytes that the pre-read buffer can accommodate.
[0080] As can be seen from the above embodiments, since the data matching process cannot be completed in one time, and in order to avoid the compressed data from being repeatedly calculated in the subsequent compression process, this embodiment also provides a preferred implementation scheme, in which the step S14 of replacing the matching string in the data to be compressed with the matching pair includes:
[0081] S141: The original data of the matching string is replaced by the matching pair byte by byte, and the original data is replaced with an empty bubble mark where it is insufficient.
[0082] S142: Correspondingly, before obtaining the compressed data, it also includes:
[0083] S143: Remove the empty bubble mark in the data to be compressed.
[0084] The air bubble is marked in Figure 2 and Figure 3 In the example, INV is used to represent the empty bubble mark. The square above INV is filled with meaningless data such as 0x0 to avoid repeated calculations during the compression process. However, before outputting the compressed data, the empty bubble mark must be removed. Figure 3 As shown in (e), after length slicing, there are two bytes in the original matching data string framed by the dotted line that are empty bubble markers, which need to be removed, and the final result is Figure 3 The 3-byte string is enclosed by the dotted box in (e).
[0085] The preferred solution provided in this embodiment replaces the compressed data with an empty bubble marker, thereby avoiding repeated calculations in subsequent matching or compression processes, thereby reducing the number of matches and saving consumption of hardware resources.
[0086] In practical applications, hash values with the same characteristics are often calculated from each data through a hash algorithm, and each data is stored in a cache according to the hash value, thereby forming a continuously updated lookup table. When subsequent reading and writing are performed, the corresponding data can be obtained according to the lookup table. Therefore, when obtaining the corresponding data according to the hash value, this embodiment also provides a preferred implementation scheme, and step S11 of obtaining the data to be compressed includes:
[0087] S111: Acquire the data to be compressed according to the hash value of the data to be compressed using four bytes as a matching unit.
[0088] Correspondingly, after obtaining the data to be compressed, it also includes:
[0089] S112: Determine whether the data to be compressed meets the matching rules, wherein the matching rules include: whether the data determined according to the DI value is within the dynamic window, and whether the acquired data to be compressed is the original text.
[0090] The data that meets the matching rules will continue to be compressed, while the data that does not meet the matching rules will not be compressed and output as is. Before the subsequent matching and compression operations are carried out, a preliminary judgment is made on the data to be compressed to avoid the problem of wasting hardware resources by performing useless matching processes on the data to be compressed that does not meet the matching rules.
[0091] To further illustrate a compression method based on LZ77 provided by the present application, the present embodiment further provides a preferred solution. When the LE value exceeds the number of bytes of data that can be processed by a single clock, the method further includes:
[0092] S22: Replace the next byte of the to-be-compressed data of the matching string with an empty bubble mark.
[0093] S23: Determine whether the previous byte of the matching string matches, if so, go to step S24, if not, go to step S25.
[0094] S24: Increase the LE value by one.
[0095] S25: restore the data to be compressed of the next byte of the matching string to the empty bubble mark closest to the byte where the LE value is located.
[0096] In order to improve the efficiency of compression, pipelined data compression is often used in practical applications. Correspondingly, pipelined data processing can only process current data, and past data cannot be modified again. Therefore, when performing the single-byte retrieval step, if Figure 5 As shown in (a), the 1-byte LI data before compression is encountered; Figure 5 As shown in (b), it is first modified into an empty bubble mark, and then the result of the forward single-byte comparison is used to determine whether the single-byte retrieval is successful or failed; if it fails, Figure 5 As shown in (b), the original LI is restored to the position of INV closest to LE; if successful, Figure 5 As shown in (c), the length of LE is increased by 1. This meets the need for pipelining, ensures the delay and performance of system processing, and improves the efficiency of data compression.
[0097] After the data is compressed, this embodiment further provides a preferred implementation scheme, further comprising:
[0098] S26: Return a prompt message to prompt the data receiver or operator that the data compression is completed.
[0099] This allows relevant personnel to be notified in a timely manner after data compression is completed, and then the data can be transmitted or stored.
[0100] In the above embodiment, a compression method based on LZ77 is described in detail, and the present application also provides an embodiment corresponding to a compression device based on LZ77. It should be noted that the present application describes the embodiments of the device part from two perspectives, one is based on the perspective of functional modules, and the other is based on the perspective of hardware.
[0101] like Figure 6 As shown, from the perspective of functional modules, this embodiment provides a compression device based on LZ77, including:
[0102] An acquisition module 31 is used to acquire data to be compressed;
[0103] The matching module 32 is used to perform data matching on the data to be compressed in double-byte steps to obtain a matching string that matches the subsequent data; wherein the subsequent data is at least two bytes;
[0104] The calculation module 33 is used to obtain the LE value and DI value of the matching string; wherein the LE value is the length of the matching string, and the DI value is the distance between the matching string and the subsequent data matching it;
[0105] A replacement module 34, configured to form a matching pair with the acquired LE value and DI value, and replace the matching string in the data to be compressed with the matching pair;
[0106] The compression module 35 is used to obtain compressed data after completing data matching for all data to be compressed and replacing matching character strings with matching pairs.
[0107] Preferably, after obtaining the matched string, the method further comprises: a single-byte retrieval module for re-matching the single bytes before and after the matched string to obtain a new matched string.
[0108] Preferably, when the LE value in the matching pair exceeds the number of bytes that the pre-read buffer can accommodate, it also includes: a splitting module, used to split the matching pair into multiple matching pairs according to the LE value, wherein the LE value of the split matching pair does not exceed the number of bytes that the pre-read buffer can accommodate.
[0109] Preferably, when the LE value exceeds the number of bytes of data that can be processed by a single clock, it also includes: a legacy original text processing module, which is used to replace the data to be compressed of the next byte of the matching string with an empty bubble mark; determine whether the previous byte of the matching string matches; if it matches, increase the LE value by one; if it does not match, restore the data to be compressed of the next byte of the matching string to the empty bubble mark closest to the byte where the LE value is located.
[0110] Preferably, after the compressed data is obtained, the method further comprises: a prompt module for returning prompt information to prompt the data receiver or operator that the data compression is completed.
[0111] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, which will not be repeated here.
[0112] The compression device based on LZ77 provided in this embodiment uses double bytes as the stepping of data matching, so that the number of repeated matching is reduced when performing data matching. When the matching units are connected in series to obtain the final matching series string, using double bytes as the stepping is also simpler than using a single byte, which further reduces the consumption of hardware resources and is easier to implement in hardware.
[0113] Figure 7 A structural diagram of a compression device based on LZ77 is provided in another embodiment of the present application, such as Figure 7 As shown, a compression device based on LZ77 includes: a memory 40 for storing a computer program;
[0114] The processor 41 is used to implement the steps of a compression method based on LZ77 in the above embodiment when executing a computer program.
[0115] The LZ77-based compression device provided in this embodiment may include but is not limited to a smart phone, a tablet computer, a laptop computer or a desktop computer.
[0116] Among them, the processor 41 may include one or more processing cores, such as a 4-core processor, an 8-core processor, etc. The processor 41 can be implemented in at least one hardware form of a digital signal processor (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The processor 41 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as a central processing unit (CPU); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 41 may be integrated with a graphics processing unit (GPU), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 41 may also include an artificial intelligence (AI) processor, which is used to process computing operations related to machine learning.
[0117] The memory 40 may include one or more computer-readable storage media, which may be non-transitory. The memory 40 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 40 is at least used to store the following computer program 401, wherein, after the computer program is loaded and executed by the processor 41, it can implement the relevant steps of a compression method based on LZ77 disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 40 may also include an operating system 402 and data 403, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 402 may include Windows, Unix, Linux, etc. Data 403 may include but is not limited to a compression method based on LZ77, etc.
[0118] In some embodiments, a compression device based on LZ77 may also include a display screen 42 , an input / output interface 43 , a communication interface 44 , a power supply 45 , and a communication bus 46 .
[0119] Those skilled in the art will understand that Figure 7 The structure shown in the figure does not constitute a limitation of a compression device based on LZ77, and may include more or fewer components than shown in the figure.
[0120] An embodiment of the present application provides a compression device based on LZ77, including a memory and a processor. When the processor executes a program stored in the memory, it can implement the following method: a compression method based on LZ77.
[0121] The present embodiment provides a compression device based on LZ77, which uses a processor to execute a compression method based on LZ77 stored in a memory to implement the use of double bytes as the stepping for data matching, thereby reducing the number of repeated matching times when performing data matching. When performing series connection of each matching unit to obtain a final matching series string, using double bytes as the stepping is also simpler than using a single byte, further reducing the consumption of hardware resources and being easier to implement in hardware.
[0122] Finally, the present application also provides an embodiment corresponding to a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps recorded in the above method embodiment are implemented.
[0123] It is understandable that if the method in the above embodiment is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application is essentially or the part that contributes to the prior art or all or part of the technical solution can be embodied in the form of a software product, and the computer software product is stored in a storage medium to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM), random access memory (RAM), disk or optical disk and other media that can store program code.
[0124] A computer-readable storage medium provided in this embodiment uses a processor to execute a compression method based on LZ77 stored in the storage medium to implement the use of double bytes as the stepping for data matching, thereby reducing the number of repeated matching times when performing data matching. When connecting the matching units in series to obtain the final matching series string, using double bytes as the stepping is also simpler than using a single byte, further reducing the consumption of hardware resources and being easier to implement in hardware.
[0125] The above is a detailed introduction to a compression method, device and medium based on LZ77 provided by the present application. The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the embodiments can be referred to each other. For the device disclosed in the embodiment, since it corresponds to the method disclosed in the embodiment, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
[0126] It should also be noted that, in this specification, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, an element defined by the statement "comprises a ..." does not exclude the presence of other identical elements in the process, method, article or device including the element.
Claims
1. A compression method based on LZ77, It is characterized in that include: Get the data to be compressed; Performing data matching on the data to be compressed in double-byte steps to obtain a matching character string that matches the subsequent data; wherein the subsequent data is at least two bytes; Obtaining the LE value and DI value of the matching string; wherein the LE value is the length of the matching string, and the DI value is the distance between the matching string and the subsequent data matching the matching string; The obtained LE value and the DI value form a matching pair, and the matching string in the data to be compressed is replaced by the matching pair; After all the data to be compressed have completed the data matching and the matching character string has been replaced with the matching pair, compressed data is obtained; Wherein, when the LE value in the matching pair exceeds the number of bytes that can be accommodated by the pre-read buffer, the method further includes: Decomposing the matching pair into a plurality of matching pairs according to the LE value, wherein the LE value of the matching pair after decomposition does not exceed the number of bytes that the pre-read buffer can accommodate; After replacing the matching character string in the to-be-compressed data with the matching pair, the method further includes: According to the Huffman coding correspondence principle, the LE value in the matching pair is mapped to 0-255, and Huffman coding is performed to obtain the encoded matching pair.
2. The LZ77-based compression method according to claim 1, It is characterized in that After obtaining the matching string, the following is further included: The single bytes before and after the matching character string are re-matched to obtain a new matching character string.
3. The LZ77-based compression method according to claim 1, It is characterized in that The step of replacing the matching character string in the to-be-compressed data with the matching pair comprises: The original data of the matching string is replaced by the matching pair, which is replaced byte by byte, and the original data is replaced with an empty bubble mark where it is insufficient; Correspondingly, before obtaining the compressed data, the step further includes: The empty bubble mark in the data to be compressed is removed.
4. The LZ77-based compression method according to any one of claims 1 to 3, It is characterized in that The obtaining of the data to be compressed comprises: Acquire the data to be compressed according to the hash value of the data to be compressed using four bytes as a matching unit; Correspondingly, after obtaining the data to be compressed, the method further includes: Determine whether the data to be compressed meets the matching rules, wherein the matching rules include: whether the data determined according to the DI value is within the dynamic window, and whether the acquired data to be compressed is the original text.
5. The LZ77-based compression method according to claim 2, It is characterized in that When the LE value exceeds the number of bytes of data that can be processed by a single clock, the method further includes: Replacing the data to be compressed at the next byte of the matching string with an empty bubble mark; Determine whether the previous byte of the matching string matches; If it matches, the LE value is increased by one; If there is no match, the data to be compressed of the next byte of the matching string is restored to the empty bubble mark closest to the byte where the LE value is located.
6. The LZ77-based compression method according to claim 1, It is characterized in that After obtaining the compressed data, the method further includes: A prompt message is returned to inform the data recipient or operator that data compression is complete.
7. A compression device based on LZ77, It is characterized in that include: An acquisition module, used for acquiring data to be compressed; A matching module, used for performing data matching on the data to be compressed in double-byte steps to obtain a matching string that matches the subsequent data; wherein the subsequent data is at least two bytes; A calculation module, used to obtain the LE value and the DI value of the matching string; wherein the LE value is the length of the matching string, and the DI value is the distance between the matching string and the subsequent data matching the matching string; A replacement module, used for forming a matching pair from the acquired LE value and the DI value, and replacing the matching string in the data to be compressed with the matching pair; A compression module, used for obtaining compressed data after all the data to be compressed have completed the data matching and replaced the matching character string with the matching pair; Wherein, when the LE value in the matching pair exceeds the number of bytes that can be accommodated by the pre-read buffer, the LZ77-based compression device is specifically used to: Decomposing the matching pair into a plurality of matching pairs according to the LE value, wherein the LE value of the matching pair after decomposition does not exceed the number of bytes that the pre-read buffer can accommodate; The LZ77-based compression device is specifically used for: According to the Huffman coding correspondence principle, the LE value in the matching pair is mapped to 0-255, and Huffman coding is performed to obtain the encoded matching pair.
8. A compression device based on LZ77, It is characterized in that include: Memory for storing computer programs; A processor, configured to implement the steps of the LZ77-based compression method as described in any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the LZ77-based compression method according to any one of claims 1 to 6 are implemented.
Citation Information
Patent Citations
Coding method and device
CN105426413A