Data Compression Method and Device

By using multiple hash tables and hash node selection rules in data compression, the problem of low data compression rate under limited hardware resources is solved, and a more efficient data compression effect is achieved.

CN116112122BActive Publication Date: 2025-07-18ALIBABA (CHINA) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310097325.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-01-20
Publication Date
2025-07-18
Estimated Expiration
2043-01-20

AI Technical Summary

Technical Problem

The existing data compression hardware implementation method is difficult to achieve the same compression rate as software implementation when hardware resources are limited, and the compression rate of the LZ algorithm is relatively low.

Method used

By setting up multiple hash tables, obtain multiple data to be compressed from the data object to be compressed based on different preset ranges, generate multiple hash values, and find the hash table to obtain the location information of the compressed data, and combine the hash node and the matching length for data compression, reducing the impact of hash conflicts and improving the compression speed and rate.

Benefits of technology

In the case of limited hardware resources, the data compression rate and speed are improved, the impact of hash conflict on compression rate is reduced, and the data compression efficiency is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116112122B_ABST
    Figure CN116112122B_ABST
Patent Text Reader

Abstract

The present application provides a data compression method and an apparatus for the data compression method. The data compression method includes: obtaining a plurality of data to be compressed from a data object to be compressed based on different preset ranges; respectively generating a plurality of hash values based on the plurality of data to be compressed; looking up a hash table corresponding to each hash value to obtain the position information of the compressed data corresponding to the hash value, wherein the hash table is used to store the position information of the compressed data and the data to be compressed in the data object to be compressed, and the position information corresponds to a hash node; obtaining the matching length of the data corresponding to the hash node based on the data to be compressed and the compressed data; and compressing the data in the data object to be compressed based on the hash node and the matching length. This data compression method can improve the data compression ratio.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This specification relates to the technical field of data compression, and more particularly, to a data compression method and an apparatus for the data compression method. Background Art

[0002] In the application of big data technology, a large amount of data is generally processed in parallel on a large number of servers. There will be a large amount of information transmission between these servers, and this information needs to be transmitted and stored through compression and encryption. The hardware implementation of data compression can reduce the workload of the server to improve the data compression speed.

[0003] However, the current hardware implementation of data compression is difficult to achieve the same compression ratio as the software implementation under limited hardware resources. Summary of the Invention

[0004] A first aspect of this specification provides a data compression method, which includes: obtaining a plurality of data to be compressed from a data object to be compressed based on different preset ranges; generating a plurality of hash values respectively based on the plurality of data to be compressed; looking up a hash table corresponding to each hash value to obtain the position information of the compressed data corresponding to the hash value, where the hash table is used to store the position information of the compressed data and the data to be compressed in the data object to be compressed, and the position information corresponds to a hash node; obtaining the matching length of the data corresponding to the hash node based on the data to be compressed and the compressed data; and compressing the data in the data object to be compressed based on the hash node and the matching length.

[0005] In the above solution, by setting a plurality of hash tables, the hash nodes used to compress the data object to be compressed can be selected based on the plurality of hash tables, which can reduce the impact of hash collisions on the compression ratio, and can also enable the selection of hash nodes with relatively larger matching lengths in the subsequent process of calculating the matching length of the data to compress the data in the data object to be compressed, so as to improve the compression speed and the data compression ratio.

[0006] In an embodiment of the first aspect of this specification, for the step of obtaining a plurality of data to be compressed from a data object to be compressed based on different preset ranges, it may include: obtaining at least two data to be compressed from the data object to be compressed based on different preset ranges, and the lengths of the preset ranges of the at least two data to be compressed are the same and the initial positions are different; and / or, obtaining at least two data to be compressed from the data object to be compressed based on different preset ranges, and the lengths of the preset ranges of the at least two data to be compressed are different and the initial positions are the same.

[0007] In an embodiment of the first aspect of this specification, regarding the step of obtaining at least two data to be compressed from the data object to be compressed based on different preset ranges, it may include: obtaining four data to be compressed from the data object to be compressed based on two different preset ranges. In this embodiment, the two data to be compressed are respectively a data block at an odd position and a data block at an even position. The serial number of the byte corresponding to the starting position of the data block at the odd position in the data object to be compressed is odd, the serial number of the byte corresponding to the starting position of the data block at the even position in the data object to be compressed is even, and the starting positions of the data block at the odd position and the data block at the even position are adjacent.

[0008] In the above solution, in the entire data object to be compressed, it is possible to simultaneously perform hash calculation on the data to be compressed starting from odd and even positions, so as to increase the probability of finding the corresponding hash value from the hash table; in addition, compared with separately performing hash calculation on the data to be compressed starting from odd or even positions, this method of selecting hash nodes can reduce the impact of hash collisions on the compression ratio, and the hash nodes selected by this method may correspond to longer matchable strings, so that in the process of performing the matching length of the data, there is a greater probability of selecting a hash node that makes the matching length relatively larger to compress the data of the data object to be compressed.

[0009] In an embodiment of the first aspect of this specification, regarding the step of obtaining at least two data to be compressed from the data object to be compressed based on different preset ranges, it may also include: obtaining four data to be compressed from the data object to be compressed based on four different preset ranges. In this embodiment, the four data to be compressed are respectively a short data block at an odd position, a long data block at an odd position, a short data block at an even position, and a long data block at an even position. The serial number of the byte corresponding to the starting position of the short data block at the odd position and the long data block at the odd position in the data object to be compressed is odd, the serial number of the byte corresponding to the starting position of the short data block at the even position and the long data block at the even position in the data object to be compressed is even. The lengths of the short data block at the odd position and the short data block at the even position are the same, the lengths of the long data block at the odd position and the long data block at the even position are the same. The starting positions of the short data block at the odd position and the long data block at the odd position are the same, the starting positions of the short data block at the even position and the long data block at the even position are the same, and the starting positions of the short data block at the odd position and the short data block at the even position are adjacent.

[0010] In an embodiment of the first aspect of this specification, the step of compressing the data in the data object to be compressed based on the hash node and the matching length may include: determining whether the data to be compressed corresponding to the hash node is located at the termination position of the data object to be compressed; if so, encapsulating the compressed data and ending the compression process; if not, in the data object to be compressed, shifting the window defined by each preset range backward, and using the data covered by the window as the subsequent data to be compressed. In this embodiment, the distance that all the windows defined by the preset ranges move each time is the same.

[0011] In an embodiment of the first aspect of this specification, the distance that the window moves each time is the length occupied by two bytes.

[0012] In an embodiment of the first aspect of this specification, the data compression method may further include: when the corresponding hash values are found in at least two hash tables, based on a preset node selection rule, screening out target nodes from all the obtained hash nodes. In this embodiment, the step of obtaining the matching length of the data corresponding to the hash node based on the data to be compressed and the compressed data may include: obtaining the matching length of the data corresponding to the target node based on the data to be compressed and the compressed data.

[0013] In the above solution, one of the multiple hash nodes found from the hash table is selected as the target node, and then the calculation of the matching length of the subsequent data is performed, so that the calculation amount of the matching length of the subsequent data is reduced to reduce the compression delay.

[0014] In an embodiment of the first aspect of this specification, the node selection rule may include: comparing the position information of the hash node whose hash value is found from the hash table with the position information of the previously selected hash node, and using the hash node whose position information is inconsistent with the position information of the previously selected hash node as the target node.

[0015] In the above solution, during the process of calculating the matching length of the data, it is possible to avoid the loss of the compression ratio caused by the repetition or a large amount of repetition of the data to be compressed obtained based on the previous and subsequent hash nodes.

[0016] In an embodiment of the first aspect of this specification, the information of the data format corresponding to the hash node includes location information and a valid flag bit. The optional flags of the valid flag bit include a first flag and a second flag. The first flag is used to indicate that the hash node is invalid, and the second flag is used to indicate that the hash node is valid. In addition, the step of generating multiple hash values respectively based on multiple data to be compressed may include: if the valid flag bit corresponding to the hash node found from the hash table is the first flag, generating a hash node with the data to be compressed; if the valid flag bit corresponding to the hash node found from the hash table is the second flag, using the hash node corresponding to the data to be compressed as the hash node of the preselected target node, where if the valid flag bits corresponding to the hash nodes found from at least two hash tables are the second flag, screening out the target node through a node selection rule, and if the valid flag bit corresponding to the hash node found from only one hash table is the second flag, using the found hash node as the target node.

[0017] In an embodiment of the first aspect of this specification, the step of obtaining multiple data to be compressed from the data object to be compressed based on different preset ranges may include: initializing the hash nodes corresponding to the hash table so that the valid flag bit in the data format corresponding to the hash node is written as the first flag.

[0018] In the above solution, it is possible to avoid the hash table updated when processing the data object to be compressed in the previous process from affecting the compression process of the current data object to be compressed.

[0019] In an embodiment of the first aspect of this specification, the step of finding the hash table corresponding to each hash value and obtaining the location information of the compressed data corresponding to the hash value may include: writing the first flag bit corresponding to each hash node as the second flag, and updating the information in the data format corresponding to each hash node to the hash table based on the hash value corresponding to each hash node.

[0020] In an embodiment of the first aspect of this specification, the information of the data format may further include first check information. For the same hash node, the data used to generate the first check information and the hash value is the same, but the calculation rules are different. In this embodiment, the step of generating multiple hash values respectively based on multiple data to be compressed may include: when the valid flag bit corresponding to the hash node found from the hash table is the second flag, comparing the first check information corresponding to the hash node corresponding to the data to be compressed with the first check information corresponding to the hash node found from the hash table; if the comparison result is different, generating a hash node with the data to be compressed; if the comparison result is consistent, using the hash node corresponding to the data to be compressed as the hash node of the preselected target node.

[0021] In the above solution, by verifying the first verification information, when looking up the hash value in the hash table, there is a greater chance of screening out the hash nodes with hash conflicts, so as to improve the compression ratio of data compression.

[0022] In an embodiment of the first aspect of this specification, the information of the data format further includes second verification information. The second verification information corresponding to the previous hash node is calculated and generated by the data corresponding to the next hash node, and the calculation rules for generating the second verification information and the hash value are different. In this embodiment, for the step of generating multiple hash values respectively based on multiple data to be compressed, it may further include: when the valid flag bit corresponding to the hash node found in the hash table is the second flag, after comparing the hash node through the first verification information, if there are at least two hash nodes as preselected target nodes, compare the second verification information corresponding to the hash node corresponding to the data to be compressed with the second verification information corresponding to the hash node found in the hash table, and use the hash node when the comparison result is consistent as the hash node of the preselected target node.

[0023] In the above solution, at the position corresponding to the hash node screened out by verifying the second verification information, a relatively longer string can be obtained in the subsequent process of calculating the matching length of the data, so as to improve the compression ratio of the data object to be compressed.

[0024] The second aspect of this specification provides a data compression method, which may include: obtaining the data to be compressed within a preset range from the data object to be compressed; generating a hash value based on the data to be compressed; looking up the hash table based on the hash value to obtain the position information of the compressed data corresponding to the hash value, and generating a hash node corresponding to the data to be compressed, where the hash table is used to store the hash nodes corresponding to the compressed data, the position information of the hash nodes in the data object to be compressed, and the verification information of the compressed data, the hash node corresponds to the position information, and the calculation rules for generating the hash value and the first verification information are different; generating the first verification information of the hash node corresponding to the data to be compressed based on the data to be compressed; when the first verification information is consistent with the second verification information, obtaining the matching length of the data corresponding to the hash node based on the data to be compressed and the compressed data; compressing the data in the data object to be compressed based on the hash node and the matching length.

[0025] In the above solution, by verifying the first verification information, when looking up the hash value in the hash table, there is a greater chance of screening out the hash nodes with hash conflicts, so as to improve the compression ratio of data compression.

[0026] In the data compression methods of the first aspect and the second aspect of this specification above, each execution step is implemented by a field-programmable gate array.

[0027] The third aspect of this specification provides an apparatus for a data compression method. The apparatus includes an acquisition module, a generation module, a search module, a matching length calculation module, and a packaging module. The acquisition module is configured to acquire a plurality of data to be compressed from a data object to be compressed based on different preset ranges. The generation module is configured to generate a plurality of hash values respectively based on the plurality of data to be compressed. The search module is configured to search a hash table corresponding to each hash value to obtain the position information of the compressed data corresponding to the hash value, where the hash table is used to store the position information of the compressed data and the data to be compressed in the data object to be compressed, and the position information corresponds to a hash node. The matching length calculation module is configured to obtain the matching length of the data corresponding to the hash node based on the data to be compressed and the compressed data. The packaging module is configured to compress the data in the data object to be compressed based on the hash node and the matching length.

[0028] The third aspect of this specification provides an apparatus for a data compression method. The apparatus includes an acquisition module, a generation module, a search module, a verification module, a matching length calculation module, and a packaging module. The acquisition module is configured to acquire the data to be compressed within a preset range from the data object to be compressed. The generation module generates a hash value based on the data to be compressed. The search module searches a hash table based on the hash value to obtain the position information of the compressed data corresponding to the hash value, and generates a hash node corresponding to the data to be compressed, where the hash table is used to store the hash node corresponding to the compressed data, the position information of the hash node in the data object to be compressed, and the verification information of the compressed data, the hash node corresponds to the position information, and the calculation rules for generating the hash value and the first verification information are different. The verification module generates the first verification information of the hash node corresponding to the data to be compressed based on the data to be compressed. The matching length calculation module, when the first verification information is consistent with the second verification information, obtains the matching length of the data corresponding to the hash node based on the data to be compressed and the compressed data. The packaging module compresses the data in the data object to be compressed based on the hash node and the matching length. Description of the Drawings

[0029] Figure 1 Shown is a flowchart of a data compression method provided by an exemplary embodiment of this specification.

[0030] Figure 2 For Figure 1 Shown is a schematic structural diagram of a data compression method corresponding to the data compression method.

[0031] Figure 3 For Figure 1 Shown is a method for acquiring the data to be compressed for calculating the hash value in the data compression method.

[0032] Figure 4 ForFigure 3 Schematic diagram after the window where the data to be compressed is located slides.

[0033] Figure 5 For Figure 1 Another method for obtaining the data to be compressed used to calculate the hash value in the data compression method shown.

[0034] Figure 6 For Figure 5 Schematic diagram of the position after the window corresponding to the data to be compressed moves.

[0035] Figure 7 Shown is a flowchart of another data compression method provided by an exemplary embodiment of this specification.

[0036] Figure 8 For Figure 7 Process diagram of the implementation process of the data compression method shown.

[0037] Figure 9 Shown is a schematic diagram of a data format of a hash node in the data compression method provided by an exemplary embodiment of this specification.

[0038] Figure 10 Shown is a process diagram of the implementation process of another data compression method provided by an exemplary embodiment of this specification.

[0039] Figure 11 Shown is a schematic diagram of another data format of a hash node in the data compression method provided by an exemplary embodiment of this specification.

[0040] Figure 12 Shown is a flowchart of another data compression method provided by an exemplary embodiment of this specification.

[0041] Figure 13 Shown is a structural block diagram of a device for a data compression method provided by an exemplary embodiment of this specification.

[0042] Figure 14 Shown is a structural block diagram of another device for a data compression method provided by an exemplary embodiment of this specification.

[0043] Figure 15 Shown is a structural block diagram of another device for a data compression method provided by an exemplary embodiment of this specification. Detailed implementation manners

[0044] The following will be combined with the drawings in the embodiments of this specification to clearly and completely describe the technical solutions in the embodiments of this specification. Obviously, the described embodiments are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this specification.

[0045] With the continuous growth of cloud storage services and the amount of data on the cloud, data compression technology has received widespread attention and application. Compressing data can not only effectively reduce storage costs, but also alleviate the bandwidth bottleneck problem in the process of data migration under the current distributed cloud architecture. Data compression can be divided into two categories: lossy compression and lossless compression. Lossy compression is mainly aimed at images, audio and video, etc., and has limited application scenarios, while lossless compression can be used in any scenario, so it has better applicability. Among the many implementation methods of lossless compression, the LZ algorithm has been more widely used due to its high encoding and decoding speed, but the compression rate of the LZ algorithm is relatively low. Therefore, how to maintain the high decoding speed of LZ while further improving the compression rate has become a technical problem that needs to be solved by the LZ algorithm. The LZ algorithm can be implemented based on software or hardware. The software-based implementation of the LZ algorithm will occupy a lot of CPU resources and is limited by the CPU processing speed; for the hardware-based implementation of the LZ algorithm, if you want to achieve the same compression rate as the software implementation, you will consume a lot of hardware resources. When the hardware resources are limited, the data compression rate is relatively low.

[0046] At least one embodiment of the present specification provides a data compression method and a device for the data compression method to at least solve the above technical problems, as described below.

[0047] In at least one embodiment of this specification, Figure 1 As shown, the data compression method may include the following steps S100 to S500. The data compression method may be implemented by hardware.

[0048] S100, obtaining a plurality of to-be-compressed data from a to-be-compressed data object based on different preset ranges.

[0049] like Figure 2As shown, before compressing the data, the data is divided into multiple data objects to be compressed for caching (input data caching), and then these data objects to be compressed are input into the hardware, and the hardware compresses these data objects to be compressed sequentially. For example, 1000KB of original data can be split into multiple data objects to be compressed such as 4KB and 16KB. When compressing 1000Kb of data, the 4KB data object to be compressed can be compressed first, and then the 16KB data object to be compressed can be compressed until all the data objects to be compressed are compressed.

[0050] The preset range may include length and initial position. As Figure 3 shown, for the data of each data object to be compressed, based on the preset range, the data to be compressed A (for example, the data from byte 0 to byte n - 1) and the data to be compressed B (for example, the data from byte 1 to byte n) are obtained. The preset ranges corresponding to the data to be compressed A and the data to be compressed B are equal (both are the length occupied by byte n), and the initial position of the data to be compressed A corresponds to byte 0, and the initial position of the data to be compressed B corresponds to byte 1. That is, the lengths of the preset ranges corresponding to the data to be compressed A and the data to be compressed B are equal but the initial positions are different.

[0051] It should be noted that in at least one embodiment of this specification, the preset ranges being different may mean that at least one of the length and the initial position is different. That is, based on different preset ranges, at least two data to be compressed are obtained from the data object to be compressed, and the lengths of the preset ranges of the at least two data to be compressed are the same and the initial positions are different; and / or, based on different preset ranges, at least two data to be compressed are obtained from the data object to be compressed, and the lengths of the preset ranges of the at least two data to be compressed are different and the initial positions are the same.

[0052] For example, as Figure 3 and Figure 4 shown, the starting positions of the preset ranges corresponding to the data to be compressed A and the data to be compressed B are different and the lengths are the same.

[0053] For example, as Figure 5 shown, the lengths of the preset ranges corresponding to the data to be compressed A and the data to be compressed C are different and the initial positions are the same.

[0054] S200, generate multiple hash values respectively based on multiple data to be compressed.

[0055] As Figure 2 and Figure 3As shown, a hash value A is generated based on the data A to be compressed, and a hash value B is generated based on the data B to be compressed. The hash values A and B are used to respectively search in the hash tables A and B to check if there are corresponding hash nodes. In other words, the hash value A is used to search in the corresponding hash table A to check if there is a hash node corresponding to the hash value A. Correspondingly, the hash value B is used to search in the corresponding hash table B to check if there is a hash node corresponding to the hash value B.

[0056] It should be noted that the window defined by the preset range slides along the characters of the data object to be compressed. As Figure 3 and Figure 4 shown, Figure 3 After the window where the data A to be compressed is located in Figure 4 slides multiple times, it moves to the position shown in Figure 3 . At this time, the data of the data A to be compressed corresponds to the data from byte n - 1 to byte m - 1. Correspondingly, Figure 4 After the data B to be compressed in

[0057] slides the same number of times and the same distance as the data A to be compressed, it moves to the position shown in

[0058] At this time, the data of the data B to be compressed corresponds to the data from byte n to byte m.

[0059] It should be noted that in this specification, the "data to be compressed", also known as the data to be encoded, is the data for which the corresponding hash value is currently to be searched in the hash table, that is, the data within the window defined by the preset range in the data object to be compressed at the current stage of searching for the hash value. And the "compressed data", also known as the encoded data, is relative to the data for which the corresponding hash value is currently to be searched in the hash table (data to be compressed). It can be the data part before the current data to be compressed, and these data have already been searched in the hash table through the corresponding hash value. The data compression of the entire data object to be compressed can be performed in the subsequent step S500. For example, as Figure 4As shown in the figure, for the hash table A, the data in bytes n - 1 to m - 1 in the window A within the preset range is the data to be compressed, while the data before byte n - 1 corresponding to the window L1 has been searched in the hash table A. The part of the data to be compressed object located in the window L1 is called the compressed data. Correspondingly, for the hash table B, the data in bytes n to m in the window B within the preset range is the data to be compressed, while the data corresponding to the window L2 has been searched in the hash table B. The part of the data to be compressed object located in the window L2 is the compressed data.

[0060] When searching for the corresponding hash value in the hash table based on the hash value corresponding to the data to be compressed, if the corresponding hash value is found in the hash table, update the information (including location information) of the current data to be compressed to the hash table and associate it with the hash value to establish a hash node. It should be noted that under the condition of ignoring hash collisions, the above situation indicates that there is data in the compressed data that matches the current data to be compressed.

[0061] S400, obtain the matching length of the data corresponding to the hash node based on the data to be compressed and the compressed data.

[0062] The hash node records the location information of the data. Based on the hash node, data with the same coincidence degree can be found from the data to be compressed object. For the coincident data, by recording the location information where it appears and the length of the overlapping part, the subsequent coincident data can be obtained based on the previously appeared coincident data. That is, the subsequent appeared coincident data can be replaced by information such as location information and matching length, so as to realize the compression of the compressed package data.

[0063] Refer to the following example string. There are two hash nodes corresponding to data both being 05060708. Thus, starting from the positions corresponding to these two hash nodes, calculate whether the subsequent data can match, so as to obtain the coincident data 0506070800000000090a03040201 (set as the target string). Thus, by comparing the offset distance (offset) of these two target strings, the relative position of the second target string relative to the first target string can be obtained (equal to obtaining the location information of the second string). Thus, the second target string can be replaced based on the first appeared target string, the offset distance and the matching length of the second appeared target string. The subsequent appeared target strings are processed with reference to the second appeared target string. Thus, in the subsequent step S500, the subsequent strings can be replaced with relatively less information (such as matching length, offset, etc.), so that the data of the data to be compressed object can be compressed.

[0064] Example string:

[0065]

[0066] S500 compresses the data in the data object to be compressed based on hash nodes and the matching length. For example, in subsequent processes, the compressed data can be further encapsulated.

[0067] In this specification, before compressing and encapsulating the data object to be compressed, step S500 may further include: determining whether the data to be compressed corresponding to the hash node is located at the termination position of the data object to be compressed; if so, encapsulating the compressed data and ending the compression process; if not, in the data object to be compressed, shifting each window defined by a preset range backward, and using the data covered by the window as the subsequent data to be compressed. In this embodiment, each window defined by all preset ranges moves the same distance each time. Exemplarily, after finding the hash values corresponding to the data to be compressed A and the data to be compressed B in Figure 3 it is determined that the data to be compressed A and the data to be compressed B are not located at the termination position of the data object to be compressed, so the data to be compressed A and the data to be compressed B are slid (at least once) based on the window corresponding to the preset range to move to the position shown in Figure 4 and the above step S200 is re-executed. After sliding, the starting positions of the data to be compressed A and the data to be compressed B are still adjacent, that is, the sliding distances of the data to be compressed A and the data to be compressed B are equal.

[0068] In at least one embodiment of this specification, the above step S500 may include: before encapsulating the data in the data object to be compressed, determining whether the data to be compressed is at the end of the data object to be compressed; if so, performing data encapsulation and output; if not, shifting the window corresponding to each preset range backward to obtain multiple data to be compressed in the next compression process from the data object to be compressed.

[0069] According to the embodiments of this specification, by setting multiple hash tables, so that the hash nodes used to compress the data object to be compressed can be selected based on the multiple hash tables, the impact of hash conflicts on the compression ratio can be reduced, and in the subsequent process of calculating the matching length of the data, hash nodes with relatively larger matching lengths can be selected to compress the data of the data object to be compressed, so as to improve the compression speed and the data compression ratio.

[0070] In the embodiments of this specification, there is no limit to the designed number of hash tables, which can be designed according to actual needs. Next, for several designed numbers of hash tables, the solutions of the data compression method will be further described.

[0071] In at least one embodiment of this specification, the above step S100 may include: obtaining four data to be compressed from the data object to be compressed based on two different preset ranges. In this embodiment, the two data to be compressed are respectively the odd-position data block and the even-position data block. The serial number of the byte corresponding to the starting position of the odd-position data block in the data object to be compressed is odd, and the serial number of the byte corresponding to the starting position of the even-position data block in the data object to be compressed is even, and the starting positions of the odd-position data block and the even-position data block are adjacent. Exemplarily, referring back to Figure 3 , the starting position of the data A to be compressed is at byte 0 (in the first position, i.e., odd position), so it can also be called the odd-position data block A. The starting position of the data B to be compressed is at byte 1 (in the second position, i.e., even position), so it can also be called the even-position data block B. In this way, in the entire data object to be compressed, the data to be compressed starting from odd positions and even positions can be hashed simultaneously to increase the probability of finding the corresponding hash value in the hash table. In addition, compared with hashing the data to be compressed starting from only odd positions or only even positions, this method of selecting hash nodes can reduce the impact of hash collisions on the compression rate, and the hash nodes selected in this way may correspond to longer matchable strings, so that in the process of calculating the match length of the data, there is a greater probability of selecting a hash node with a relatively larger match length to compress the data of the data object to be compressed. The specific explanation is as follows.

[0072] Referring to the following example string, assume that the length of the preset ranges corresponding to the odd-position data block and the even-position data block is 4 characters. Assuming that the data corresponding to the odd-position data block is the string 8a8a0506, the same string is not found in the previous data, that is, the hash value corresponding to the string 8a8a0506 is not found in the hash table. In this case, the data corresponding to the even-position data block is the second string 05060708, and the same string 05060708 can be found in the previous data, that is, the hash value corresponding to the second string 05060708 can be found in the corresponding hash table to increase the compression rate of the data object to be compressed. In addition, in this case, when calculating the subsequent match length based on the hash node corresponding to the second string 05060708, a longer repeatable string 0506070800000000090a03040201 will be obtained, thereby further increasing the compression rate of the data object to be compressed.

[0073] Example string:

[0074] ……09090a0304 05060708 00000000090a030402018a8a8a 0506070800000000090a030402019999……

[0075] In at least one embodiment of this specification, four hash tables can be designed. Thus, the above step S100 may include: obtaining four data to be compressed from the data object to be compressed based on four different preset ranges. As Figure 5 shown, the four data to be compressed are the short data block A at odd positions, the short data block B at even positions, the long data block C at odd positions, and the long data block D at even positions. The serial numbers of the bytes corresponding to the starting positions of the short data block A at odd positions and the long data block C at odd positions in the data object to be compressed are odd, and the serial numbers of the bytes corresponding to the starting positions of the short data block B at even positions and the long data block D at even positions in the data object to be compressed are even. The lengths of the short data block A at odd positions and the short data block B at even positions are the same, the lengths of the long data block C at odd positions and the long data block D at even positions are the same, the starting positions of the short data block A at odd positions and the long data block C at odd positions are the same, the starting positions of the short data block B at even positions and the long data block D at even positions are the same, and the starting positions of the short data block A at odd positions and the short data block B at even positions are adjacent. In the above solution, the four data to be compressed correspond to all cases with different preset ranges, thereby further increasing the probability of finding the corresponding hash value from the hash table, and making it more likely to select a hash node with a relatively larger matching length to compress the data of the data object to be compressed during the process of the matching length of the data.

[0076] In an implementation manner of the first aspect of this specification, the distance that the window moves each time is the length occupied by two bytes. For example, taking the four data to be compressed as shown in Figure 5 as an example, after the window moves once, the positions corresponding to the four data to be compressed are as shown in Figure 6 shown. The starting positions of the short data block A at odd positions and the long data block C at odd positions become byte 2, and the starting positions of the short data block at even positions and the long data block D at even positions will become byte 3.

[0077] In at least one embodiment of this specification, as shown in Figure 7 shown, the data compression method may further include the following step S600.

[0078] S600, in the case of finding the corresponding hash value in at least two hash tables, based on the preset node selection rule, screening out the target node from all the obtained hash nodes.

[0079] In the case where step S600 appears in the data compression method, the above step S400 may include: obtaining the matching length of the data corresponding to the target node based on the data to be compressed and the compressed data. In this way, one of the multiple hash nodes found in the hash table is selected as the target node, and then the subsequent calculation of the matching length of the data is performed, reducing the calculation amount of the subsequent matching length of the data and reducing the compression delay at the same time.

[0080] In at least one embodiment of the present specification, the node selection rule may include: comparing the position information of the hash node whose hash value is found in the hash table with the position information of the previously selected hash node, and taking the hash node whose position information is inconsistent with the position information of the previously selected hash node as the target node. In this way, in the process of calculating the matching length of the data, it is possible to avoid loss of compression ratio caused by repetition or a large amount of repetition of the data to be compressed obtained based on two consecutive hash nodes, as follows.

[0081] Referring to the following example strings, for the case where the length corresponding to the preset range is 4 bytes, there are two strings 05060708 that can be matched, and there are three strings 090a0304 that can be matched. In the subsequent calculation of the matching length, the matchable string obtained from the hash node corresponding to the second string 05060708 is 0506070800000000090a03040201 (this string appears repeatedly). However, the matchable string obtained from the hash node corresponding to the third string 090a0304 is 090a03040201. Obviously, the length of the matchable string 090a03040201 is much smaller than the length of the matchable string 0506070800000000090a03040201, and the matchable string 0506070800000000090a03040201 includes the matchable string 090a03040201, that is, there is a situation of repeated matching. In this case, it is not necessary to select the hash node corresponding to the string 090a0304 as the target node. In this situation, due to the design of different preset ranges, the matchable strings corresponding to the hash nodes of other hash tables may be different. Therefore, it is possible to select the hash nodes in other hash tables as the target nodes, thereby reducing the probability of repeated matching.

[0082] Example strings:

[0083]

[0084] In at least one embodiment of this specification, the information of the data format corresponding to the hash node includes location information and a valid flag bit. The optional flags of the valid flag bit include a first flag and a second flag. The first flag is used to indicate that the hash node is invalid, and the second flag is used to indicate that the hash node is valid. In this case, the above step S200 may include the following step S210: If the valid flag bit corresponding to the hash node found from the hash table is the first flag, generate a hash node with the data to be compressed; if the valid flag bit corresponding to the hash node found from the hash table is the second flag, use the hash node corresponding to the data to be compressed as the hash node of the preselected target node. Among them, if the valid flag bits corresponding to the hash nodes found from at least two hash tables are the second flag, filter out the target node through the node selection rule; if the valid flag bit corresponding to the hash node found from only one hash table is the second flag, use the found hash node as the target node. For example, the first flag may be denoted as 0, and the second flag may be denoted as 1.

[0085] In at least one embodiment of this specification, the above step S100 may further include: initializing the hash nodes corresponding to the hash table so that the valid flag bits in the data format corresponding to the hash nodes are written as the first flag. In this way, it is possible to avoid the hash table updated when compressing the data object to be compressed in the previous process from affecting the compression process of the current data object to be compressed.

[0086] In at least one embodiment of this specification, the above step 300 may further include: writing the first flag bit corresponding to each hash node as the second flag, and updating the information in the data format corresponding to each hash node to the hash table based on the hash value corresponding to each hash node.

[0087] Next, in combination with Figure 8 the flowchart shown below, for the case where the data format of the hash node includes a valid flag bit, the data compression process will be described. The process is specifically as follows S11 to S17.

[0088] S11, Initialize the hash nodes in the hash table so that the valid flag bits in its data format are written as 0 (the first flag), and then input the data object to be compressed.

[0089] S12, Obtain multiple data blocks based on different preset ranges, and calculate the hash values corresponding to each data block.

[0090] S13, Search the hash table based on the hash value, update the hash value to the hash table, and at the same time write the valid flag bits in the data format of the hash nodes corresponding to each data block as 1 (the second flag)

[0091] S14. Determine whether the valid flag bit of the hash node found in the hash table is 1 to determine whether data that can match the current data block can be found through the hash table. Specifically, it can be divided into the following situations.

[0092] If the valid flags of the hash nodes corresponding to the hash values found in each hash table are all 0, that is, data that matches the current data block is not found, and the data of the current data block appears for the first time in the data object to be compressed, then use the hash node corresponding to the current data block as the target node, and then jump to S16 below.

[0093] If only the valid flag bit of the hash node corresponding to the hash value found in one hash table is 1, and the valid flags of the hash nodes corresponding to the hash values found in other hash tables are all 0, that is, the number of hash nodes with a valid flag bit of 1 is not greater than 1, then use the hash node corresponding to the valid flag bit of 1 as the target node, and then jump to S15 below.

[0094] If there are at least two hash nodes corresponding to the hash values found in the hash tables whose valid flag bits are 1, that is, the number of hash nodes with a valid flag bit of 1 is greater than 1, select the target node from these hash nodes based on the node selection rules mentioned above, and then jump to S15 below.

[0095] S15. When the valid flag bits of the hash nodes corresponding to the hash values found in at least one hash table are 1, calculate the matching length of the target node.

[0096] S16. Determine whether the current data to be compressed (data block) is at the termination position of the data object to be compressed. If the judgment result is yes, jump to S17 below. If the judgment result is no, move the data block backward (for example, move two bytes) to re-enter step S12 above.

[0097] S17. Under the condition that in step S16, it is satisfied that the current data to be compressed (data block) is at the termination position of the data object to be compressed, compress and encapsulate the data of the data object to be compressed for output.

[0098] In the embodiments of this specification, check information can be added to the data format of the hash node to reduce the risk of hash collision during the process of finding the hash node in the hash table. For example, Figure 9As shown, the information on the data format of the hash node may further include first verification information. For the same hash node, the data used to generate the first verification information and the hash value is the same, but the calculation rules are different. In this embodiment, step S200 described above may include: when the valid flag bit corresponding to the hash node found from the hash table is the second flag, comparing the first verification information corresponding to the hash node corresponding to the data to be compressed with the first verification information corresponding to the hash node found from the hash table; if the comparison result is different, generating a hash node with the data to be compressed; if the comparison result is the same, using the hash node corresponding to the data to be compressed as the hash node of the preselected target node. In this way, by verifying the first verification information, there is a greater chance of screening out the hash nodes with hash collisions when looking up the hash value from the hash table, so as to improve the compression ratio of data compression.

[0099] Next, with reference to Figure 10 the flowchart shown, for the case where the data format of the hash node includes a valid flag bit, the process of data compression will be described. The process is specifically as follows, S21 to S27.

[0100] S21 to S23 can refer to S11 to S13 described above, and will not be elaborated here.

[0101] S24, determine whether the valid flag bit of the hash node found from the hash table is 1, and combine the first verification information to determine whether data that can match the current data block can be found through the hash table. Specifically, it can be divided into the following several cases.

[0102] If the valid flag of each hash node corresponding to the hash value found from the hash table is 0, that is, data that matches the current data block is not found, and the data of the current data block appears for the first time in the data object to be compressed, then use the hash node corresponding to the current data block as the target node, and then jump to S26 below.

[0103] If there is at least one hash node corresponding to the hash value found from the hash table whose valid flag bit is 1, then compare the first verification information of the hash node corresponding to the current data block with the first verification information of the hash node found from the hash table. If the verification information of the two is different, it indicates that data that matches the current data block is not found, and the data of the current data block appears for the first time in the data object to be compressed, then use the hash node corresponding to the current data block as the target node, and then jump to S26 below; if the verification information of the two is the same, screen out the hash nodes that meet this condition and make the following two choices.

[0104] If only the valid flag bit of the hash node corresponding to the hash value found in one hash table is 1, the valid flags of the hash nodes corresponding to the hash values found in other hash tables are all 0, and the number of hash nodes with the valid flag bit of 1 is not greater than 1, then use the hash node corresponding to the valid flag bit of 1 as the target node, and then jump to S25 below.

[0105] If there are at least two hash nodes corresponding to the hash values found in the hash tables with the valid flag bit of 1, that is, the number of hash nodes with the valid flag bit of 1 is greater than 1, select the target node from these hash nodes based on the node selection rule mentioned above, and then jump to S25 below.

[0106] S25 to S27 can refer to S15 to S17 above and will not be elaborated here.

[0107] In at least one embodiment of this specification, as Figure 11 shown, the information of the data format further includes second verification information. The second verification information corresponding to the previous hash node is generated by calculating the data corresponding to the next hash node, and the calculation rules for generating the second verification information and the hash value are different. In this embodiment, step S200 above may further include: when the valid flag bit corresponding to the hash node found from the hash table is the second flag, after comparing the hash node with the first verification information, if there are at least two hash nodes as preselected target nodes, compare the second verification information corresponding to the hash node corresponding to the data to be compressed with the second verification information corresponding to the hash node found from the hash table, and use the hash node in the case where the comparison result is consistent as the hash node of the preselected target node. In this way, the position corresponding to the hash node screened by verifying the second verification information can obtain a relatively longer string in the subsequent process of calculating the matching length of the data, so as to improve the compression ratio of the data object to be compressed, as follows.

[0108] Refer to the following example strings. For the case where the length corresponding to the preset range is 4 bytes, there are three strings 05060708 that can be matched, and there are three strings 090a0304 that can be matched. For the data block corresponding to the string 05060708, after shifting the current data block two characters backward, in the subsequent data blocks, the window of the preset range at the first string 05060708 will cover 07080011, while the windows of the preset range at the second and third strings 05060708 will cover 07080000. The second check information calculated from 70080011 corresponds to the hash node based on the first string 05060708, and the second check information calculated from 07080000 corresponds to the hash nodes based on the second and third strings 05060708. Obviously, the second check information of the hash node corresponding to the first string 05060708 is different from the second check information of the hash node corresponding to the second string 05060708, and the second check information of the hash node corresponding to the second string 05060708 is the same as the second check information of the hash node corresponding to the third string 05060708. Thus, the hash node corresponding to the second string 05060708 can be selected as the target node, so that in the subsequent process of calculating the matching length of the data, a relatively longer string can be obtained.

[0109] Example strings:

[0110]

[0111]

[0112] In the embodiments of this specification, the applicant performed simulation calculations on the data compression method based on the silesia dataset. Among them, in the data compression method used for simulation, the hash table is designed as Figure 5 and Figure 6 shown in four cases, and the length of the preset range of the short data block A in the odd position and the short data block B in the even position is 4 bytes, and the length of the preset range of the long data block C in the odd position and the long data block D in the even position is 6 bytes. In this case, based on the above node selection rule, the data compression rate is increased by 0.26%. In addition, when the data format includes the above first check information and second check information, the compression rate will also increase by 2%.

[0113] At least one embodiment of this specification provides a data compression method, as Figure 12 shown, and this data compression method may include the following steps S121 to S126.

[0114] S121. Obtain the data to be compressed within a preset range from the data object to be compressed.

[0115] S122. Generate a hash value based on the data to be compressed.

[0116] S123. Search the hash table based on the hash value to obtain the position information of the compressed data corresponding to the hash value, and generate a hash node corresponding to the data to be compressed. The hash table is used to store the hash nodes corresponding to the compressed data, the position information of the hash node in the data object to be compressed, and the verification information of the compressed data. The hash node corresponds to the position information, and the calculation rules for generating the hash value and the first verification information are different.

[0117] S124. Generate the first verification information of the hash node corresponding to the data to be compressed based on the data to be compressed.

[0118] S125. When the first verification information is consistent with the second verification information, obtain the matching length of the data corresponding to the hash node based on the data to be compressed and the compressed data.

[0119] S126. Compress the data in the data object to be compressed based on the hash node and the matching length.

[0120] For the implementation manners of the above steps S121 to S126, reference may be made to the relevant descriptions in the foregoing embodiments. Among them, by verifying the first verification information, when searching for the hash value in the hash table, it is possible to more likely screen out the hash nodes with hash conflicts, so as to improve the compression ratio of data compression.

[0121] At least one embodiment of this specification provides a device for a data compression method, as Figure 13 shown. The device includes an acquisition module 10, a generation module 20, a search module 30, a matching length calculation module 40, and a packaging module 50. The acquisition module 10 is configured to obtain multiple pieces of data to be compressed from the data object to be compressed based on different preset ranges. The preset ranges include length and initial position. The generation module 20 is configured to generate multiple hash values respectively based on the multiple pieces of data to be compressed. The search module 30 is configured to search the hash table corresponding to each hash value and obtain the position information of the compressed data corresponding to the hash value. The hash table is used to store the position information of the compressed data and the data to be compressed in the data object to be compressed, and the position information corresponds to the hash node. The matching length calculation module 40 is configured to obtain the matching length of the data corresponding to the hash node based on the data to be compressed and the compressed data. The packaging module 50 is configured to compress the data in the data object to be compressed based on the hash node and the matching length. The implementation manners of each module in the device may refer to the relevant descriptions in the foregoing embodiments of the data compression method, and will not be elaborated here.

[0122] In the embodiments of this specification, the data compression method is implemented through a hardware circuit. Thus, in the apparatus for the data compression method, each module such as the acquisition module 10, the generation module 20, the search module 30, the match length calculation module 40, and the encapsulation module 50 can be each circuit module in the hardware circuit.

[0123] In the apparatus provided in at least one embodiment of this specification, the acquisition module 10 can also be used to: obtain at least two data to be compressed from the data object to be compressed based on different preset ranges, and the lengths of the preset ranges of the at least two data to be compressed are the same and the initial positions are different; and / or, obtain at least two data to be compressed from the data object to be compressed based on different preset ranges, and the lengths of the preset ranges of the at least two data to be compressed are different and the initial positions are the same.

[0124] In the apparatus provided in at least one embodiment of this specification, the acquisition module 10 can also be used to: obtain four data to be compressed from the data object to be compressed based on two different preset ranges. In this embodiment, the two data to be compressed are respectively an odd-position data block and an even-position data block. The serial number of the byte corresponding to the starting position of the odd-position data block in the data object to be compressed is odd, the serial number of the byte corresponding to the starting position of the even-position data block in the data object to be compressed is even, and the starting positions of the odd-position data block and the even-position data block are adjacent.

[0125] In the apparatus provided in at least one embodiment of this specification, the acquisition module 10 can also be used to: obtain four data to be compressed from the data object to be compressed based on four different preset ranges. In this embodiment, the four data to be compressed are respectively an odd-position short data block, an odd-position long data block, an even-position short data block, and an even-position long data block. The serial numbers of the bytes corresponding to the starting positions of the odd-position short data block and the odd-position long data block in the data object to be compressed are odd, the serial numbers of the bytes corresponding to the starting positions of the even-position short data block and the even-position long data block in the data object to be compressed are even. The lengths of the odd-position short data block and the even-position short data block are the same, the lengths of the odd-position long data block and the even-position long data block are the same. The starting positions of the odd-position short data block and the odd-position long data block are the same, the starting positions of the even-position short data block and the even-position long data block are the same, and the starting positions of the odd-position short data block and the even-position short data block are adjacent.

[0126] In the device provided by at least one embodiment of this specification, the encapsulation module 50 can also be used to: determine whether the data to be compressed corresponding to the hash node is located at the termination position of the data object to be compressed; if so, encapsulate the compressed data and end the compression process; if not, in the data object to be compressed, move the window defined by each preset range backward, and use the data covered by the window as the subsequent data to be compressed. In this embodiment, the distance that all the windows defined by the preset ranges move each time is the same.

[0127] In the device provided by at least one embodiment of this specification, the distance that the window moves each time is the length occupied by two bytes.

[0128] As Figure 14 shown, the device provided by at least one embodiment of this specification may further include a node selection module 60. The node selection module 60 is used to: when the corresponding hash value is found in at least two hash tables, based on a preset node selection rule, screen out the target node from all the obtained hash nodes. In this embodiment, the step of obtaining the matching length of the data corresponding to the hash node based on the data to be compressed and the compressed data may include: obtaining the matching length of the data corresponding to the target node based on the data to be compressed and the compressed data.

[0129] In the device provided by at least one embodiment of this specification, the node selection rule may include: comparing the position information of the hash node whose hash value is found in the hash table with the position information of the previously selected hash node, and taking the hash node whose position information is inconsistent with the position information of the previously selected hash node as the target node.

[0130] In the device provided by at least one embodiment of this specification, the information on the data format corresponding to the hash node includes position information and a valid flag bit. The optional flags of the valid flag bit include a first flag and a second flag. The first flag is used to indicate that the hash node is invalid, and the second flag is used to indicate that the hash node is valid. In this case, the generation module 20 can be used to: if the valid flag bit corresponding to the hash node found in the hash table is the first flag, generate a hash node with the data to be compressed; if the valid flag bit corresponding to the hash node found in the hash table is the second flag, take the hash node corresponding to the data to be compressed as the hash node of the preselected target node. Among them, if the valid flag bits corresponding to the hash nodes found in at least two hash tables are the second flag, screen out the target node through the node selection rule; if the valid flag bit corresponding to the hash node found in only one hash table is the second flag, take the found hash node as the target node.

[0131] In the device provided by at least one embodiment of this specification, the obtaining module 10 may further be configured to: initialize a hash node corresponding to a hash table, so that a valid flag bit in the data format corresponding to the hash node is written as a first flag.

[0132] In the device provided by at least one embodiment of this specification, the searching module 30 may further be configured to: write a first flag bit corresponding to each hash node as a second flag, and update the information in the data format corresponding to each hash node to the hash table based on the hash value corresponding to each hash node.

[0133] In the device provided by at least one embodiment of this specification, the information of the data format may further include first check information. For the same hash node, the data used to generate the first check information and the hash value is the same but the calculation rules are different. In this case, the generating module 20 may further be configured to: when the valid flag bit corresponding to the hash node found from the hash table is the second flag, compare the first check information corresponding to the hash node corresponding to the data to be compressed with the first check information corresponding to the hash node found from the hash table; if the comparison result is different, generate a hash node with the data to be compressed; if the comparison result is consistent, use the hash node corresponding to the data to be compressed as the hash node of the preselected target node.

[0134] In the device provided by at least one embodiment of this specification, the information of the data format further includes second check information. The second check information corresponding to the previous hash node is generated by calculating the data corresponding to the next hash node, and the calculation rules for generating the second check information and the hash value are different. In this case, the generating module 20 may further be configured to: when the valid flag bit corresponding to the hash node found from the hash table is the second flag, after comparing the hash nodes through the first check information, if there are at least two hash nodes as preselected target nodes, compare the second check information corresponding to the hash node corresponding to the data to be compressed with the second check information corresponding to the hash node found from the hash table, and use the hash node with the consistent comparison result as the hash node of the preselected target node.

[0135] In the device provided by at least one embodiment of this specification, the encapsulating module 50 may further be configured to: before encapsulating the data in the data object to be compressed, determine whether the data to be compressed is at the end of the data object to be compressed; if so, perform data encapsulation and output; if not, move the window corresponding to each preset range backward to obtain multiple data to be compressed in the next compression process from the data object to be compressed.

[0136] This specification also provides a device for a data compression method, such as Figure 15As shown, the device may include an acquisition module 1, a generation module 2, a lookup module 3, a verification module 4, a matching length calculation module 5, and a packaging module 6. The acquisition module 1 is used to acquire the data to be compressed within a preset range from the data object to be compressed. The generation module 2 generates a hash value based on the data to be compressed. The lookup module 3 looks up a hash table based on the hash value to obtain the position information of the compressed data corresponding to the hash value, and generates a hash node corresponding to the data to be compressed. Among them, the hash table is used to store the hash nodes corresponding to the compressed data, the position information of the hash nodes in the data object to be compressed, and the verification information of the compressed data. The hash node corresponds to the position information, and the calculation rules for generating the hash value and the first verification information are different. The verification module 4 generates the first verification information of the hash node corresponding to the data to be compressed based on the data to be compressed. The matching length calculation module 5 obtains the matching length of the data corresponding to the hash node based on the data to be compressed and the compressed data when the first verification information is consistent with the second verification information. The packaging module 6 compresses the data in the data object to be compressed based on the hash node and the matching length. The implementation manners of the various modules in the device can refer to the relevant descriptions in the foregoing embodiments of the data compression method, which will not be elaborated here.

[0137] In the embodiments of this specification, the data compression method is implemented through a hardware circuit. Thus, in the device for the data compression method, each module such as the acquisition module 1, the generation module 2, the lookup module 3, the verification module 4, the matching length calculation module 5, and the packaging module 6 can be each circuit module in the hardware circuit.

[0138] In the description of this specification, the description with reference to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples" means that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples. In addition, without contradiction, those skilled in the art can combine and combine the different embodiments or examples described in this specification and the features of different embodiments or examples.

[0139] In addition, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" can explicitly or implicitly include at least one of the features. In the description of the present application, "a plurality" means at least two, such as two, three, etc., unless otherwise specifically defined.

[0140] Any process or method description, whether in a flowchart or otherwise described herein, can be understood to represent a module, segment, or portion of code including one or more executable instructions for implementing a specific logical function or process. The scope of the preferred embodiments of the present application includes additional implementations, where functions may be performed in a substantially simultaneous manner or in an order opposite to that shown or discussed, according to the functions involved, which should be understood by those skilled in the technical field to which the embodiments of the present application pertain.

[0141] The logic and / or steps represented in a flowchart or otherwise described herein, for example, can be considered a sequenced list of executable instructions for implementing a logical function and can be embodied specifically in any computer-readable medium for use by or in connection with an instruction execution system, apparatus, or device, such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in connection with the instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of the computer-readable medium include the following: an electrical connection having one or more wires (electronic device), a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable medium can even be paper or other suitable medium on which a program can be printed, as the program can be obtained, for example, electronically by optical scanning of the paper or other medium, followed by editing, interpretation, or other suitable processing as necessary, and then stored in a computer memory.

[0142] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0143] Those of ordinary skill in the art can understand that all or part of the steps carried out in implementing the above method embodiments can be completed by instructing relevant hardware through a program. The program can be stored in a computer-readable storage medium. When the program is executed, it includes one or a combination of the steps of the method embodiments.

[0144] In addition, in each of the embodiments of the present application, the functional units can be integrated into a processing module, or each unit can exist physically alone, or two or more units can be integrated into one module. The above integrated module can be implemented in the form of hardware or in the form of a software functional module. When the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0145] The storage medium mentioned above can be a read-only memory, a magnetic disk, an optical disk, etc. Although the embodiments of the present application have been shown and described above, it can be understood that the above embodiments are exemplary and should not be construed as limiting the present application. Those of ordinary skill in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present application.

[0146] The above are only the preferred embodiments of this specification and are not intended to limit this specification. Any modifications, equivalent substitutions, etc. made within the spirit and principles of this specification shall be included within the protection scope of this specification.

[0147] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this specification are all information and data that have been authorized by the user or fully authorized by all parties. The collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.

Claims

1. A data compression method, characterized in that, Including: Obtaining a plurality of data to be compressed from a data object to be compressed based on different preset ranges; Generating a plurality of hash values respectively based on the plurality of data to be compressed; Searching for a hash table corresponding to each hash value to obtain position information of the compressed data corresponding to the hash value, wherein the hash table is used to store position information of the compressed data and the data to be compressed in the data object to be compressed, and the position information corresponds to a hash node; Obtaining a matching length of data corresponding to the hash node based on the data to be compressed and the compressed data; Compressing the data in the data object to be compressed based on the hash node and the matching length.

2. The data compression method according to claim 1, wherein The obtaining a plurality of data to be compressed from a data object to be compressed based on different preset ranges includes: Obtaining at least two data to be compressed from the data object to be compressed based on different preset ranges, and the lengths of the preset ranges of the at least two data to be compressed are the same and the initial positions are different; and / or Obtaining at least two data to be compressed from the data object to be compressed based on different preset ranges, and the lengths of the preset ranges of the at least two data to be compressed are different and the initial positions are the same.

3. The data compression method according to claim 2, wherein The obtaining at least two data to be compressed from the data object to be compressed based on different preset ranges includes: Obtaining four data to be compressed from the data object to be compressed based on two different preset ranges; Wherein, two of the data to be compressed are an odd-position data block and an even-position data block respectively, the serial number of the byte corresponding to the starting position of the odd-position data block in the data object to be compressed is odd, the serial number of the byte corresponding to the starting position of the even-position data block in the data object to be compressed is even, and the starting positions of the odd-position data block and the even-position data block are adjacent.

4. The data compression method according to claim 2, wherein The obtaining at least two data to be compressed from the data object to be compressed based on different preset ranges includes: Obtaining four data to be compressed from the data object to be compressed based on four different preset ranges; Wherein, the four data to be compressed are an odd-position short data block, an odd-position long data block, an even-position short data block and an even-position long data block respectively, the serial numbers of the bytes corresponding to the starting positions of the odd-position short data block and the odd-position long data block in the data object to be compressed are odd, the serial numbers of the bytes corresponding to the starting positions of the even-position short data block and the even-position long data block in the data object to be compressed are even, the lengths of the odd-position short data block and the even-position short data block are the same, the lengths of the odd-position long data block and the even-position long data block are the same, and the starting positions of the odd-position short data block and the odd-position long data block are the same, the starting positions of the even-position short data block and the even-position long data block are the same, and the starting positions of the odd-position short data block and the even-position short data block are adjacent.

5. The data compression method according to claim 2, wherein Compressing the data in the data object to be compressed based on the hash node and the matching length includes: Determining whether the data to be compressed corresponding to the hash node is located at the termination position of the data object to be compressed; If so, encapsulating the compressed data and ending the compression process; If not, in the data object to be compressed, moving each window defined by the preset range backward, and using the data covered by the window as the subsequent data to be compressed; Wherein, all the windows defined by the preset range move the same distance each time.

6. The data compression method according to claim 5, characterized in that The distance that the window moves each time is the length occupied by two bytes.

7. The data compression method according to any one of claims 1 to 6, characterized in that It also includes: In the case where the corresponding hash values are found in at least two of the hash tables, based on a preset node selection rule, screening out target nodes from all the obtained hash nodes; Wherein, obtaining the matching length of the data corresponding to the hash node based on the data to be compressed and the compressed data includes: Obtaining the matching length of the data corresponding to the target node based on the data to be compressed and the compressed data.

8. The data compression method according to claim 7, wherein The node selection rule includes: Comparing the position information of the hash node whose hash value is found in the hash table with the position information of the previously selected hash node, and using the hash node whose position information is inconsistent with the position information of the previously selected hash node as the target node.

9. The data compression method according to claim 7, wherein The information on the data format corresponding to the hash node includes the position information and a valid flag bit. The optional flags of the valid flag bit include a first flag and a second flag. The first flag is used to indicate that the hash node is invalid, and the second flag is used to indicate that the hash node is valid. Among them, generating multiple hash values respectively based on the multiple data to be compressed includes: If the valid flag bit corresponding to the hash node found in the hash table is the first flag, generating the hash node with the data to be compressed; If the valid flag bit corresponding to the hash node found in the hash table is the second flag, using the hash node corresponding to the data to be compressed as the hash node for preselecting the target node. Among them, if the valid flag bits corresponding to the hash nodes found in at least two of the hash tables are the second flag, screening out the target node through the node selection rule. If only one of the hash tables has a valid flag bit corresponding to the hash node found being the second flag, using the found hash node as the target node.

10. The data compression method according to claim 9, wherein Obtaining multiple data to be compressed from the data object to be compressed based on different preset ranges includes: Initializing the hash node corresponding to the hash table so that the valid flag bit in the data format corresponding to the hash node is written as the first flag.

11. The data compression method according to claim 10, wherein Searching each hash table corresponding to each hash value and obtaining the position information of the compressed data corresponding to the hash value includes: Write the first flag bit corresponding to each of the hash nodes as the second flag, and based on the hash value corresponding to each of the hash nodes, update the information in the data format corresponding to each of the hash nodes to the hash table.

12. The data compression method according to claim 10, wherein The information in the data format further includes first check information. For the same hash node, the data used to generate the first check information and the hash value is the same but the calculation rules are different. Among them, the step of separately generating multiple hash values based on the multiple data to be compressed includes: When the valid flag bit corresponding to the hash node found from the hash table is the second flag, compare the first check information corresponding to the hash node corresponding to the data to be compressed with the first check information corresponding to the hash node found from the hash table; If the comparison result is different, generate the hash node with the data to be compressed; If the comparison result is consistent, use the hash node corresponding to the data to be compressed as the hash node of the preselected target node.

13. The data compression method according to claim 12, wherein The information in the data format further includes second check information. The second check information corresponding to the previous hash node is generated by calculating the data corresponding to the next hash node, and the calculation rules for generating the second check information and the hash value are different. Among them, the step of separately generating multiple hash values based on the multiple data to be compressed further includes: When the valid flag bit corresponding to the hash node found from the hash table is the second flag, after comparing the hash node by the first check information, if there are at least two hash nodes as the preselected target nodes, compare the second check information corresponding to the hash node corresponding to the data to be compressed with the second check information corresponding to the hash node found from the hash table, and use the hash node when the comparison result is consistent as the hash node of the preselected target node.

14. A data compression method, characterized in that, Includes: Obtain the data to be compressed within a preset range from the data object to be compressed; Generate a hash value based on the data to be compressed; Search the hash table based on the hash value, obtain the position information of the compressed data corresponding to the hash value, and generate a hash node corresponding to the data to be compressed. The hash table is used to store the hash nodes corresponding to the compressed data and the data to be compressed, the position information of the hash nodes in the data object to be compressed, and the first check information corresponding to the hash nodes corresponding to the compressed data. The hash node corresponds to the position information; Generate the first check information corresponding to the hash node corresponding to the data to be compressed based on the data to be compressed. The calculation rule for generating the hash value is different from the calculation rule for generating the first check information corresponding to the hash node corresponding to the data to be compressed; When the first check information of the hash node corresponding to the data to be compressed is consistent with the first check information of the hash node corresponding to the compressed data, obtain the matching length of the data corresponding to the hash node based on the data to be compressed and the compressed data; and Compress the data in the data object to be compressed based on the hash node and the matching length.

15. A data compression method as claimed in claim 1 or 14, characterized in that, The execution steps in the data compression method are implemented by a field programmable gate array.

16. A device for data compression, characterized in that, Including: An acquisition module, configured to obtain a plurality of data to be compressed from the data object to be compressed based on different preset ranges, where the preset ranges include length and initial position; A generation module, configured to generate a plurality of hash values respectively based on the plurality of data to be compressed; A search module, configured to search a hash table corresponding to each hash value, and obtain the position information of the compressed data corresponding to the hash value, where the hash table is used to store the position information of the compressed data and the data to be compressed in the data object to be compressed, and the position information corresponds to a hash node; A matching length calculation module, configured to obtain the matching length of the data corresponding to the hash node based on the data to be compressed and the compressed data; and A packaging module, configured to compress the data in the data object to be compressed based on the hash node and the matching length.

Citation Information

Patent Citations

  • Method for quickly realizing GZIP compression based on hardware and application thereof

    CN114157305A

  • Data processing method, device and equipment and readable storage medium

    CN115577149A