Data compression method and data decompression method

By introducing chunking processing and dual hash matching search technology into the data compression method, the problem of difficult to balance speed and rate in existing compression algorithms is solved, and efficient data compression and memory resource management are achieved.

CN120238134APending Publication Date: 2025-07-01SHENZHEN CORERAIN TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510222769.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

Existing compression algorithms are difficult to balance between pursuing high compression rates and high compression speeds, and there are problems such as large matching search overhead and low memory usage efficiency.

Method used

By introducing chunking processing technology into the data compression method, the data to be compressed is decomposed into multiple independent subtasks, and the compression process is carried out in parallel. The duplicate data between multiple blocks is identified by using the dual hash matching search technology to optimize memory resource allocation.

Benefits of technology

It significantly improves the data compression rate and compression speed, reduces data storage space, and improves memory usage efficiency, solving the problem that traditional algorithms are difficult to balance speed and rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120238134A_ABST
    Figure CN120238134A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of data processing, and provides a data compression method and a data decompression method.The data compression method comprises the steps that in response to a data compression instruction, data to be compressed are read from a data input stream to a main buffer area; when the data volume of the to-be-compressed data in the main buffer area reaches a preset data volume or a data end mark is detected, performing block processing on the to-be-compressed data in the main buffer area according to a preset size to obtain a plurality of blocks; and performing compression processing on each block in sequence to obtain a compression result of each block, and writing the compression result of each block into the temporary buffer area. The method can realize the technical effect of considering both the compression speed and the compression rate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of data processing, and more specifically, relates to a data compression method and a data decompression method. Background Art

[0002] In the digital age, data compression is crucial for storage and transmission. However, there are many deficiencies in current mainstream compression algorithms.

[0003] For dictionary-based compression algorithms, although they achieve compression by searching for repeated strings in historical data, they have slow compression speeds due to the need for a large-scale search and matching. For statistics-based compression algorithms, they rely on counting the frequencies of character occurrences to construct optimal prefix codes, but the compression ratio is limited when used alone. Although hybrid compression algorithms can obtain a higher compression ratio by integrating multiple technologies, their compression speed is significantly reduced. Thus, current high-compression-ratio algorithms have poor speeds, and fast compression algorithms have unsatisfactory compression ratios. It is difficult to achieve a balance between compression speed and compression ratio.

[0004] Moreover, traditional compression algorithms generally have the problem of large matching search overhead. When searching for repeated strings on a large scale, the computational complexity increases significantly. To ensure the operation of the algorithm, a large amount of status information and cache data need to be maintained, which not only causes huge memory overhead but also extremely low memory usage efficiency. Furthermore, most compression algorithm parameters are fixed, and it is difficult to flexibly adjust performance in the face of diverse actual scenarios, lacking the ability to handle different situations. Summary of the Invention

[0005] The purpose of the embodiments of the present invention is to provide a data compression method and a data decompression method, aiming to solve the technical problem that it is difficult to achieve a balance between the compression speed and compression ratio of existing compression algorithms.

[0006] To achieve the above purpose, according to the first aspect of the present invention, a data compression method is provided. The method includes:

[0007] In response to a data compression instruction, read the data to be compressed from the data input stream into the main buffer;

[0008] When the amount of data to be compressed in the main buffer reaches a predetermined amount of data or a data end flag is detected, perform block processing on the data to be compressed in the main buffer according to a predetermined size to obtain multiple blocks;

[0009] Perform compression processing on each block in sequence to obtain the compression result of each block, and write the compression result of each block into the temporary buffer respectively.

[0010] According to the second aspect of the present invention, a data decoding method is provided. The method includes:

[0011] In response to a data decompression instruction, read the data to be decompressed from the data input stream into the decompression buffer;

[0012] Perform decompression on the data to be decompressed in the decompression buffer to restore the symbol stream in the data to be decompressed;

[0013] Perform an inverse symbol sorting transformation process on the symbol stream in the data to be decompressed, and parse to obtain matching information or original characters;

[0014] If the matching information is parsed, locate the source data position corresponding to the data to be decompressed according to the matching information, and write the source data at the source data position as the decompressed data into the output data stream;

[0015] If the original characters are parsed, write the original characters as the decompressed data into the output data stream.

[0016] According to the third aspect of the present invention, a data compression device is provided, and the device includes:

[0017] A first reading unit, configured to read the data to be compressed from the data input stream into the main buffer in response to a data compression instruction;

[0018] A block processing unit, configured to, when the data volume of the data to be compressed in the main buffer reaches a predetermined data volume or a data end flag is detected, perform block processing on the data to be compressed in the main buffer according to a predetermined size to obtain a plurality of blocks;

[0019] A compression processing unit, configured to sequentially perform compression processing on each block to obtain a compression result of each block, and write the compression result of each block into the temporary buffer respectively.

[0020] The third aspect and any implementation manner of the third aspect respectively correspond to the first aspect and any implementation manner of the first aspect. For the technical effects corresponding to the third aspect and any implementation manner of the third aspect, reference can be made to the technical effects corresponding to the first aspect and any implementation manner of the first aspect above, and details are not described herein again.

[0021] According to the fourth aspect of the present invention, a data decompression device is provided, and the device includes:

[0022] A second reading unit, configured to read the data to be decompressed from the data input stream into the decompression buffer in response to a data decompression instruction;

[0023] A decompression processing unit, configured to perform decompression on the data to be decompressed in the decompression buffer to restore the symbol stream in the data to be decompressed;

[0024] A parsing processing unit for performing an inverse symbol sorting transformation on a symbol stream in the data to be decompressed to parse and obtain matching information or original characters;

[0025] A data output unit, if the matching information is parsed, for locating the source data position corresponding to the data to be decompressed according to the matching information, and writing the source data at the source data position as decompressed data into the output data stream; and if the original characters are parsed, for writing the original characters as decompressed data into the output data stream.

[0026] The fourth aspect and any two implementation manners of the fourth aspect respectively correspond to the second aspect and any two implementation manners of the second aspect. For the technical effects corresponding to the fourth aspect and any two implementation manners of the fourth aspect, reference may be made to the technical effects corresponding to the second aspect and any two implementation manners of the second aspect above, which will not be elaborated herein.

[0027] According to a fifth aspect of the present invention, there are provided two types of electronic devices, including: a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the electronic device implements the method according to any one of the above.

[0028] According to a sixth aspect of the present invention, there is provided a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the method according to any one of the above is implemented.

[0029] According to a seventh aspect of the present invention, there is provided a computer program product, and when the computer program product runs on an electronic device, the electronic device is caused to execute the method according to any one of the above in the first aspect.

[0030] It can be understood that for the beneficial effects of the above second aspect to seventh aspect, reference may be made to the relevant descriptions in the above first aspect, which will not be elaborated herein.

[0031] An embodiment of the present invention provides a data compression method and a data decompression method. The data compression method reads data to be compressed from a data input stream into a main buffer in response to a data compression instruction; when the data volume of the data to be compressed in the main buffer reaches a predetermined data volume or a data end flag is detected, the data to be compressed in the main buffer is block-processed according to a predetermined size to obtain a plurality of blocks; compression processing is sequentially performed on each block to obtain a compression result for each block, and the compression result for each block is respectively written into a temporary buffer.

[0032] Through the examples of the present invention, by reasonably setting the main buffer and the temporary buffer, and adopting the method of block processing, the main buffer and the temporary buffer are used for data storage and processing, memory resources are reasonably allocated, and excessive memory occupation is avoided. Moreover, by adopting the block processing method, the data processing task is decomposed into multiple independent subtasks, and compression processing is performed in parallel. The compression processing of each block can be completed independently, thereby significantly improving the data compression ratio and data compression speed, and reducing the data storage space. BRIEF DESCRIPTION OF THE DRAWINGS

[0033] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for use in the embodiments or the description of the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings.

[0034] Figure 1 is a flowchart showing a data compression method provided by an embodiment of the present invention;

[0035] Figure 2 is a flowchart showing an alternative data compression method provided by an embodiment of the present invention;

[0036] Figure 3 is a flowchart showing an alternative data compression method provided by an embodiment of the present invention;

[0037] Figure 4 is a flowchart showing an alternative data compression method provided by an embodiment of the present invention;

[0038] Figure 5 is a flowchart showing an alternative data compression method provided by an embodiment of the present invention;

[0039] Figure 6 is a flowchart showing an alternative data compression method provided by an embodiment of the present invention;

[0040] Figure 7 is a flowchart showing a data decompression method provided by another embodiment of the present invention;

[0041] Figure 8 is a flowchart showing an alternative data decompression method provided by an embodiment of the present invention;

[0042] Figure 9 is a flowchart showing an alternative data decompression method provided by an embodiment of the present invention;

[0043] Figure 10 is a structural diagram showing a data compression device provided by an embodiment of the present invention;

[0044] Figure 11 is a schematic structural diagram of a data decompression device provided by an embodiment of the present invention;

[0045] Figure 12 is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. Detailed implementation manners

[0046] In the following description, for the purpose of illustration rather than limitation, specific details such as specific system architectures, technologies, etc. are presented to thoroughly understand the embodiments of the present invention. However, those skilled in the art should clearly understand that the present invention can also be implemented in other embodiments without these specific details. In other cases, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid unnecessary details from interfering with the description of the present invention.

[0047] It should be understood that when used in the specification and claims of the present invention, the term "comprising" indicates the presence of the described features, wholes, steps, operations, elements, and / or components, but does not exclude the presence or addition of one or more other features, wholes, steps, operations, elements, components, and / or their combinations.

[0048] It should also be understood that the term "and / or" as used in the specification and claims of the present invention refers to any combination and all possible combinations of one or more of the associated listed items, and includes these combinations.

[0049] As used in the specification and claims of the present invention, the term "if" can be interpreted as "when", "once", "in response to determining", or "in response to detecting" according to the context. Similarly, the phrase "if determined" or "if detecting [the described condition or event]" can be interpreted as meaning "once determined", "in response to determining", "once detecting [the described condition or event]", or "in response to detecting [the described condition or event]" according to the context.

[0050] In addition, in the description of the specification and claims of the present invention, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be understood as indicating or implying relative importance.

[0051] References to "one embodiment" or "some embodiments" etc. described in the specification of the present invention mean that specific features, structures, or characteristics described in connection with that embodiment are included in one or more embodiments of the present invention. Thus, statements such as "in one embodiment", "in some embodiments", "in other some embodiments", "in still other embodiments", etc. that appear in different places in this specification do not necessarily all refer to the same embodiment, but mean "one or more but not all embodiments", unless otherwise specifically emphasized. The terms "comprising", "including", "having" and their variants all mean "including but not limited to", unless otherwise specifically emphasized.

[0052] First, some terms in the embodiments of the present invention are explained to facilitate understanding by those skilled in the art.

[0053] The Double-Hashed Compression (DHC) algorithm is a data compression method designed to improve the deficiencies of traditional compression algorithms. It combines two hash calculations to enhance compression efficiency and effect.

[0054] The ROLZ algorithm, namely the Reduced Offset Lempel Ziv algorithm, is an improved variant of the LZ77 data compression algorithm.

[0055] The Core RZ algorithm, where RZ refers to the Reduced Offset Lempel Ziv (ROLZ) algorithm, is a further optimized or improved version based on the ROLZ algorithm. It makes core improvements in aspects such as the utilization of context information, table construction, or matching methods on the basis of the ROLZ algorithm to further improve the compression speed while maintaining a high compression ratio.

[0056] Huffman coding is an entropy coding algorithm widely used in the field of data compression, proposed by David A. Huffman in 1952. This algorithm assigns codes of unequal lengths to each character based on the probability of the character's occurrence to achieve the purpose of data compression.

[0057] The above is a simple introduction to the terms involved in the embodiments of the present invention, which will not be elaborated further below.

[0058] The example of the present invention provides an example of a data compression method. Please refer to Figure 1 as shown Figure 1 The following shows a schematic flowchart of a data compression method provided by the present invention. As an example and not a limitation, this method can be applied to or run in a data compression system or a data compression device. The method includes:

[0059] S101, in response to a data compression instruction, read the data to be compressed from the data input stream into the main buffer.

[0060] S102, when the amount of data to be compressed in the main buffer reaches a predetermined amount of data or a data end flag is detected, process the data to be compressed in the main buffer in chunks according to a predetermined size to obtain multiple chunks.

[0061] S103, perform compression processing on each chunk in sequence to obtain the compression result of each chunk, and write the compression result of each chunk into the temporary buffer respectively.

[0062] Optionally, in the example of the present invention, the data input stream can come from multiple channels, such as files, network transmissions, device data, etc. When the data compression system (i.e., the compression system) receives a data compression instruction, a connection channel with the data source is established according to the data source specified in the data compression instruction to obtain the data to be compressed.

[0063] In one example, the compression system reads the data to be compressed byte by byte or in chunks from the data input stream through a read operation, and stores the read data to be compressed in a pre-allocated main buffer. During the reading process, continuously track the amount of data to be compressed that has been read, and at the same time check whether a data end flag (such as the EOF flag, indicating the end of the file) is encountered.

[0064] After each time the data to be compressed is read from the data input stream, the compression system checks whether the amount of data to be compressed in the main buffer has reached a pre-set data volume threshold (e.g., 32MB), or whether a data end flag has been detected. If either of these two conditions is met, trigger the execution of the chunking operation: divide the data to be compressed in the main buffer according to a predetermined chunk size (e.g., 1MB). The division process is specifically based on the physical storage order of the data to be compressed, and continuous data segments are cut into fixed-size chunks.

[0065] It should be noted that for the last chunk, if the amount of data in this chunk is less than the predetermined chunk size, it is divided according to the actual amount of data.

[0066] After that, for each chunk, use compression algorithms such as ROLZ for processing, analyze the data characteristics of the data within the chunk, try to find duplicate patterns, rules, etc. in the data within the chunk, and use corresponding coding methods to compress the data to reduce the storage space of the data. Write the compression result obtained after each chunk is compressed into the temporary buffer in sequence. This temporary buffer is used to temporarily store the compressed chunk data, that is, the compression result of the chunk.

[0067] By using the data compression method provided in the examples of the present invention, through reasonable setting of the main buffer and the temporary buffer, and adopting a block processing method, the main buffer and the temporary buffer are used for data storage and processing, and the memory resources are reasonably allocated to avoid excessive occupation of memory. Moreover, by adopting the block processing method, the data processing task is decomposed into multiple independent subtasks, and parallel compression processing is carried out. The compression processing of each block can be completed independently, thereby significantly improving the data compression ratio and the data compression speed, and reducing the data storage space.

[0068] In a possible implementation manner, before reading the data to be compressed from the data input stream into the main buffer in response to a data compression instruction, the method further includes:

[0069] S001, obtaining the compression level specified by the user based on command line parameters or interface passed parameters;

[0070] S002, creating a corresponding compressor instance in the compression system according to the compression level, wherein the working mode of the compressor instance includes one of a fast compression mode, a balanced mode, and a high compression ratio mode;

[0071] S003, initializing the memory buffer in the compression system, wherein the memory buffer includes: a main buffer for storing the data to be compressed, a temporary buffer for block processing, a table storage space, and a status information storage space.

[0072] In one example, when the user calls the compression program through the command line, the compression program obtains the compression level (total_level) specified by the user by parsing the command line parameters. In different programming languages, there are corresponding libraries to handle command line parameters. For example, in the Python programming language, the argparse module in the Python standard library can be used to write a command line interface for parsing command line parameters and options. When the user inputs an instruction like python compressor.py --level 2 in the command line, the compression program can extract the compression level specified by the user as 2 from the command line parameters.

[0073] In another example, if the compression function is provided to other programs for calling in the form of an interface, when calling the interface, the calling party needs to pass the compression level as a parameter to the interface. For example, in a function call: compress data(data, level = 2), the level parameter is the compression level specified by the user.

[0074] In the example of the present invention, the compression system pre - defines compressor configurations corresponding to different compression levels, and a dictionary or configuration file can be used to store the compression levels and the corresponding configuration information. Each compression level corresponds to a working mode, such as a fast compression mode, a balanced mode, and a high compression ratio mode. It should be understood that different compression working modes have different configuration information in terms of the complexity of the compression algorithm, the matching depth, the delayed matching strategy, etc.

[0075] After obtaining the compression level specified by the user, according to the compression level specified by the user, find the configuration information corresponding to this compression level from the configuration mapping, and then use this configuration information to create a compressor instance. For example, in the Python programming language, a Compressor class can be defined and an instance of this class can be initialized according to different configurations.

[0076] Moreover, initialize the memory buffer in the compression system, in which:

[0077] The main buffer is used to store the data to be compressed read from the data input stream. According to the requirements and performance considerations of the compression system, a certain amount of memory space is allocated as the main buffer. For example, an array or a byte array can be used to implement the main buffer, and the size of the main buffer can be determined according to the data block size and the system memory limit.

[0078] The temporary buffer is used to temporarily store intermediate results during the chunk - processing. Therefore, the size of the temporary buffer is usually related to the size of the chunks. When chunking the data in the main buffer, the data of each chunk is first stored in the temporary buffer for subsequent compression processing.

[0079] The table storage space allocates a memory space of appropriate size to store data structures such as hash tables and coding tables according to the scale and complexity of the data. The hash table is used for double - hash matching search, and the coding table is used for multi - level coding processing. For example, specifically, an ASCII feature hash table can be stored: used to quickly locate possible matching positions; a cascaded feature hash table: used to improve the matching accuracy; store the encoding - decoding tables required for dynamic coding optimization; a ROID coding table: records the encoding mapping relationship of offsets; a ROID decoding table: records the decoding mapping relationship of offsets; store the coding tree required for Huffman coding; a symbol frequency statistics table; a Huffman coding bit - length table; a Huffman coding mapping table, etc.

[0080] The status information storage space is used to record various status information during the compression process, such as the current matching position, matching length, delayed matching flag, recorded context information, type of the most recently processed character, current matching status, delayed matching flag, storage of matching position information, historical matching position record, matching length statistics, offset data, saving temporary calculation results, hash value cache, intermediate matching results, encoded temporary data, and so on. This information is very important for the correct execution and optimization of the compression algorithm. Similarly, according to the quantity and type of the status information, the corresponding memory space is allocated. The purpose of pre-allocating this space is to avoid frequent memory allocation and release during the compression process, improve data access speed and cache hit rate, and ensure the continuous and efficient operation of the compression algorithm.

[0081] In the examples of the present invention, different users have different compression requirements in different scenarios. Some need fast compression, while some need a high compression ratio. By allowing users to specify the compression level, users can select different compression levels according to their compression requirements, so as to make a trade-off between compression speed and compression ratio. For example, the fast compression mode can complete the compression task in a shorter time and is suitable for scenarios with high time requirements; the high compression ratio mode can obtain a higher compression ratio and is suitable for scenarios with high requirements for storage space.

[0082] By initializing different memory buffers, different types of data and information are stored separately, avoiding the chaotic use of memory and improving the memory utilization efficiency. At the same time, a reasonable buffer size setting can reduce memory waste. By separating the creation of the compressor and the initialization of the memory buffer from the specific compression process, the compression system is made easier to expand and maintain.

[0083] In a possible implementation, after the compression results of each block are respectively written into the temporary buffer, the method further includes: performing a double-hash matching search on the compression results of multiple blocks to compress the duplicate data between multiple blocks to obtain compressed data.

[0084] Optionally, in the examples of the present invention, a double-hash matching search can be performed on the compression results of multiple blocks to find the duplicate data segments between multiple blocks. For these duplicate data, they only need to be stored once, and information such as the position and length of the duplicate data in a certain block is recorded, so as to achieve the compression of duplicate data.

[0085] In one example, for each data segment of the compression result in each block, the corresponding first-level hash value can be calculated, and the approximate matching range can be quickly located in the hash table. Based on the approximate matching range located by the first-level hash matching, using data context information such as the feature correlation of adjacent positions of each data segment, the second-level hash value of each data segment is further calculated, and the matching data of each data segment is screened out using the second-level hash value to improve the matching accuracy.

[0086] In large-scale data, there are often a large number of duplicate data. Traditional compression methods cannot effectively identify and process these duplicate data, resulting in a low compression ratio. In the example of the present invention, based on the block compression process, the double-hash matching search processing technology is used to identify the duplicate data between the compression results of multiple blocks and perform effective compression, avoiding the repeated storage of the same data, thereby significantly improving the data compression ratio.

[0087] In one possible implementation, as Figure 3 shown, perform double-hash matching search on the compression results of multiple blocks to compress the duplicate data between multiple blocks and obtain compressed data, including:

[0088] S201, when processing the compression results of multiple blocks, determine the data segment currently being processed, where the data segment includes the current character and the previous character of the current character.

[0089] S202, generate the first-level hash value and the second-level hash value of the current character according to the current character and the previous character.

[0090] S203, locate the best matching position of the current character according to the first-level hash value and the second-level hash value of the current character, and use the data at the best matching position as the hash matching result of the current character.

[0091] S204, determine the compressed data of the current character according to the hash matching result of the current character.

[0092] Optionally, in the example of the present invention, when processing the compression results of multiple blocks, the data in each block can be traversed in a certain order. During the traversal, clarify the position currently being processed, use the character at this position as the "current character", and at the same time select the previous character of this character, and combine these two characters into a data segment. For example, when processing text data, if the 10th character is currently being processed, the 10th character and the 9th character are used as the current data segment. The above processing method combines the current information and context information of the character at the current processing position, which helps to more accurately identify duplicate data.

[0093] When generating the first-level hash value and the second-level hash value of the current character based on the current character and the previous character, the first-level hash value is mainly generated based on the basic characteristics of the character. Since the lower 7 bits of the ASCII value can represent the basic encoding information of the character. For the current character, the lower 7 bits of the ASCII value of the current character are extracted, and at the same time, the type characteristics of the previous character are analyzed to determine whether the previous character is a letter or a number. Then, these two parts of information are combined through bit operations to generate a unique first-level hash value. For example, the lower 7 bits of the ASCII of the current character are shifted left by 1 bit, and then 0 or 1 is set in the lowest bit according to whether the previous character is a letter or a number, and finally the first-level hash value is obtained.

[0094] The first-level hash value can quickly locate the approximate matching range. The second-level hash value is based on the first-level hash value, and further utilizes the contextual relevance of the data to improve the matching accuracy. Specifically, it extracts a more comprehensive feature value of the current character and combines it with the first-level hash result of the previous position. The first-level hash result of the previous position is shifted left by several bits through a bit shift operation to make room for the feature value of the current character, and then the two are bitwise ored, and the feature information at different positions is recombined to generate a second-level hash value. Furthermore, each bit of the second-level hash value can be fully utilized to represent different features, so that the second-level hash value can more accurately reflect the characteristics of the data.

[0095] After that, based on the first-level hash value and the second-level hash value, the best matching position of the current character is located. First, the first-level hash value of the current character is used as an index to access the corresponding hash bucket. A hash bucket is a data structure that stores historical position information. Each hash bucket stores the locations where similar features appear in previously processed data. The first-level hash value can be used to quickly locate the hash bucket where a match may exist, thereby obtaining a list of candidate matching positions.

[0096] Then, based on the second-level hash value, the best matching position is further screened from the candidate matching position list. Since the second-level hash value contains richer context information, it can more accurately determine which candidate position is truly matched with the current data fragment. By comparing the current data and the data of the candidate position, calculating the maximum matching length and evaluating the matching quality, the best matching position (corresponding to the best matching data) can be finally selected. Furthermore, based on the hash matching result of the current character, the compressed data of the current character is determined.

[0097] Through the double-hash matching search technology, the first-level hash value is first used to quickly narrow down to the approximate matching range, and then the second-level hash value is used for fine screening, greatly improving the accuracy of finding duplicate data. Traditional compression methods sacrifice the matching accuracy when pursuing compression efficiency, or result in low processing efficiency when ensuring accuracy. The example of the present invention adopts the double-hash matching search technology, combining the advantages of fast positioning and precise screening, ensuring the matching speed and its accuracy while improving the compression efficiency, and achieving a balance between compression efficiency and matching speed.

[0098] In the example of the present invention, for duplicate data, only the relevant matching information needs to be stored once, instead of repeatedly storing the same data content, thereby improving the compression efficiency and reducing the space required for data storage. Moreover, in the process of generating hash values and matching, if only the characteristics of a single character are considered, it is often impossible to fully mine the duplicate patterns in the data. Therefore, in the example of the present invention, by considering the context information of the current character and the previous character, and using the hash result of the previous character, the compression algorithm can better understand the semantics and structure of the data to be compressed, further optimizing the compression effect.

[0099] In a possible implementation, as Figure 4 shown, S202, generating the first-level hash value and the second-level hash value of the current character according to the current character and the previous character, includes:

[0100] S301, generating the first-level hash value of the current character according to the information exchange standard code value of the current character and the character type of the previous character.

[0101] S302, generating the second-level hash value of the current character according to the feature correlation between the current character and adjacent characters, and the hash matching result of the previous character.

[0102] It should be understood that in the field of computer technology, characters are usually stored in a specific encoding form. For example, in forms such as ASCII code, UTF-8 encoding, etc. In the example of the present invention, the information exchange standard code value refers to the encoding value corresponding to the character. For example, in ASCII encoding, the code value of the letter 'A' is 65. In implementation, the encoding value of the currently processed character can be directly obtained.

[0103] Optionally, character types can be simply divided into several categories such as letters, numbers, punctuation marks, etc. In the example of the present invention, the type to which the previous character belongs can be judged by a preset rule. For example, in Python, the str.isalpha() statement can be used to judge whether the type of the character is a letter, and the str.isdigit() statement can be used to judge whether the type of the character is a number.

[0104] In the example of the present invention, the ASCII code value of the current character and the character type of the previous character can be combined through bitwise operations or other mathematical operations to generate a unique first-level hash value. For example, the lower 7 bits of the ASCII code value of the current character are shifted left by 1 bit, and then 0 or 1 is set at the lowest bit according to whether the previous character is a letter or a number (0 for letters and 1 for numbers).

[0105] After that, the feature correlation between the current character and adjacent characters and the hash matching result of the previous character are combined through a specific algorithm. For example, after quantifying the feature correlation, a numerical value is obtained. A bitwise exclusive OR operation is performed on a partial information of the hash matching result of the previous character (such as the lower bits of the matching position) and this numerical value, and then through a series of shift and modulo operations, a second-level hash value is finally generated.

[0106] It should be noted that there are generally semantic, syntactic, etc. correlations between adjacent characters. For example, in English texts, combinations such as "th", "er", etc. often appear. This kind of correlation can be quantified by analyzing features such as the combination pattern and occurrence frequency of characters. For example, the occurrence frequency of the current character and the triplet formed by the characters before and after this current character in the entire dataset is statistically calculated. It still needs to be noted that the previous character has completed hash matching in the previous processing, and the matching result of the previous character contains some useful information, such as the matching position, matching length, etc.

[0107] In the example of the present invention, the generation of the first-level hash value is relatively simple and the calculation speed is fast. By combining the encoding value of the current character and the type of the previous character, the approximate range where a possible match may exist can be quickly located in the hash table, reducing the time for global search and improving the preliminary screening efficiency of the match. The second-level hash value takes into account the feature correlation between the current character and adjacent characters and the hash matching result of the previous character, so that the hash value can more accurately reflect the context information of the data. By analyzing the correlation between adjacent characters, hidden patterns and rules in the data can be mined, improving the processing ability of the compression algorithm for complex data. Thus, on the basis of the first-level hash positioning, the truly matching data can be further screened out, reducing the situation of mis-matching.

[0108] In a possible implementation manner, according to the first-level hash value and the second-level hash value of the current character, the best matching position of the current character is located, including:

[0109] Using the first-level hash value of the current character as an index to access the corresponding hash bucket to obtain a list of candidate matching positions, where each candidate matching position in the list of candidate matching positions is used to represent the position where similar data features appear in the processed data;

[0110] Locate the best matching position from the candidate matching position list according to the second-level hash value.

[0111] In the example of the present invention, before performing data compression processing, a hash table is created in advance. The hash table consists of multiple hash buckets. Each hash bucket can be regarded as a storage unit for storing data position information with the same or similar first-level hash values.

[0112] After calculating the first-level hash value of the current character, use this hash value as an index to access the corresponding hash bucket in the hash table. Due to the characteristics of the hash function, data with similar data characteristics (i.e., the same or close to the first-level hash value) are mapped to the same hash bucket. That is, in the corresponding hash bucket, the position information of the data with similar data characteristics to the current character in the previously processed data is stored. By extracting this position information, a candidate matching position list is formed. This greatly reduces the search space and improves the preliminary screening efficiency of the matching.

[0113] The second-level hash value contains more detailed information about the current character and its context, and is more accurate than the first-level hash value. Using the second-level hash value, each candidate position in the candidate matching position list is further screened, so that the best matching position that matches the current character can be more accurately found among the candidate matching positions. By considering more context information and the evaluation of the matching quality, the possibility of mis-matching is reduced and the accuracy of the matching is improved.

[0114] Specifically, for each candidate matching position, recalculate the second-level hash value of the data near this position and compare it with the second-level hash value of the current character. It should be understood that during the screening process, not only compare whether the second-level hash values are equal, but also evaluate the quality of the matching. For example, calculate the matching length of the current data segment and the data segment at the candidate position. The longer the matching length, the higher the matching quality. At the same time, the distance factor of the position can also be considered, and the matching closer to the current position is preferred to reduce the additional overhead required to store the matching information.

[0115] In large-scale data, if a global search is performed for each character to find a match, the time complexity will be very high, resulting in low processing efficiency. Using the first-level hash value for fast indexing can narrow down the search scope to a specific hash bucket, solving the problem of low data search efficiency. However, although the first-level hash value can quickly locate the approximate matching range, due to its relatively simple calculation, there may be cases where multiple different data segments have the same or similar first-level hash values, resulting in false matches. By introducing the second-level hash value to introduce more context information and match quality evaluation, screening and evaluating all candidate positions in the candidate position list, and selecting the position with the highest match quality as the best match position, the accuracy of the match is improved, solving the problem of inaccurate matching.

[0116] In one possible implementation, as Figure 5 shown, S204, determine the compressed data of the current character according to the hash matching result of the current character, including:

[0117] S501, if the best match position is not located, directly encode the current character as compressed data and update the match context information, where the match context information includes at least one of: the type of the most recently processed character, the current match status, and the delayed match flag.

[0118] S502, if the best match position is located, evaluate the matching information between the data at the best match position and the current character, where the matching information includes: the matching length and position optimality; after executing S502, execute S503.

[0119] S503, perform delayed match optimization and encode the matching information as compressed data.

[0120] During the double-hash matching search process, if the best position that matches the current character is not found, it indicates that the current character does not appear repeatedly in the processed data, and the match context information is updated. At this time, the current character is directly encoded as compressed data according to the established encoding rule. For example, if ASCII encoding is used, the ASCII code value corresponding to the character is used as part of the compressed data.

[0121] Optionally, in the example of the present invention, updating the match context information mainly includes at least one of the type of the most recently processed character, the current match status, and the delayed match flag.

[0122] For example, recording whether the most recently processed character is a letter, digit, punctuation mark, or other type helps analyze the characteristics and patterns of the data. For example, if consecutive processed characters are all letters, it indicates that the currently processed content is a text segment. Another example is to clarify the current matching status so that different decisions can be made based on this matching status during subsequent processing. Another example is that the delayed matching flag can be used when no match is found currently, but it is possible that subsequent data forms a match with the current character. A delayed matching flag can be set to remind the subsequent processing process to pay attention to this possibility.

[0123] By updating the matching context information, the compression algorithm can make corresponding adjustments according to the dynamic changes of the data, enhancing the adaptability of the compression algorithm to different types of data and solving the problem that a single compression strategy cannot adapt to complex data.

[0124] In addition, during the double-hash matching search process, if the best matching position for the current character is successfully located, the matching information between the data at this best matching position and the current character is evaluated. This matching information mainly includes: matching length and position optimality.

[0125] Optionally, in the example of the present invention, the matching length refers to calculating the number of identical characters between the data segment where the current character is located and the data segment at the best matching position. The longer the matching length, the more duplicate data there is, and the better the compression effect.

[0126] Optionally, in the example of the present invention, the position optimality considers the rationality of the best matching position, such as whether the position is within a reasonable distance range. If the matching position is too far from the current position, it will increase the overhead of storing the matching information and affect the compression efficiency.

[0127] To obtain a better compression ratio, delayed matching optimization is performed. That is, instead of immediately determining the current matching result, subsequent data is continuously checked to determine whether there is a longer or better match. For example, based on the current match, look at a few more characters backward to determine whether there is a better match at other positions.

[0128] After evaluation and optimization, the matching information (such as the starting position of the match, matching length, etc.) is encoded into compressed data according to specific encoding rules. In this way, during decompression, the original data can be accurately restored based on this matching information.

[0129] Through the example of the present invention, when the best matching position is not located, the current character is directly encoded, ensuring that all data can be correctly compressed without losing any information and guaranteeing the integrity of the data. When the best matching position is located, by evaluating the matching information and performing delayed matching optimization, duplicate data can be more accurately identified and processed, avoiding unnecessary duplicate storage, thereby improving the compression ratio of the data.

[0130] In one possible implementation, as Figure 6 shown, after determining the compressed data of the current character according to the hash matching result of the current character in S204, it further includes:

[0131] S601, performing multi-level encoding processing on multiple compressed data to obtain pre-encoded data.

[0132] Among them, the multi-level encoding processing is used to indicate the sequential execution of symbol sorting transformation, dynamic bit length encoding, and Huffman entropy encoding.

[0133] S602, writing the block header information into the pre-encoded data to obtain the final encoded data.

[0134] Among them, the block header information includes: data size, check value.

[0135] S603, writing the final encoded data into the data output stream, and writing the end marker into the data output stream after the multi-level encoding processing ends.

[0136] Optionally, the symbol sorting transformation is an operation of rearranging data, aiming to make the data have better distribution characteristics for subsequent encoding. Specifically, the symbols can be sorted according to the order and frequency of symbol occurrences in the data, so that similar symbols are as adjacent as possible.

[0137] In one optional implementation, first count the number of occurrences of each symbol in the compressed data, and then sort the symbols according to the occurrence frequency. For example, for a piece of text data, count the number of occurrences of each character, and arrange the characters with high occurrence frequencies in the front. Finally, rearrange the compressed data according to the sorted symbol order.

[0138] Optionally, the dynamic bit length encoding assigns different lengths of encodings to different data according to the statistical characteristics of the data. For example, data with high occurrence frequency is assigned a shorter encoding, and data with low occurrence frequency is assigned a longer encoding, so as to reduce the overall encoding length.

[0139] For example, after performing symbol sorting transformation processing on multiple compressed data, the frequency of symbols can be counted again. Calculate the number of encoding bits required for each symbol according to the frequency. Symbols with high frequency have fewer encoding bits, and symbols with low frequency have more encoding bits. Then encode the compressed data according to the new number of encoding bits.

[0140] Optionally, the Huffman entropy encoding is a variable length encoding algorithm based on the occurrence probability of characters. By constructing a Huffman tree, characters with high occurrence probability are placed in the upper layer of the tree with short encoding lengths; characters with low occurrence probability are placed in the lower layer of the tree with long encoding lengths.

[0141] For example, by counting the occurrence probabilities of symbols, a Huffman tree is constructed based on the probabilities. Starting from the root node, moving left is encoded as 0 and moving right is encoded as 1 to obtain the Huffman encoding for each symbol. Finally, the compressed data is converted into Huffman encoding to obtain the pre-encoded data.

[0142] Optionally, in the example of the present invention, the block header information mainly includes: data size and check value. The data size refers to the number of bytes of the pre-encoded data, which is obtained by calculating the length of the pre-encoded data. The check value is used to verify the integrity of the data, and a check algorithm such as CRC (Cyclic Redundancy Check) can be used for calculation.

[0143] By writing the data size and the check value into the beginning part of the pre-encoded data in a certain format, the final encoded data is formed. For example, a fixed-length field can be written first to represent the data size, and then a fixed-length field can be written to represent the check value. After that, the final encoded data is written into the data output stream, and an end marker EOF is written.

[0144] In one example, the final encoded data containing the block header information is first written into a specified data output stream. Optionally, the data output stream can be a file, a network connection, etc. For example, in the Python programming language, the write() method of a file object can be used to write data to a file. After the multi-level encoding process is completed, an end marker is written into the data output stream to indicate the end of data transmission or storage. The end marker can be a specific byte sequence, and when the decompression program reads the data and encounters the end marker, it determines that all the data has been received.

[0145] Technical effects

[0146] Improve the compression ratio: The symbol sorting transformation, dynamic bit length encoding, and Huffman entropy encoding in the multi-level encoding process cooperate with each other, making full use of the statistical characteristics of the data, reducing data redundancy, and thus improving the overall compression ratio.

[0147] After double-hash matching search and preliminary compression, there may still be some redundancy in the compressed data. The multi-level encoding process further explores the compression potential of the data through a combination of multiple encoding methods, solving the problem of how to improve the compression ratio. Also, during data transmission and storage, the data may be damaged due to various interferences. Since the check value in the block header information can be used to verify the integrity of the data during decompression. If the check value does not match, it indicates that an error may have occurred during data transmission or storage, ensuring the reliability of the data.

[0148] The example of the present invention also provides an example of a data decompression method. Please refer to Figure 7 as shown Figure 7FIG. 0 shows a schematic flowchart of a data decompression method provided by the present invention. As an example but not a limitation, this method can be applied to or run in a data decompression system or a data decompression device. The method includes:

[0149] S701, in response to a data decompression instruction, read data to be decompressed from a data input stream into a decompression buffer.

[0150] S702, perform decompression on the data to be decompressed in the decompression buffer to restore the symbol stream in the data to be decompressed.

[0151] S703, perform an inverse symbol sorting transformation process on the symbol stream in the data to be decompressed, and parse to obtain matching information or original characters.

[0152] S704, if matching information is parsed, locate the source data position corresponding to the data to be decompressed according to the matching information, and write the source data at the source data position as decompressed data into an output data stream.

[0153] S705, if original characters are parsed, write the original characters as decompressed data into the output data stream.

[0154] Optionally, in an example of the present invention, the decompression system waits for and captures a data decompression instruction triggered by a user through command line input, graphical interface operation, program call, or the like. Then, according to the data source specified in the data decompression instruction, a connection with the data input stream is established. Optionally, the data input stream can be a file, network transmission, or the like.

[0155] Read the data to be decompressed from the data input stream, and store the data to be decompressed in a pre-allocated decompression buffer. During the process of reading the data to be decompressed, the integrity and legality of the read data to be decompressed can be checked. For example, verify whether the check value in the block header information of the data to be decompressed is correct, and so on.

[0156] Since multi-level encoding (such as symbol sorting transformation, dynamic bit length encoding, Huffman entropy encoding) is used during the compression process, when decompressing the data to be decompressed, decoding needs to be performed in the reverse order. And, decoding is performed according to the byte type. For example, for single-byte data to be decompressed, single-byte decoding is used, and for double-byte data to be decompressed, double-byte decoding is used.

[0157] For example, Huffman entropy decoding: First, restore the encoded binary data to a symbol sequence according to the Huffman tree. The Huffman tree has been constructed during compression and may be stored at a specific location in the compressed data or can be reconstructed through specific rules. Dynamic bit-length decoding: According to the rules of dynamic bit-length encoding, restore the encodings of different lengths to the original symbols. Symbol sorting inverse transformation: Restore the data that has undergone symbol sorting transformation to the original order to obtain a symbol stream.

[0158] In the example of the present invention, by performing symbol sorting inverse transformation processing on the symbol stream, matching information or original characters are parsed. Since symbol sorting transformation is performed during compression, reverse operations are required during decompression. According to the sorting rules recorded during the compression process, restore the symbol stream to the original arrangement order.

[0159] In the restored symbol stream, distinguish between matching information and original characters. It should be understood that the matching information usually includes information such as the position and length of data duplication, while the original characters are individual characters that have not undergone matching compression.

[0160] If matching information is parsed, according to the source data position recorded in the matching information, locate the corresponding source data in the already decompressed data. Then write this part of the source data as decompressed data to the output data stream. For example, if the matching information indicates that the data at the current position is the same as the data at a previous position with a length of 10, then copy these 10 data from the previous position to the current position. If the original character is parsed, directly write the original character as decompressed data to the output data stream.

[0161] By decompressing according to the reverse operation of the compression process, the problem of how to restore the data compressed through multi-level encoding and double-hash matching compression to the original data is solved, and the compressed data can be accurately restored to the original data, ensuring the integrity and accuracy of the data. Using multi-level decoding processing and fast positioning of matching information, through reasonable decoding steps and utilization of matching information, repeated processing of data is avoided, improving the decompression efficiency, and the decompression task can be completed in a relatively short time.

[0162] In a possible implementation, as Figure 8 shown, in S701, before reading the data to be decompressed from the data output stream into the decompression buffer in response to the data decompression instruction, it further includes:

[0163] S801, obtain the compression level specified by the user based on command-line parameters or interface parameter passing.

[0164] S802, create a corresponding decompressor instance in the decompression system according to the compression level, where the working mode of the decompressor instance includes one of the fast decompression mode, balanced mode, and high decompression rate mode.

[0165] S803 initializes the output buffer in the decompression system. The output buffer includes: a decompression buffer, a table storage space, and a status information storage space.

[0166] Optionally, in the example of the present invention, when the user calls the decompression program in the decompression system through the command line, the decompression program uses the corresponding command line parameter parsing tool to obtain the compression level specified by the user. For example, in the Python programming language, it can be implemented with the help of the argparse module. For example, when the user enters an instruction like python decompressor.py --level 1 in the command line, the decompression program extracts the corresponding compression level of 1 from the command line parameters. Another example is that if the decompression function of the decompression program is provided in the form of an interface for other programs to call, the caller can pass the compression level as a parameter to the interface when calling the interface. For example, in the function call decompress_data(data, level = 1), the level parameter is the compression level specified by the user.

[0167] Since the decompression system pre-defines the decompressor configurations corresponding to different compression levels. Each compression level corresponds to a working mode, such as a fast decompression mode, a balanced mode, and a high decompression rate mode. These working modes have differences in configuration information such as the complexity of the decompression algorithm, the data reading method, and the matching search strategy. The decompression system can use a dictionary or a configuration file to store the compression levels and the corresponding configuration information.

[0168] According to the obtained compression level, find the corresponding configuration information from the configuration mapping, and then create a decompressor instance with this configuration information. For example, in the Python programming language, define a Decompressor class and initialize an instance of this class according to different configurations.

[0169] Moreover, in the example of the present invention, the output buffer in the decompression system is initialized in the following manner. The output buffer includes:

[0170] Decompression buffer: Used to store the data to be decompressed read from the data input stream. For example, according to the compressed data block size and the system memory limit, allocate a certain amount of memory space as the decompression buffer. The decompression buffer can be implemented using an array or a byte array.

[0171] Table storage space: Used to store various tables required during the decompression process, such as Huffman tree tables, symbol sorting tables, etc. These tables are generated during compression and are used for data decoding and restoration during decompression. Allocate an appropriate amount of memory space to store these tables according to the scale and complexity of the data.

[0172] Status information storage space: used to record various status information during the decompression process, such as the current decompression position, matching status, error flags, etc. This information is very important for the correct execution and error handling of the decompression algorithm. Allocate corresponding memory space according to the quantity and type of status information.

[0173] Through the examples of the present invention, users can select different compression levels according to their own needs, and then make a trade-off between the decompression speed and the decompression effect. For example, the fast decompression mode can complete the decompression task in a short time and is suitable for scenarios with high time requirements; the high decompression rate mode can ensure more accurate data restoration and is suitable for scenarios with high requirements for data accuracy.

[0174] Moreover, separating the creation of the decompressor and the initialization of the output buffer from the specific decompression process makes the system easier to expand and maintain. During the decompression process, a large amount of data and intermediate results need to be processed. By initializing different output buffers, different types of data and information are stored separately, avoiding chaotic use of memory and improving the memory utilization efficiency.

[0175] In a possible implementation, as Figure 9 shown, S701, in response to a data decompression instruction, read the data to be decompressed from the data output stream into the decompression buffer, including:

[0176] S901, in response to a data decompression instruction, read the data to be decompressed from the data output stream.

[0177] S902, verify the block header information of the data to be decompressed and check the data integrity of the data to be decompressed.

[0178] S903, if the block header information and data integrity of the data to be decompressed do not meet the decompression requirements, output a decompression error message.

[0179] S904, if both the block header information and data integrity of the data to be decompressed meet the decompression requirements, read the data to be decompressed into the decompression buffer.

[0180] Optionally, in the examples of the present invention, after the decompression system receives a data decompression instruction issued by the user, it will establish a connection with the data output stream according to the data source specified in the instruction. It should be understood that this data output stream is actually the storage location of the compressed data. For example, a compressed file or a data channel in network transmission. Then, reading operations can be performed byte by byte, block by block, or in a specific format (specifically depending on the storage method of the data and the requirements of the compression algorithm) to read the data to be decompressed from the data output stream.

[0181] In an alternative example, after reading the data to be decompressed from the data output stream, the header information of the data to be decompressed is verified to check data integrity. Specifically, the header information of the data to be decompressed usually contains the size information of the compressed data (data to be decompressed). After reading the size information of the data to be decompressed, this size information is compared with the actual amount of data read. If the two do not match, it indicates that there may be missing or redundant data.

[0182] In another alternative example, since the checksum value of the data to be decompressed is calculated during the data compression process, it can be used to verify whether the data has changed during transmission or storage. For example, alternative checksum algorithms include: CRC (Cyclic Redundancy Check), etc. The decompression program recalculates the checksum value of the read data to be decompressed and compares this checksum value with the checksum value recorded in the header. If the two are not equal, it indicates that the data may be corrupted.

[0183] In addition, in another alternative example, in addition to the checksum value verification method, other integrity checks can also be performed. For example, check whether the format of the data meets the requirements of the compression algorithm, whether there are illegal characters or data truncation, etc.

[0184] If both the header information of the data to be decompressed and the data integrity meet the decompression requirements, it indicates that the data can be decompressed normally. At this time, all or a certain rule of the read data to be decompressed is read into the pre-allocated decompression buffer to prepare for subsequent decompression operations.

[0185] If the header information of the data to be decompressed or the data integrity does not meet the decompression requirements, the decompression program outputs a decompression error message. This decompression error message can include specific error types (such as data size mismatch, checksum value error, etc.) so that the user can understand the problem with this decompression error. At the same time, the decompression program can also directly terminate the decompression process to avoid generating incorrect decompression results.

[0186] Through the examples of the present invention, verifying the header information of the data to be decompressed and checking the data integrity of the data to be decompressed can effectively avoid decompressing damaged or incomplete data, thereby ensuring the accuracy and reliability of the decompression result. When the data does not meet the decompression requirements, detailed error information is output in a timely manner so that the user can quickly understand the problem and take corresponding measures.

[0187] In an alternative example, a 100MB file is used as the data to be compressed or decompressed for testing, and the following benchmark test results are obtained:

[0188]

[0189] Among them, xz is a high compression ratio compression algorithm, namely based on LZMA2 (Lempel-Ziv-Markov chain algorithm, generation 2). CoreRZ-l2, CoreRZ-l1, and CoreRZ-l0 are compression processing algorithms provided by the examples of the present invention. Zstd (Zstandard) is a standard compression algorithm, a real-time compression algorithm developed by Facebook. Bzip2 is the second generation of compression algorithms, a compression algorithm based on the Burrows-Wheeler transform. Brotli is the Google Brotli compression algorithm, a general compression algorithm developed by Google. Lzfse is the Apple LZFSE compression algorithm, a fast compression algorithm developed by Apple. Gzip (GNU zip) is a GNU compression tool, a compression tool based on the DEFLATE algorithm. Note: The numeric suffixes (-6, -9, etc.) represent the compression level. The larger the number, the higher the compression ratio, but the slower the compression speed.

[0190] The following is an analysis of the performance advantages of the compression processing algorithm CoreRZ-l2 provided by the examples of the present invention compared with some mainstream algorithms among the above algorithms:

[0191] 1. Compression ratio comparison

[0192] - Compared with xz-6:

[0193] - CoreRZ-l2: 26,893,684, xz: 26,665,156

[0194] - The compression ratio of CoreRZ-l2 and xz only differs by 0.86%

[0195] 2. Compression speed comparison

[0196] - Compared with xz-6:

[0197] - CoreRZ-l2: 8.245 s, xz: 69.815 s

[0198] - The compression speed of CoreRZ-l2 is 8.47 times faster than that of xz

[0199] - Compared with other compression tools:

[0200] - CoreRZ-l2 is faster than brotli-9 (36.147 s)

[0201] - CoreRZ-l2 is faster than zstd-19 (62.931 s)

[0202] 3. Decompression speed comparison

[0203] - CoreRZ-l2: 1.414 s

[0204] - Comparable to mainstream tools:

[0205] - xz: 1.309 s

[0206] - bzip2: 3.538 s

[0207] In summary, through innovative algorithm design and efficient engineering implementation, the example of the present invention significantly improves the compression speed while maintaining a compression ratio close to the optimal one.

[0208] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0209] Corresponding to the data compression method in the above embodiments, Figure 10 is a schematic structural diagram of a data compression device provided by an embodiment of the present invention. This device can be implemented by software, hardware, or a combination of both to become part or all of a computer device, and this computer device can be Figure 12 the electronic device shown.

[0210] Referring to Figure 10 , this data compression device includes:

[0211] A first reading unit 1001, configured to read data to be compressed from a data input stream into a main buffer in response to a data compression instruction.

[0212] A block processing unit 1002, configured to, when the amount of data to be compressed in the main buffer reaches a predetermined amount of data or a data end flag is detected, perform block processing on the data to be compressed in the main buffer according to a predetermined size to obtain a plurality of blocks.

[0213] A compression processing unit 1003, configured to sequentially perform compression processing on each block to obtain a compression result for each block, and write the compression result of each block into a temporary buffer respectively.

[0214] Corresponding to the data decompression method in the above embodiments, Figure 11 is a schematic structural diagram of a data decompression device provided by an embodiment of the present invention. This device can be implemented by software, hardware, or a combination of both to become part or all of a computer device, and this computer device can be Figure 12 the electronic device shown.

[0215] Referring to Figure 11 , this data decompression device includes:

[0216] A second reading unit 1101, configured to read data to be decompressed from a data input stream into a decompression buffer in response to a data decompression instruction, in response to a data decompression instruction.

[0217] The decompression processing unit 1102 is configured to perform decompression on the data to be decompressed in the decompression buffer to restore the symbol stream in the data to be decompressed.

[0218] The parsing processing unit 1103 is configured to perform an inverse symbol sorting transformation process on the symbol stream in the data to be decompressed, and parse to obtain matching information or original characters.

[0219] The data output unit 1104 is configured to, if the matching information is parsed, locate the source data position corresponding to the data to be decompressed according to the matching information, and write the source data at the source data position as decompressed data into the output data stream; and if the original characters are parsed, write the original characters as decompressed data into the output data stream.

[0220] It should be noted that: the data compression device and data decompression device provided in the above embodiments are only illustrated by dividing the above functional modules. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the device is divided into different functional modules to complete all or part of the functions described above.

[0221] Each functional unit and module in the above embodiments can be integrated in one processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit. In addition, the specific names of the functional units and modules are only for the convenience of distinguishing each other, and are not used to limit the protection scope of the embodiments of the present invention.

[0222] It should be noted that, for the content such as information interaction and execution process between the above devices / units, since it is based on the same concept as the method embodiments of the present invention, its specific functions and the technical effects brought, please refer to the method embodiments for details, and will not be elaborated here.

[0223] The embodiments of the present invention further provide an electronic device, which includes one or more processors and a memory;

[0224] The memory is coupled to one or more processors, and the memory is used to store computer program code. The computer program code includes computer instructions, and one or more processors call the computer instructions to enable the electronic device to execute the data compression method and data decompression method shown above.

[0225] Figure 12A schematic structural diagram of an electronic device provided by an embodiment of the present invention. The electronic device 1200 may be a mobile phone, a smart screen, a tablet computer, a wearable electronic device, a vehicle-mounted electronic device, an augmented reality (AR) device, a virtual reality (VR) device, a notebook computer, an ultra-mobile personal computer (UMPC), a netbook, a personal digital assistant (PDA), a projector, or a communication device such as a server, a memory, a base station, or an intelligent vehicle, etc. The embodiment of the present invention does not impose any restrictions on the specific type of the electronic device.

[0226] The memory 1201 can be used to store a computer software program 1202 and modules. The processor 1203 executes various functional applications and data processing of the electronic device by running the software program and modules stored in the memory 1201. The memory 1201 mainly includes a program storage area and a data storage area. Among them, the program storage area can store an operating system, application programs required for at least one function (such as a sound playback function, an image playback function, etc.), etc.; the data storage area can store data created according to the use of the electronic device (such as audio data, a phone book, etc.). In addition, the memory 1201 may include a high-speed random access memory, and may also include a non-volatile memory, such as at least one magnetic disk storage device, a flash memory device, or other volatile solid-state storage devices.

[0227] Among them, the processor 1203 may include one or more of a central processing unit, an application processor (AP), a baseband processor, etc. The processor can be the nerve center and command center of a wireless router. The processor 1203 can generate operation control signals according to the instruction operation code and timing signal to complete the control of fetching instructions and executing instructions. The memory 1201 can be used to store computer-executable program codes, and the executable program codes include instructions. The processor 1203 executes various functional applications and data processing of the network device by running the instructions stored in the memory. The memory 1201 may include a program storage area and a data storage area, such as data storing a sound signal to be played. For example, the memory may be a double data rate synchronous dynamic random access memory DDR or a flash memory Flash, etc.

[0228] The embodiment of the present invention also provides a computer-readable storage medium, in which computer instructions are stored; when the computer-readable storage medium runs on an electronic device, the electronic device is enabled to execute the data compression method and data decompression method shown above.

[0229] Computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired manner (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or a data center that includes one or more media integrated therein. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium, or a semiconductor medium (such as a solid state disk (SSD)).

[0230] An embodiment of the present invention also provides a computer program product containing computer instructions. When the computer program product runs on an electronic device, the electronic device can execute the data compression method and data decompression method shown above.

[0231] The computer storage medium and computer program product provided by the above embodiments of the present invention are both used to execute the methods provided above. Therefore, the beneficial effects that can be achieved can refer to the beneficial effects corresponding to the methods provided above, and will not be elaborated here.

[0232] In the above embodiments, it can also be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, the processes or functions according to the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another, for example, the computer instructions can be transmitted from a website, a computer, a server, or a data center to another website, a computer, a server, or a data center in a wired manner (such as coaxial cable, optical fiber, Digital Subscriber Line (DSL)) or a wireless manner (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or a data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, a hard disk, a magnetic tape), an optical medium (such as a Digital Versatile Disc (DVD)), or a semiconductor medium (such as a Solid State Disk (SSD)), etc.

[0233] In the above embodiments, the descriptions of the respective embodiments have their own emphases. For parts not detailed or recorded in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0234] Those of ordinary skill in the art can realize that the units and algorithm steps of the examples described in combination with the embodiments of the present invention can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. A professional technician can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0235] In the embodiments provided by the present invention, it should be understood that the disclosed device / network device and method can be implemented in other ways. For example, the device / network device embodiments described above are merely illustrative. For example, the division of the modules or units is only a logical function division. In actual implementation, there may be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of the device or unit can be in electrical, mechanical or other forms.

[0236] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0237] The above-described embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments or equivalently replace some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention and should all be included in the protection scope of the present invention.

Claims

1. A data compression method, characterized in that: include: In response to the data compression instruction, read the data to be compressed from the data input stream into the main buffer; When the amount of the data to be compressed in the primary buffer reaches a predetermined amount or a data end mark is detected, the data to be compressed in the primary buffer is divided into blocks according to a predetermined size to obtain a plurality of blocks; Compression processing is performed on each of the blocks in turn to obtain a compression result of each of the blocks, and the compression result of each of the blocks is written into a temporary buffer respectively.

2. The method according to claim 1, characterized in that After writing the compression results of each of the blocks into a temporary buffer respectively, the method further comprises: A double hash matching search is performed on the compression results of the multiple blocks to compress the duplicate data between the multiple blocks to obtain compressed data.

3. The method according to claim 1, characterized in that Performing double hash matching search on the compression results of the plurality of the blocks to compress duplicate data between the plurality of the blocks to obtain compressed data, including: When processing the compression results of the plurality of said blocks, determining a data segment currently being processed, wherein the data segment includes a current character and a character preceding the current character; Generate a first-level hash value and a second-level hash value of the current character according to the current character and the previous character; Locate and obtain the best matching position of the current character according to the first-level hash value and the second-level hash value of the current character, and use the data of the best matching position as the hash matching result of the current character; Determine the compressed data of the current character according to the hash matching result of the current character.

4. The method according to claim 3, characterized in that The step of generating a first-level hash value and a second-level hash value of the current character according to the current character and the previous character comprises: Generate a first level hash value of the current character according to the information exchange standard code value of the current character and the character type of the previous character; A second level hash value of the current character is generated according to the feature association between the current character and adjacent characters, and the hash matching result of the previous character.

5. The method according to claim 4, characterized in that The locating the best matching position of the current character according to the first level hash value and the second level hash value of the current character comprises: Using the first-level hash value of the current character as an index to access the corresponding hash bucket, obtaining a candidate matching position list, wherein each candidate matching position in the candidate matching position list is used to represent a position where a similar data feature appears in the processed data; The best matching position is located from the candidate matching position list according to the second-level Hash value.

6. The method according to claim 5, characterized in that The step of determining the compressed data of the current character according to the hash matching result of the current character comprises: If the best matching position is not located, the current character is directly encoded into compressed data, and matching context information is updated, wherein the matching context information includes: at least one of the most recently processed character type, the current matching state, and a delayed matching flag; If the best matching position is located, then the matching information between the data of the best matching position and the current character is evaluated, wherein the matching information includes: matching length and position optimality; Delay matching optimization is performed, and the matching information is encoded as compressed data.

7. The method according to claim 6, characterized in that After determining the compressed data of the current character according to the hash matching result of the current character, the method further includes: Performing multi-stage encoding processing on the plurality of compressed data to obtain pre-coded data, wherein the multi-stage encoding processing is used to indicate sequentially performing symbol sorting transformation, dynamic bit length encoding, and Huffman entropy encoding; Writing block header information into the pre-coded data to obtain final coded data, wherein the block header information includes: data size and check value; The final encoded data is written into the data output stream, and an end marker is written into the data output stream after the multi-stage encoding process is completed.

8. The method according to any one of claims 1 to 7, characterized in that Before reading the data to be compressed from the data input stream to the primary buffer in response to the data compression instruction, the method further includes: Get the compression level specified by the user based on command line parameters or interface parameters; Creating a corresponding compressor instance in the compression system according to the compression level, wherein the working mode of the compressor instance includes one of a fast compression mode, a balanced mode, and a high compression ratio mode; Initialize a memory buffer in the compression system, wherein the memory buffer includes: a main buffer for storing data to be compressed, a temporary buffer for block processing, a table storage space, and a status information storage space.

9. A data decoding method, characterized in that: include: In response to the data decompression instruction, read the to-be-decompressed data from the data input stream into the decompression buffer; Decompressing the data to be decompressed in the decompression buffer to restore the symbol stream in the data to be decompressed; Performing symbol sorting inverse transformation processing on the symbol stream in the to-be-decompressed data, and parsing to obtain matching information or original characters; If the matching information is obtained by parsing, the source data position corresponding to the data to be decompressed is located according to the matching information, so as to write the source data at the source data position as the decompressed data into the output data stream; If the original characters are obtained through parsing, the original characters are written into the output data stream as decompressed data.

10. The method according to claim 9, characterized in that Before reading the to-be-decompressed data from the data output stream into the decompression buffer in response to the data decompression instruction, the method further includes: Get the compression level specified by the user based on command line parameters or interface parameters; Creating a corresponding decompressor instance in the decompression system according to the compression level, wherein the working mode of the decompressor instance includes one of a fast decompression mode, a balanced mode, and a high decompression rate mode; Initialize an output buffer in the decompression system, wherein the output buffer includes: a decompression buffer, a table storage space, and a status information storage space.

11. The method according to claim 9, characterized in that The step of reading the to-be-decompressed data from the data output stream into the decompression buffer in response to the data decompression instruction comprises: In response to the data decompression instruction, reading the to-be-decompressed data from the data output stream; Verifying block header information of the data to be decompressed, and checking data integrity of the data to be decompressed; If the block header information and the data integrity of the data to be decompressed both meet the decompression requirements, reading the data to be decompressed into the decompression buffer; If the block header information and the data integrity of the data to be decompressed do not meet the decompression requirements, a decompression error message is output.

Citation Information

Cited By

  • Information scheduling method, data compression transmission method and data decompression method

    CN121367705A