A method for converting LZ4 format files to GZIP format files
By parsing and constructing syntax elements and directly coding Huffman, the problem of slow conversion of LZ4 format files to GZIP format files is solved, and efficient format conversion is achieved, greatly improving the conversion speed.
Patent Information
- Application Number
- CN202111166379.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2041-09-30
AI Technical Summary
In the prior art, the conversion of LZ4 format files to GZIP format files is slower, and it is necessary to store the fully decompressed source files.
By analyzing the frame header and end of the LZ4 format file, the target syntax element is obtained and the GZIP format file header and end are constructed. Then parse the data blocks of the LZ4 format file, directly obtain the original text, the matching length and the offset distance, and Huffman encode the data to generate the deflate format data blocks, and finally encapsulate and generate the GZIP format file.
This method greatly improves the format conversion speed, almost skips the encoding process of the traditional conversion scheme, and directly uses the matching pairing of LZ4 format files to encode, thus avoiding the process of complete decoding and re-encoding.
Smart Images

Figure CN113986820B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of format conversion, and in particular to a method for converting an LZ4 format file into a GZIP format file; and also to an apparatus, a device and a computer-readable storage medium for converting an LZ4 format file into a GZIP format file. Background Art
[0002] With the rapid development of cutting-edge technologies such as big data, AI, and blockchain, data has exploded in growth, and massive data will put tremendous pressure on existing storage devices. In addition, with the replacement of traditional computing architectures by cloud computing, the structure of data storage is also changing. Computing resources and storage resources will further converge to the top data centers, further putting pressure on server storage. In the face of these continuously increasing massive data, data compression has become one of the effective ways to reduce the storage burden on servers and reduce storage costs. Data compression technology is mainly reflected in the compression processing of repeated redundant data, which can be implemented in the following steps: finding duplicate data, determining whether there are paragraphs in the previous text that are the same as the current data, and obtaining the address of the previous text; representing duplicate data, representing duplicate data according to certain rules, usually using run-length encoding; generating compressed files according to a specific compression format.
[0003] The mainstream data compression standards in the industry are GZIP data compression standard and LZ4 data compression standard. GZIP data compression standard is usually used in PC and server, while LZ4 data compression standard is usually used in mobile and IoT terminals. Data from various terminals need to be uploaded and stored in the data center. LZ4 algorithm can quickly compress and encapsulate the data generated by the terminal. Data center pays more attention to compression rate and usually uses GZIP algorithm to compress and encapsulate data. In this way, there is a "format gap" between terminal data and server. Therefore, data format conversion is required before data is stored on disk, that is, converting LZ4 format files to GZIP format files. At present, most solutions for converting LZ4 format files to GZIP format files adopt the method of complete decompression and then compression. Such conversion scheme has a slow conversion speed and needs to store the source files obtained after complete decompression.
[0004] In view of this, how to improve the conversion speed has become a technical problem that needs to be solved urgently by those skilled in the art. Summary of the invention
[0005] The purpose of the present application is to provide a method for converting an LZ4 format file to a GZIP format file, which can greatly improve the format conversion speed. Another purpose of the present application is to provide an apparatus, device and computer-readable storage medium for converting an LZ4 format file to a GZIP format file, all of which have the above technical effects.
[0006] To solve the above technical problems, the present application provides a method for converting an LZ4 format file into a GZIP format file, comprising:
[0007] Parse the frame header and frame tail of the LZ4 format file to obtain the target syntax element;
[0008] Construct a GZIP format file header, and construct a GZIP format file trailer according to the target syntax element obtained by parsing;
[0009] Parse the data block of the LZ4 format file to obtain the original text, matching length and offset distance;
[0010] Performing Huffman coding on the original text, the matching length and the offset distance to obtain a data block in a deflate format;
[0011] The GZIP format file header, the GZIP format file trailer, and the deflate format data block are encapsulated to obtain the GZIP format file.
[0012] Optionally, the parsing of the frame header and frame tail of the LZ4 format file to obtain the target syntax element includes:
[0013] Parse the frame header of the LZ4 format file to obtain a frame descriptor;
[0014] Parse the frame tail of the LZ4 format file to obtain the source data check code.
[0015] Optionally, constructing a GZIP format file header includes:
[0016] The first byte of the GZIP format check code in the GZIP format file header is set to a first preset value, and the second byte of the GZIP format check code is set to a second preset value;
[0017] Setting the compression algorithm identifier in the GZIP format file header to a third preset value;
[0018] Setting each bit of the flag bit in the GZIP format file header to zero;
[0019] Setting the source file timestamp in the GZIP format file header to the current time;
[0020] The additional flag and the operating system flag in the GZIP format file header are both set to a fourth preset value.
[0021] Optionally, constructing a GZIP format file trailer according to the target syntax element includes:
[0022] The source data check code at the end of the GZIP format file is set to be consistent with the source data check code at the end of the frame of the LZ4 format file;
[0023] The number of source data characters at the end of the GZIP format file is set to be consistent with the source file data length in the frame descriptor of the LZ4 format file.
[0024] Optionally, also include:
[0025] The block header information of the corresponding data block in the deflate format is set according to the data block in the LZ4 format file.
[0026] In order to solve the above technical problems, the present application also provides a device for converting an LZ4 format file into a GZIP format file, comprising:
[0027] The first parsing module is used to parse the frame header and frame tail of the LZ4 format file to obtain the target syntax element;
[0028] A construction module is used to construct a GZIP format file header, and construct a GZIP format file tail according to the target syntax element obtained by parsing;
[0029] A second parsing module is used to parse the data block of the LZ4 format file to obtain the original text, the matching length and the offset distance;
[0030] An encoding module, used for performing Huffman encoding on the original text, the matching length and the offset distance to obtain a data block in a deflate format;
[0031] The encapsulation module is used to encapsulate the GZIP format file header, the GZIP format file footer and the deflate format data block to obtain the GZIP format file.
[0032] Optionally, the first parsing module is specifically used for:
[0033] Parse the frame header of the LZ4 format file to obtain a frame descriptor;
[0034] Parse the frame tail of the LZ4 format file to obtain the source data check code.
[0035] Optionally, the building block is specifically used for:
[0036] The first byte of the GZIP format check code in the GZIP format file header is set to a first preset value, and the second byte of the GZIP format check code is set to a second preset value;
[0037] Setting the compression algorithm identifier in the GZIP format file header to a third preset value;
[0038] Set each bit of the flag bits in the GZIP format file header to zero;
[0039] Set the source file timestamp in the GZIP format file header to the current time;
[0040] Set both the additional flag and the operating system flag in the GZIP format file header to a fourth preset value.
[0041] To solve the above technical problems, the present application also provides a device for converting an LZ4 format file into a GZIP format file, including:
[0042] A memory for storing a computer program;
[0043] A processor for implementing the steps of the method for converting an LZ4 format file into a GZIP format file as described in any one of the above when executing the computer program.
[0044] To solve the above technical problems, the present application also provides a computer-readable storage medium, on which a computer program is stored, and the computer program implements the steps of the method for converting an LZ4 format file into a GZIP format file as described in any one of the above when executed by a processor.
[0045] The method for converting an LZ4 format file into a GZIP format file provided by the present application includes: parsing the frame header and frame tail of the LZ4 format file to obtain target syntax elements; constructing a GZIP format file header and constructing a GZIP format file tail according to the parsed target syntax elements; parsing the data block of the LZ4 format file to obtain the original text, the match length, and the offset distance; performing Huffman coding on the original text, the match length, and the offset distance to obtain a deflate format data block; encapsulating the GZIP format file header, the GZIP format file tail, and the deflate format data block to obtain the GZIP format file.
[0046] It can be seen that compared with the traditional conversion scheme of completely decoding and then re-encoding, the method for converting an LZ4 format file into a GZIP format file provided by the present application parses the data block of the LZ4 format file, directly obtains the original text, the match length, and the offset distance, and performs Huffman coding on the original text, the match length, and the offset distance to obtain a deflate format data block. Thus, it directly utilizes the match pairs (match length and offset distance) of the LZ4 format file, rather than re-searching for match pairs after completely decoding to obtain the source file, almost skipping the encoding process of the traditional conversion scheme, thereby greatly improving the conversion speed.
[0047] The device, equipment and computer-readable storage medium for converting LZ4 format files into GZIP format files provided in this application all have the above-mentioned technical effects. BRIEF DESCRIPTION OF THE DRAWINGS
[0048] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the prior art and the drawings required for use in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.
[0049] Figure 1 A schematic diagram of a process for converting an LZ4 format file into a GZIP format file provided in an embodiment of the present application;
[0050] Figure 2 A schematic diagram of the structure of a GZIP format file header provided in an embodiment of the present application;
[0051] Figure 3 A schematic diagram of the structure of a GZIP format file tail provided in an embodiment of the present application;
[0052] Figure 4 A schematic diagram of the structure of a compressed data block provided in an embodiment of the present application;
[0053] Figure 5 A schematic diagram of a device for converting an LZ4 format file into a GZIP format file provided in an embodiment of the present application;
[0054] Figure 6 A schematic diagram of the process of converting an LZ4 format file into a GZIP format file provided in an embodiment of the present application. DETAILED DESCRIPTION
[0055] The core of this application is to provide a method for converting LZ4 format files to GZIP format files, which can greatly improve the format conversion speed. Another core of this application is to provide an apparatus, device and computer-readable storage medium for converting LZ4 format files to GZIP format files, all of which have the above technical effects.
[0056] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0057] Please refer to Figure 1 , Figure 1 A flowchart of a method for converting an LZ4 format file to a GZIP format file provided in an embodiment of the present application is provided. Figure 1 As shown, the method includes:
[0058] S101: Parse the frame header and frame tail of the LZ4 format file to obtain the target syntax element;
[0059] S102: construct a GZIP format file header, and construct a GZIP format file trailer according to the target syntax element obtained by parsing;
[0060] Specifically, the conversion of an LZ4 format file to a GZIP format file involves transcoding of syntax elements and re-encoding of compressed data. Steps S101 and S102 are intended to complete the transcoding of syntax elements. The target syntax elements refer to the syntax elements required for the GZIP format file. The number and types of syntax elements in the GZIP format file are not exactly the same as the number and types of syntax elements in the LZ4 format file. Some syntax elements in the LZ4 format file are not syntax elements in the GZIP format file. Therefore, when converting an LZ4 format file to a GZIP format file, the target syntax elements in the LZ4 format file are parsed to obtain the syntax elements required for the GZIP format file. Then, a GZIP format file header is constructed, and a GZIP format file tail is constructed based on the target syntax elements obtained by parsing.
[0061] In a specific implementation manner, the parsing of the frame header and frame tail of the LZ4 format file to obtain the target syntax element includes:
[0062] Parse the frame header of the LZ4 format file to obtain a frame descriptor;
[0063] Parse the frame tail of the LZ4 format file to obtain the source data check code.
[0064] Specifically, the syntax elements involved in the LZ4 format are mainly divided into three levels: frame level, block level, and sequence level. Usually, a compressed file in LZ4 format is a frame, and each frame is roughly divided into a frame header, a compressed data block, and a frame tail. Among them, the frame header consists of two syntax elements: Magic number and Frame Descriptor. Magic number is the identification code of the LZ4 format file, which is a 32-bit number and its value must be equal to 0x184D2204. Frame Descriptor is a frame descriptor, which is a collection of control parameters required to decode the LZ4 file. The data structure of the frame descriptor is shown in Table 1:
[0065] Table 1
[0066] FLG BD (Content Size) (Dictionary ID) HC 1byte 1byte 0-8 bytes 0-4 bytes 1byte
[0067] It can be broken down into the following grammatical elements:
[0068] BD: A single-byte number used to identify the maximum length of the block data within the frame.
[0069] Content Size: 8 bytes, optional, used to indicate the data length of the source file, that is, the decompressed length.
[0070] Dictionary ID: 4 bytes, optional. Used to indicate the ID of the dictionary that the decoded data depends on.
[0071] HC: 1 byte, used to indicate the maximum value in the compressed block.
[0072] FLG: 1 byte. The structure of FLG is shown in Table 2:
[0073] Table 2
[0074] BitNb 7-6 5 4 3 2 1 0 FieldName Version B.Indep B. Checksum C.Size C.Checksum Reserved DictID
[0075] The meaning of each bit of FLG is as follows (from high to low):
[0076] Version Number: It is composed of bits 6-7 of FLG. The current Version Number value must be 01.
[0077] Block-Independence-flag: indicates the link relationship between blocks. "1" indicates no associated independent decoding, and "0" indicates a non-independent relationship. The decoding of the subsequent block needs to rely on the data of the previous block.
[0078] Block-checksum-flag: indicates whether the compressed block data ends with a checksum of the compressed data.
[0079] Content-Size-flag: Indicates whether the frame descriptor contains the Content Size element.
[0080] Content-checksum-flag: Indicates whether the end of the frame data contains the checksum of the source data.
[0081] Dictionary-ID-flag: If the value of this item is 1, the dictionary ID required for decoding the data needs to be specified in the frame descriptor.
[0082] The frame tail contains two syntax elements: End mask and CRC32, both of which are 32-bit numbers. The value of End mask is always 0x00000000, and CRC32 is the frame data check code.
[0083] The Content Size in the frame descriptor in the frame header of the LZ4 format file and the source data check code in the frame footer of the LZ4 format file are required by the GZIP format file, so only the frame descriptor in the frame header of the LZ4 format file and the source data check code in the frame footer of the LZ4 format file can be parsed. Other syntax elements not required by GZIP can be omitted to reduce the number of parsing times and improve the conversion speed.
[0084] In addition, in a specific implementation, the constructing of the GZIP format file header includes:
[0085] The first byte of the GZIP format check code in the GZIP format file header is set to a first preset value, and the second byte of the GZIP format check code is set to a second preset value;
[0086] Setting the compression algorithm identifier in the GZIP format file header to a third preset value;
[0087] Setting each bit of the flag bit in the GZIP format file header to zero;
[0088] Setting the source file timestamp in the GZIP format file header to the current time;
[0089] The additional flag and the operating system flag in the GZIP format file header are both set to a fourth preset value.
[0090] The constructing of a GZIP format file tail according to the target syntax element comprises:
[0091] The source data check code at the end of the GZIP format file is set to be consistent with the source data check code at the end of the frame of the LZ4 format file;
[0092] The number of source data characters at the end of the GZIP format file is set to be consistent with the source file data length in the frame descriptor of the LZ4 format file.
[0093] Specifically, a GZIP format file can be roughly divided into three parts: a GZIP format file header, several compressed data blocks (the compressed blocks in the GZIP protocol are encapsulated using Deflate), and a GZIP format file trailer.
[0094] The structure of the GZIP format file header is as follows Figure 2 As shown, it contains the following syntax elements:
[0095] 1) GZIP format checksum, a total of 2 bytes (ID1 and ID2), both bytes must be fixed values, where ID1 = 31 (0x1F), ID2 = 139 (0x8B).
[0096] 2) Compression algorithm identifier CM, a total of 1 byte. Currently, the GZIP compression algorithm only supports Deflate, so CM can be regarded as a fixed value (0x08).
[0097] 3) Flag bit FLG, a total of 1 byte. Each bit of FLG represents the following information:
[0098] bit 0FTEXT-indicates whether the text data (source file) is a text file. The LZ4 protocol has no relevant flag bit, so set this item to "0"
[0099] bit 1FHCRC-indicates the presence of a CRC16 header check field. Since the latest GZIP protocol does not support this item, the FHCRC bit can be set to "0"
[0100] bit 2 FEXTRA - indicates the presence of an optional field. The LZ4 protocol has no related flag bit, so FEXTRA is set to "0".
[0101] Bit 3 FNAME-indicates the existence of the original file name field. The LZ4 protocol does not have a related flag bit, so FNAME is set to "0" and the file field does not exist in the transcoded file.
[0102] Bit 4 FCOMMENT - indicates the presence of a comment field. The LZ4 protocol does not have a related flag bit, so FCOMMENT is set to "0". FCOMMENT does not exist in the transcoded file.
[0103] Bits 5-7 reserved are all set to 0.
[0104] 4) Source file timestamp MTIME, 4 bytes. Since there is no relevant content in the LZ4 protocol format, our transcoder will set the value of MTIME to the current time.
[0105] 5) Additional flag XFL and operating system (file system) flag OS. XFL and OS are both represented by 1 byte. Since there is no relevant content in the LZ4 protocol format, the transcoder will set both of these items to "0x00".
[0106] The structure of the GZIP file tail is as follows Figure 3 As shown, it contains two syntax elements: the source data check code CRC32 and the number of source data contents (number of bytes).
[0107] When constructing the GZIP file trailer, the source data checksum CRC32 in the GZIP file trailer is directly assigned to be consistent with the Content checksum in the LZ4 format. The source data character number ISIZE in the GZIP file trailer is directly assigned to be consistent with the content size in the LZ4 format.
[0108] S103: parsing the data block of the LZ4 format file to obtain the original text, matching length and offset distance;
[0109] S104: performing Huffman coding on the original text, the matching length and the offset distance to obtain a data block in a deflate format;
[0110] Specifically, steps S103 and S104 are intended to complete the re-encoding of compressed data. The compressed data block in the LZ4 protocol consists of three parts: a block header, sequences, and a block tail. Among them, the block header has only one syntax element: BlockSize, which is used to express the number of bytes of the compressed block. The highest bit (31 bits) of Block Size indicates the form of the current block: "1" indicates uncompressed form, and "0" indicates compressed form. The Block Checksum at the end of the block is a 32-bit number that represents the checksum value of the block compressed data. Block Checksum is a conditional option that depends on the Block-checksum-flag in the FLG in the frame header of the LZ4 format file. Usually there are multiple sequences in each compressed data block, and its format is as follows: Figure 4 shown.
[0111] Sequence is the smallest data unit in the LZ4 protocol. Sequence can be divided into five parts: Token, literal-length bytes, literals (original text), offset (offset), and Match length bytes. Its format is as follows:
[0112] Token is the first byte of Sequence, which is equivalent to Sequence identifier. The upper 4 bits of Token are related to the length of the original text, and the lower 4 bits are related to the size of the length. Literal length bytes are optional. If the value of the upper 4 bits of Token is less than 15, there are no literal length bytes; if the value of the upper 4 bits of Token is 15, it means that there are literal length bytes. When parsing, parse byte by byte. If the current byte is not 255, stop parsing. Literals: Several original characters. The number of characters is indicated by the Literal length header. Offset: Two bytes are used to indicate the offset of repeated data. Match length bytes: If the value of the lower 4 bits of Token is less than 15, there are no Match length bytes. If the value of the lower 4 bits of Token is 15, it means that there are Match length bytes.
[0113] The present application aims to directly use the Match copy (matching pair) encoding in the LZ4 format file to generate a GZIP format file, so as to skip the "search Match copy" process in the encoding algorithm. In fact, most of the encoder's computing resources are used in the "search Match copy" process. Therefore, directly using the Match copy (matching pair) encoding in the LZ4 format file to generate a GZIP format file is almost equivalent to skipping the encoding process, which can greatly improve the format conversion speed. To this end, the present application parses the syntax elements of all Sequences in the compressed data block of the LZ4 format file, including Literal-length, Literals, offset, and match-length. Among them, Literals corresponds to the original text symbol in the Deflate protocol, match-length corresponds to the length symbol in the Deflate protocol, and offset corresponds to the Distance symbol in the Deflate protocol. There is no symbol corresponding to Literal-length in the Deflate protocol, so it is only parsed without conversion.
[0114] In the Deflate protocol, Huffman codes are used to encode the above symbols. First, these symbols are classified and counted, and Huffman codes are generated by constructing a Huffman tree. Then, the three types of symbols obtained by parsing, namely Literals, offset, and match-length, are encoded using the Huffman code. According to the Deflate protocol, Literals and length are encoded using the same Huffman tree, and offset is encoded using an independent Huffman tree.
[0115] When the decoder decodes Deflate data, a Huffman code table is required, which is consistent with the code table of the encoder. Therefore, the Huffman code table needs to be encapsulated and encoded in the Deflate protocol. The Deflate protocol uses run-length encoding to encode the code length of each leaf of the Huffman tree, and the codeword is generated according to the protocol. Using the Huffman code generated above, the symbols are encoded in natural order. After the block data is encoded, a terminator needs to be followed at the end position.
[0116] This application does not elaborate on Huffman coding technology, such as how to construct a Huffman tree and how to perform Huffman coding. You can refer to the existing Huffman coding technology.
[0117] Each compressed data block in LZ4 format will generate a data block in Deflate format after conversion. The Deflate header contains only 3 bits of data, as follows:
[0118] a) BFINAL 1 bit: 1 indicates that the current deflate is the last block.
[0119] b) BTYPE 2bit: data compression encoding method, BTYPE value and meaning: 0 means no compression; 1 means static Huffman encoding; 2 means dynamic Huffman encoding.
[0120] Furthermore, it also includes:
[0121] The block header information of the corresponding data block in the deflate format is set according to the data block in the LZ4 format file.
[0122] Specifically, if the LZ4-format compressed data block is the last block, then the BFINAL bit of the Deflate header is set to "1"; otherwise, the BFINAL bit is set to "0". If the highest bit of the Block-size in the header of the LZ4-format compressed data block is "1", then the BTYPE bits of the Deflate header are set to "00"; otherwise, the BTYPE bit is set to "10".
[0123] S105: Encapsulate the GZIP format file header, the GZIP format file trailer, and the deflate format data block to obtain the GZIP format file.
[0124] Specifically, after constructing a GZIP format file header, a GZIP format file tail, and directly using the Match copy (matching pair) encoding in the LZ4 format file to generate a deflate format data block, further, the GZIP format file header, the GZIP format file tail, and the deflate format data block are encapsulated to obtain the GZIP format file.
[0125] To sum up, compared with the traditional conversion scheme of complete decoding and then re-encoding, the method for converting LZ4 format files into GZIP format files provided by the present application parses the data blocks of LZ4 format files, directly obtains the original text, matching length and offset distance, and performs Huffman encoding on the original text, the matching length and the offset distance to obtain data blocks in deflate format. Therefore, the matching pairs (matching length and offset distance) of the LZ4 format file are directly used, instead of re-searching for matching pairs after completely decoding the source file, which is almost equivalent to skipping the encoding process of the traditional conversion scheme, thereby greatly improving the conversion speed.
[0126] The present application also provides a device for converting an LZ4 format file into a GZIP format file. The device described below can be referenced in correspondence with the method described above. Figure 5 , Figure 5 A schematic diagram of a device for converting an LZ4 format file into a GZIP format file provided in an embodiment of the present application, combined with Figure 5 As shown, the device comprises:
[0127] The first parsing module 10 is used to parse the frame header and frame tail of the LZ4 format file to obtain the target syntax element;
[0128] A construction module 20 is used to construct a GZIP format file header, and construct a GZIP format file tail according to the target syntax element obtained by parsing;
[0129] A second parsing module 30 is used to parse the data block of the LZ4 format file to obtain the original text, the matching length and the offset distance;
[0130] The encoding module 40 is used to perform Huffman encoding on the original text, the matching length and the offset distance to obtain a data block in a deflate format;
[0131] The encapsulation module 50 is used to encapsulate the GZIP format file header, the GZIP format file trailer and the deflate format data block to obtain the GZIP format file.
[0132] Based on the above embodiment, optionally, the first parsing module 10 is specifically used for:
[0133] Parse the frame header of the LZ4 format file to obtain a frame descriptor;
[0134] Parse the frame tail of the LZ4 format file to obtain the source data check code.
[0135] Based on the above embodiment, optionally, the building module 20 is specifically used for:
[0136] The first byte of the GZIP format check code in the GZIP format file header is set to a first preset value, and the second byte of the GZIP format check code is set to a second preset value;
[0137] Setting the compression algorithm identifier in the GZIP format file header to a third preset value;
[0138] Setting each bit of the flag bit in the GZIP format file header to zero;
[0139] Setting the source file timestamp in the GZIP format file header to the current time;
[0140] The additional flag and the operating system flag in the GZIP format file header are both set to a fourth preset value.
[0141] Based on the above embodiment, optionally, the building module 20 is specifically used for:
[0142] The source data check code at the end of the GZIP format file is set to be consistent with the source data check code at the end of the frame of the LZ4 format file;
[0143] The number of source data characters at the end of the GZIP format file is set to be consistent with the source file data length in the frame descriptor of the LZ4 format file.
[0144] Based on the above embodiment, optionally, the method further includes:
[0145] The setting module is used to set the header information of the corresponding data block in the deflate format according to the data block of the LZ4 format file.
[0146] The device for converting an LZ4 format file into a GZIP format file provided by the present application parses the data block of the LZ4 format file to directly obtain the original text, the matching length and the offset distance, and performs Huffman encoding on the original text, the matching length and the offset distance to obtain the data block in the deflate format. Therefore, the matching pairs (matching length and offset distance) of the LZ4 format file are directly used instead of re-searching for the matching pairs after completely decoding the source file, which is almost equivalent to skipping the encoding process of the traditional conversion scheme, thereby greatly improving the conversion speed.
[0147] This application also provides a device for converting LZ4 format files into GZIP format files. Figure 6 As shown, the device includes a memory 1 and a processor 2 .
[0148] Memory 1, used for storing computer programs;
[0149] Processor 2 is used to execute the computer program to implement the following steps:
[0150] Parse the frame header and frame trailer of the LZ4 format file to obtain the target syntax element; construct a GZIP format file header, and construct a GZIP format file trailer according to the target syntax element obtained by parsing; parse the data block of the LZ4 format file to obtain the original text, the matching length and the offset distance; perform Huffman encoding on the original text, the matching length and the offset distance to obtain a deflate format data block; encapsulate the GZIP format file header, the GZIP format file trailer and the deflate format data block to obtain the GZIP format file.
[0151] For an introduction to the equipment provided in this application, please refer to the above method embodiments, and this application will not go into details here.
[0152] The present application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps can be implemented:
[0153] Parse the frame header and frame trailer of the LZ4 format file to obtain the target syntax element; construct a GZIP format file header, and construct a GZIP format file trailer according to the target syntax element obtained by parsing; parse the data block of the LZ4 format file to obtain the original text, the matching length and the offset distance; perform Huffman encoding on the original text, the matching length and the offset distance to obtain a deflate format data block; encapsulate the GZIP format file header, the GZIP format file trailer and the deflate format data block to obtain the GZIP format file.
[0154] The computer-readable storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, and other media that can store program codes.
[0155] For an introduction to the computer-readable storage medium provided in this application, please refer to the above method embodiment, and this application will not go into details here.
[0156] The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same or similar parts between the various embodiments can be referred to each other. For the devices, equipment, and computer-readable storage media disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part description.
[0157] Professionals may further appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of the two. In order to clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described in the above description according to function. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professionals and technicians may use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of this application.
[0158] The steps of the method or algorithm described in conjunction with the embodiments disclosed herein may be implemented directly using hardware, a software module executed by a processor, or a combination of the two. The software module may be placed in a random access memory (RAM), a memory, a read-only memory (ROM), an electrically programmable ROM, an electrically erasable programmable ROM, a register, a hard disk, a removable disk, a CD-ROM, or any other form of storage medium known in the art.
[0159] The technical solution provided by the present application is described in detail above. Specific examples are used herein to illustrate the principle and implementation method of the present application, and the description of the above embodiments is only used to help understand the method and core idea of the present application. It should be pointed out that for ordinary technicians in this technical field, without departing from the principle of the present application, several improvements and modifications can be made to the present application, and these improvements and modifications also fall within the scope of protection of the claims of the present application.
Claims
1. A method for converting LZ4 format files to GZIP format files. It is characterized in that include: Parse the frame header and frame tail of the LZ4 format file to obtain the target syntax element; Construct a GZIP format file header, and construct a GZIP format file trailer according to the target syntax element obtained by parsing; Parse the data block of the LZ4 format file to obtain the original text, matching length and offset distance; Performing Huffman coding on the original text, the matching length and the offset distance to obtain a data block in a deflate format; The GZIP format file header, the GZIP format file trailer, and the deflate format data block are encapsulated to obtain the GZIP format file.
2. The method for converting an LZ4 format file into a GZIP format file according to claim 1, It is characterized in that The frame header and frame tail of the LZ4 format file are parsed to obtain the target syntax elements including: Parse the frame header of the LZ4 format file to obtain a frame descriptor; Parse the frame tail of the LZ4 format file to obtain the source data check code.
3. The method for converting an LZ4 format file into a GZIP format file according to claim 2, It is characterized in that The construction of the GZIP format file header includes: The first byte of the GZIP format check code in the GZIP format file header is set to a first preset value, and the second byte of the GZIP format check code is set to a second preset value; Setting the compression algorithm identifier in the GZIP format file header to a third preset value; Setting each bit of the flag bit in the GZIP format file header to zero; Setting the source file timestamp in the GZIP format file header to the current time; The additional flag and the operating system flag in the GZIP format file header are both set to a fourth preset value.
4. The method for converting an LZ4 format file into a GZIP format file according to claim 2, It is characterized in that The constructing of a GZIP format file tail according to the target syntax element comprises: The source data check code at the end of the GZIP format file is set to be consistent with the source data check code at the end of the frame of the LZ4 format file; The number of source data characters at the end of the GZIP format file is set to be consistent with the source file data length in the frame descriptor of the LZ4 format file.
5. The method for converting an LZ4 format file into a GZIP format file according to claim 4, It is characterized in that Also includes: The block header information of the corresponding data block in the deflate format is set according to the data block in the LZ4 format file.
6. A device for converting a LZ4 format file into a GZIP format file, It is characterized in that include: The first parsing module is used to parse the frame header and frame tail of the LZ4 format file to obtain the target syntax element; A construction module is used to construct a GZIP format file header, and construct a GZIP format file tail according to the target syntax element obtained by parsing; A second parsing module is used to parse the data block of the LZ4 format file to obtain the original text, the matching length and the offset distance; An encoding module, used for performing Huffman encoding on the original text, the matching length and the offset distance to obtain a data block in a deflate format; The encapsulation module is used to encapsulate the GZIP format file header, the GZIP format file footer and the deflate format data block to obtain the GZIP format file.
7. The device for converting an LZ4 format file into a GZIP format file according to claim 6, It is characterized in that The first parsing module is specifically used for: Parse the frame header of the LZ4 format file to obtain a frame descriptor; Parse the frame tail of the LZ4 format file to obtain the source data check code.
8. The device for converting a LZ4 format file into a GZIP format file according to claim 6, It is characterized in that The building blocks are specifically used for: The first byte of the GZIP format check code in the GZIP format file header is set to a first preset value, and the second byte of the GZIP format check code is set to a second preset value; Setting the compression algorithm identifier in the GZIP format file header to a third preset value; Setting each bit of the flag bit in the GZIP format file header to zero; Setting the source file timestamp in the GZIP format file header to the current time; The additional flag and the operating system flag in the GZIP format file header are both set to a fourth preset value.
9. A device for converting LZ4 format files to GZIP format files, It is characterized in that include: Memory for storing computer programs; A processor, configured to implement the steps of the method for converting an LZ4 format file into a GZIP format file as described in any one of claims 1 to 5 when executing the computer program.
10. A computer-readable storage medium, It is characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the steps of the method for converting an LZ4 format file into a GZIP format file according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Data compression method and apparatus
CN107565971A
Business file storage method and device based on block chain
CN110032581A