A method for converting a GZIP format file to an LZ4 format file
By parsing and encoding GZIP files to directly generate LZ4 format files using Huffman trees, the method addresses the slow conversion issue, achieving faster format conversion.
Patent Information
- Application Number
- CN202111166387.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-09-30
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-09-30
AI Technical Summary
In the prior art, the conversion of GZIP format files to LZ4 format files is slower and cannot meet the application needs of fast conversion.
By analyzing the file end of the GZIP format file, obtaining the value of the target syntax element, and assigning it to the frame header and frame end of the LZ4 format file, building a Huffman tree to generate a Huffman code table, directly parsing the encoded data, encoded into a sequence of the LZ4 format file, encapsulating the frame header, frame end and sequence, and realizing direct conversion.
This greatly improves the format conversion speed, almost skips the recoding process in traditional conversion solutions, and realizes fast format conversion.
Smart Images

Figure CN114006619B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of format conversion, and particularly to a method for converting a GZIP format file into an LZ4 format file; it also relates to an apparatus, a device, and a computer-readable storage medium for converting a GZIP format file into an LZ4 format file. Background Art
[0002] Facing the continuously increasing massive data, data compression has become one of the effective methods to reduce the storage burden of servers and lower the storage cost. Data compression refers to reducing the data volume without losing useful information to reduce the storage space and improve its transmission, storage, and processing efficiency; or reorganizing the data according to a certain algorithm to reduce data redundancy and storage space. Currently, the industry mainly adopts two data compression standards: GZIP and LZ4. The GZIP data compression standard is usually adopted in PCs and servers, while the LZ4 data compression standard is usually adopted in mobile and Internet of Things terminals. When there is data interaction between the terminal and the server, the compressed data between them cannot be directly connected, and usually, the compressed data needs to be format-converted.
[0003] Currently, most of the conversion methods between different format data adopt the method of decoding and then encoding. That is, the data in one compression format is completely decoded to obtain the source data, and then the source data is encoded to obtain the data in another compression format. The conversion speed of this decoding and then encoding method is slow and can no longer meet the application requirements of fast conversion. Therefore, how to improve the format conversion speed and quickly complete the format conversion has become a technical problem that needs to be solved urgently by those skilled in the art. Summary of the Invention
[0004] The purpose of the present application is to provide a method for converting a GZIP format file into an LZ4 format file, which can improve the format conversion speed and quickly complete the format conversion. Another purpose of the present application is to provide an apparatus, a device, and a computer-readable storage medium for converting a GZIP format file into an LZ4 format file, all of which have the above technical effects.
[0005] To solve the above technical problems, the present application provides a method for converting a GZIP format file into an LZ4 format file, including:
[0006] Parsing the file tail of the GZIP format file to obtain the value of the target syntax element;
[0007] Encoding the frame header and frame tail of the LZ4 format file, and assigning the value of the target syntax element to the corresponding syntax elements in the frame header and the frame tail;
[0008] Constructing a Huffman tree and generating a Huffman code table according to the Huffman tree;
[0009] Parse the GZIP format file according to the Huffman code table to obtain encoded data;
[0010] Encode the encoded data to obtain a sequence of LZ4 format files;
[0011] Package the frame header, the frame tail, and the sequence to obtain the LZ4 format file.
[0012] Optionally, the obtaining the values of the target syntax elements by parsing the file tail of the GZIP format file includes:
[0013] Parse the file tail of the GZIP format file to obtain the value of the source data check code and the value of the source data byte count.
[0014] Optionally, the assigning the values of the target syntax elements to the corresponding syntax elements in the frame header and the frame tail includes:
[0015] Assign the value of the source data check code to the frame data check code of the frame tail of the LZ4 format file;
[0016] Assign the value of the source data byte count to the decompression length of the frame header of the LZ4 format file.
[0017] Optionally, the constructing the Huffman tree and generating the Huffman code table according to the Huffman tree includes:
[0018] Construct a first Huffman tree and generate a first Huffman code table according to the first Huffman tree; the first Huffman code table is used to parse the original text and the length;
[0019] Construct a second Huffman tree and generate a second Huffman table according to the second Huffman tree; the second Huffman code table is used to parse the displacement.
[0020] Optionally, it further includes:
[0021] Parse the block header of the data block of the GZIP format file to identify the last data block of the GZIP format file;
[0022] Add flag information indicating that the sequence is the last sequence after the sequence corresponding to the last data block.
[0023] To solve the above technical problems, the present application also provides a device for converting a GZIP format file into an LZ4 format file, including:
[0024] A first parsing module, configured to parse the file tail of the GZIP format file to obtain the values of the target syntax elements;
[0025] A first encoding module, configured to encode the frame header and frame tail of an LZ4 format file, and assign the value of the target syntax element to the corresponding syntax element in the frame header and the frame tail;
[0026] A construction module, configured to construct a Huffman tree and generate a Huffman code table according to the Huffman tree;
[0027] A second parsing module, configured to parse the GZIP format file according to the Huffman code table to obtain encoded data;
[0028] A second encoding module, configured to encode the encoded data to obtain a sequence of an LZ4 format file;
[0029] An encapsulation module, configured to encapsulate the frame header, the frame tail, and the sequence to obtain the LZ4 format file.
[0030] Optionally, the first parsing module is specifically configured to:
[0031] Parse the file tail of the GZIP format file to obtain the value of the source data check code and the value of the source data byte count.
[0032] Optionally, the first encoding module is specifically configured to:
[0033] Assign the value of the source data check code to the frame data check code of the frame tail of the LZ4 format file;
[0034] Assign the value of the source data byte count to the decompression length of the frame header of the LZ4 format file.
[0035] To solve the above technical problems, the present application further provides a device for converting a GZIP format file into an LZ4 format file, including:
[0036] A memory, configured to store a computer program;
[0037] A processor, configured to implement the steps of the method for converting a GZIP format file into an LZ4 format file as described in any one of the above when executing the computer program.
[0038] To solve the above technical problems, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method for converting a GZIP format file into an LZ4 format file as described in any one of the above are implemented.
[0039] The method for converting a GZIP format file into an LZ4 format file provided by this application includes: parsing the file tail of the GZIP format file to obtain the value of the target syntax element; encoding the frame header and frame tail of the LZ4 format file, and assigning the value of the target syntax element to the corresponding syntax elements in the frame header and the frame tail; constructing a Huffman tree, and generating a Huffman code table according to the Huffman tree; parsing the GZIP format file according to the Huffman code table to obtain encoded data; encoding the encoded data to obtain a sequence of the LZ4 format file; encapsulating the frame header, the frame tail, and the sequence to obtain the LZ4 format file.
[0040] It can be seen that, for the method for converting a GZIP format file into an LZ4 format file provided by this application, by parsing the GZIP format file, the encoded data is directly obtained, and the encoded data is encoded to obtain a sequence of the LZ4 format file. Thus, the encoded data in the GZIP format file is directly utilized, rather than re-searching for matching pairs in the encoded data after completely decoding to obtain the source file, which almost skips the re-encoding process of the traditional conversion scheme, thereby greatly improving the conversion speed.
[0041] The device, equipment, and computer-readable storage medium for converting a GZIP format file into an LZ4 format file provided by this application all have the above technical effects. Description of the Drawings
[0042] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings required for use in the prior art and the embodiments will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some embodiments of this application. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained according to these drawings.
[0043] Figure 1 It is a schematic flowchart of a method for converting a GZIP format file into an LZ4 format file provided by an embodiment of this application;
[0044] Figure 2 It is a schematic diagram of a device for converting a GZIP format file into an LZ4 format file provided by an embodiment of this application;
[0045] Figure 3 It is a schematic diagram of a device for converting a GZIP format file into an LZ4 format file provided by an embodiment of this application. Detailed Embodiments
[0046] The core of this application is to provide a method for converting a GZIP - formatted file into an LZ4 - formatted file, which can improve the format conversion speed and quickly complete the format conversion. Another core of this application is to provide a device, equipment, and computer - readable storage medium for converting a GZIP - formatted file into an LZ4 - formatted file, all of which have the above - mentioned technical effects.
[0047] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are part of the embodiments of this application, rather than all of them. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.
[0048] Please refer to Figure 1 , Figure 1 which is a schematic flowchart of a method for converting a GZIP - formatted file into an LZ4 - formatted file provided by an embodiment of this application. Referring to Figure 1 as shown, the method includes:
[0049] S101: Parse the file tail of the GZIP - formatted file to obtain the value of the target syntax element;
[0050] S102: Encode the frame header and frame tail of the LZ4 - formatted file, and assign the value of the target syntax element to the corresponding syntax elements in the frame header and the frame tail;
[0051] Specifically, the target syntax element refers to the syntax element required for the LZ4 - formatted file. The number and type of syntax elements in the GZIP - formatted file are not exactly the same as those in the LZ4 - formatted file. Some syntax elements in the GZIP - formatted file are not syntax elements in the LZ4 - formatted file. Therefore, when converting a GZIP - formatted file to an LZ4 - formatted file, it is necessary to extract the value of the syntax element required for the LZ4 - formatted file from the GZIP - formatted file.
[0052] In a specific implementation manner, the parsing of the file tail of the GZIP - formatted file to obtain the value of the target syntax element includes:
[0053] Parse the file tail of the GZIP - formatted file to obtain the value of the source data checksum and the value of the source data byte count.
[0054] The assigning of the value of the target syntax element to the corresponding syntax elements in the frame header and the frame tail includes:
[0055] Assign the value of the source data checksum to the frame data checksum of the frame tail of the LZ4 - formatted file;
[0056] Assign the value of the number of source data bytes to the decompression length of the frame header of the LZ4 format file.
[0057] Specifically, the structure of the GZIP format file is as Figure 2 shown, including a GZIP file header, several compressed data blocks encapsulated by Deflate, and a GZIP file trailer. Among them, the GZIP file header contains the following syntax elements: GZIP format check code, compression algorithm identifier, flag bits, source file timestamp, additional identifier, and operating system identifier.
[0058] The GZIP format check code totals 2 bytes (ID1 and ID2), and both bytes are fixed values. Among them, ID1 = 31 (0x1F), ID2 = 139 (0x8B). The corresponding LZ4 file check code is 0x184D2204.
[0059] The compression algorithm identifier CM totals 1 byte. The current GZIP compression algorithm only supports the Deflate compression algorithm. Therefore, the compression algorithm identifier CM can be regarded as a fixed value of 8, totaling 1 byte.
[0060] The flag bit FLG totals 1 byte. Among them, the information represented by each bit of the flag bit FLG is as follows:
[0061] bit 0 FTEXT - Indicates text data;
[0062] bit 1 FHCRC - Indicates the existence of a CRC16 header check field;
[0063] bit 2 FEXTRA - Indicates the existence of an optional field;
[0064] bit 3 FNAME - Indicates the existence of an original file name field;
[0065] bit 4 FCOMMENT - Indicates the existence of a comment field;
[0066] bits 5 - 7 reserved are all set to 0.
[0067] Since the LZ4 format does not involve the above flag bits and the attached information, the data attached to the above flag bits is directly discarded after being parsed.
[0068] The source file timestamp MTIME totals 4 bytes. Since the LZ4 format does not involve the source file timestamp, the process of parsing the source file timestamp can be skipped.
[0069] The additional flag XFL and the operating system identifier OS are both represented by 1 byte. The LZ4 format does not involve these two syntax elements, so the parsing process for these two syntax elements can be skipped.
[0070] The GZIP file trailer contains two syntax elements: the source data checksum CRC32 and the quantity (number of bytes) of the source data content. The frame header of the LZ4 format file consists of two syntax elements, the Magic number and the Frame Descriptor. The Magic number is the identification code of the LZ4 format file, which is a 32-bit number and its value must be equal to 0x184D2204. The Frame Descriptor is the frame descriptor, which is a set of control parameters required to decode the LZ4 file. The frame descriptor can be decomposed into the following syntax elements:
[0071] BD: Identified by a single-byte number, used to identify the maximum length of the block data inside the frame.
[0072] Content Size: 8 bytes, optional, used to indicate the data length of the source file, that is, the decompression length.
[0073] Dictionary ID: 4 bytes, optional. Used to indicate the ID of the dictionary on which the decoded data depends.
[0074] HC: 1 byte, used to represent the maximum value in the compressed block.
[0075] FLG: 1 byte.
[0076] The frame trailer of the LZ4 format file contains two syntax elements: End mask and CRC32, both of which are 32-bit numbers. Among them, the value of the End mask is always 0x00000000, and the CRC32 is the frame data checksum.
[0077] The source data checksum CRC32 in the GZIP format file is the same as the frame data checksum of the frame trailer of the LZ4 format file; the quantity of the source data content in the GZIP format file is the same as the decompression length of the frame header of the LZ4 format file. Therefore, parse the file trailer of the GZIP format file to obtain the value of the source data checksum and the value of the source data byte count, and assign the value of the source data checksum to the frame data checksum of the frame trailer of the LZ4 format file; assign the value of the source data byte count to the decompression length of the frame header of the LZ4 format file. Other syntax elements in the frame header and frame trailer of the LZ4 format file are assigned accordingly according to their meanings.
[0078] S103: Construct a Huffman tree and generate a Huffman code table according to the Huffman tree;
[0079] S104: Parse the GZIP format file according to the Huffman code table to obtain the encoded data;
[0080] Specifically, the data blocks in the GZIP format file are encapsulated with deflate. Therefore, to parse and obtain the encoded data, first parse the Huffman code information of the CZIP format file, construct a Huffman tree based on the parsed Huffman code information, and then generate a Huffman code table according to the Huffman tree. On the basis of constructing the Huffman code table, use the constructed Huffman code table to parse the encoded data in the GZIP file to obtain the original text, length, and offset.
[0081] The constructing of the Huffman tree and generating the Huffman code table according to the Huffman tree includes:
[0082] Construct a first Huffman tree and generate a first Huffman code table according to the first Huffman tree; the first Huffman code table is used to parse the original text and length;
[0083] Construct a second Huffman tree and generate a second Huffman table according to the second Huffman tree; the second Huffman code table is used to parse the displacement.
[0084] Specifically, construct a Huffman tree for the original text and length according to the parsed Huffman code information, and generate a corresponding Huffman code table according to the Huffman tree for the original text and length. The length of the generated Huffman code table is 286. Among them, 0 to 255 represent the original text, 257 to 286 represent the length, and 256 is the block end symbol. Construct a Huffman tree for the displacement according to the parsed Huffman information, and generate a corresponding Huffman code table according to the Huffman tree for the displacement.
[0085] S105: Encode the encoded data to obtain a sequence of the LZ4 format file;
[0086] Specifically, obtain the encoded data, including the original text, length, and offset, and encode the original text and the matching pair (including length and offset) to obtain a Sequence that conforms to the LZ4 format specification, that is, a sequence. A sequence refers to the smallest data unit in the LZ4 format.
[0087] The sequence is divided into five parts, including: Token, literal length bytes, literals (original text), offset (offset), Match length bytes.
[0088] The Token is the first byte of the Sequence and is equivalent to the identifier of the Sequence. The high 4 bits of the Token are related to the length of the original text, and the low 4 bits are related to the size of the length. Literal length bytes (additional original text length bytes) are optional. If the value of the high 4 bits of the Token is less than 15, there are no literal length bytes; if the value of the high 4 bits of the Token is 15, it means there are literal length bytes. During parsing, parse byte by byte and stop parsing if the current byte is not 255. Literals (original text), several original text characters. Offset (offset), represented by two bytes for the offset of the repeated data. Match length bytes (additional match length bytes), if the value of the low 4 bits of the Token is less than 15, there are no Match length bytes; if the value of the low 4 bits of the Token is 15, it means there are Match length bytes.
[0089] S106: Encapsulate the frame header, the frame tail, and the sequence to obtain the LZ4 format file.
[0090] Specifically, after obtaining the frame header, the frame tail, and the sequence of the LZ4 format file, further encapsulate the frame header, the frame tail, and the sequence to obtain the LZ4 format file, completing the conversion of the GZIP format file to the LZ4 format file.
[0091] Furthermore, it also includes:
[0092] Parse the block header of the data block of the GZIP format file to identify the last data block of the GZIP format file;
[0093] Add flag information indicating it is the last data block to the data block in the LZ4 format file corresponding to the last data block of the GZIP format file.
[0094] Specifically, there are only 3 bits of data in the Deflate header, as follows:
[0095] 1) BFINAL, a total of 1 bit. When the value of this bit is 1, it indicates that the compressed data block currently encapsulated using deflate is the last data block. Correspondingly, add flag information indicating it is the last data block to the last data block of the LZ4 format file.
[0096] 2) BTYPE, a total of 2 bits, used to represent the data compression coding method. The values and meanings of BTYPE:
[0097] 0 means no compression; 1 means static Huffman coding; 2 means dynamic Huffman coding.
[0098] In summary, the method for converting a GZIP format file into an LZ4 format file provided by this application directly obtains encoded data by parsing the GZIP format file, and encodes the encoded data to obtain a sequence of the LZ4 format file. Thus, the encoded data in the GZIP format file is directly utilized, rather than re-searching for matching pairs in the encoded data after fully decoding to obtain the source file, which almost skips the re-encoding process of the traditional conversion scheme, thereby greatly improving the conversion speed.
[0099] This application also provides a device for converting a GZIP format file into an LZ4 format file. The device described below can be correspondingly referred to the method described above. Please refer to Figure 2 , Figure 2 is a schematic diagram of a device for converting a GZIP format file into an LZ4 format file provided by an embodiment of this application. In combination with Figure 2 as shown, the device includes:
[0100] A first parsing module 10, configured to parse the file tail of the GZIP format file to obtain the value of a target syntax element;
[0101] A first encoding module 20, configured to encode the frame header and frame tail of the LZ4 format file, and assign the value of the target syntax element to corresponding syntax elements in the frame header and the frame tail;
[0102] A construction module 30, configured to construct a Huffman tree and generate a Huffman code table according to the Huffman tree;
[0103] A second parsing module 40, configured to parse the GZIP format file according to the Huffman code table to obtain encoded data;
[0104] A second encoding module 50, configured to encode the encoded data to obtain a sequence of the LZ4 format file;
[0105] An encapsulation module 60, configured to encapsulate the frame header, the frame tail, and the sequence to obtain the LZ4 format file.
[0106] Based on the above embodiment, optionally, the first parsing module 10 is specifically configured to:
[0107] Parse the file tail of the GZIP format file to obtain the value of the source data check code and the value of the source data byte count.
[0108] Based on the above embodiment, optionally, the first encoding module 20 is specifically configured to:
[0109] Assign the value of the source data check code to the frame data check code at the end of the frame of the LZ4 format file;
[0110] Assign the value of the source data byte count to the decompression length of the frame header of the LZ4 format file.
[0111] Based on the above embodiments, optionally, the building module 30 includes:
[0112] A first building unit, configured to build a first Huffman tree and generate a first Huffman code table according to the first Huffman tree; the first Huffman code table is used to parse the original text and the length;
[0113] A second building unit, configured to build a second Huffman tree and generate a second Huffman table according to the second Huffman tree; the second Huffman code table is used to parse the displacement.
[0114] Based on the above embodiments, optionally, it further includes:
[0115] A third parsing module, configured to parse the block header of the data block of the GZIP format file and identify the last data block of the GZIP format file;
[0116] An adding module, configured to add flag information indicating that the sequence is the last sequence after the sequence corresponding to the last data block.
[0117] The device for converting a GZIP format file into an LZ4 format file provided by the present application directly obtains encoded data by parsing the GZIP format file, and encodes the encoded data to obtain a sequence of the LZ4 format file. Thus, the encoded data in the GZIP format file is directly utilized, rather than re-searching for matching pairs in the encoded data after completely decoding to obtain the source file, which almost skips the re-encoding process of the traditional conversion scheme, thereby greatly improving the conversion speed.
[0118] The present application also provides a device for converting a GZIP format file into an LZ4 format file. Referring to Figure 3 as shown, the device includes a memory 1 and a processor 2.
[0119] The memory 1 is used to store a computer program;
[0120] The processor 2 is configured to execute the computer program to implement the following steps:
[0121] Parse the end of a GZIP format file to obtain the value of a target syntax element; encode the header and tail of an LZ4 format file, and assign the value of the target syntax element to the corresponding syntax elements in the header and the tail; construct a Huffman tree and generate a Huffman code table according to the Huffman tree; parse the GZIP format file according to the Huffman code table to obtain encoded data; encode the encoded data to obtain a sequence of the LZ4 format file; encapsulate the header, the tail, and the sequence to obtain the LZ4 format file.
[0122] For the introduction of the device provided in this application, please refer to the above method embodiments, and this application will not elaborate here.
[0123] This application also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the following steps can be implemented:
[0124] Parse the end of a GZIP format file to obtain the value of a target syntax element; encode the header and tail of an LZ4 format file, and assign the value of the target syntax element to the corresponding syntax elements in the header and the tail; construct a Huffman tree and generate a Huffman code table according to the Huffman tree; parse the GZIP format file according to the Huffman code table to obtain encoded data; encode the encoded data to obtain a sequence of the LZ4 format file; encapsulate the header, the tail, and the sequence to obtain the LZ4 format file.
[0125] The computer-readable storage medium may include: various media such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disc that can store program codes.
[0126] For the introduction of the computer-readable storage medium provided in this application, please refer to the above method embodiments, and this application will not elaborate here.
[0127] The various embodiments in the specification are described in a progressive manner. Each embodiment focuses on the differences from other embodiments. The same or similar parts among the various embodiments can be referred to each other. For the devices, equipment, and computer-readable storage media disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the description of the method part.
[0128] Those skilled in the art may further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the components and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of this application.
[0129] The steps of the methods or algorithms described in combination with the embodiments disclosed herein can be directly implemented by hardware, software modules executed by a processor, or a combination of the two. The software modules can be placed in a random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium well known in the technical field.
[0130] The technical solutions provided in this application have been introduced in detail above. Specific examples have been used herein to elaborate on the principles and implementation manners of this application. The description of the above embodiments is only used to help understand the method and its core idea of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the protection scope of the claims of this application.
Claims
1. A method for converting a GZIP format file to an LZ4 format file, characterized in that, including: Parsing the end of the GZIP format file to obtain the value of the target syntax element; Encoding the frame header and frame tail of the LZ4 format file, and assigning the value of the target syntax element to the corresponding syntax element in the frame header and the frame tail; Constructing a Huffman tree and generating a Huffman code table according to the Huffman tree; Parsing the GZIP format file according to the Huffman code table to obtain encoded data; Encoding the encoded data to obtain a sequence of the LZ4 format file; Encapsulating the frame header, the frame tail, and the sequence to obtain the LZ4 format file.
2. The method for converting a GZIP format file into an LZ4 format file according to claim 1, characterized in that, The parsing the end of the GZIP format file to obtain the value of the target syntax element includes: Parsing the end of the GZIP format file to obtain the value of the source data checksum and the value of the source data byte count.
3. The method for converting a GZIP format file into an LZ4 format file according to claim 2, wherein The assigning the value of the target syntax element to the corresponding syntax element in the frame header and the frame tail includes: Assigning the value of the source data checksum to the frame data checksum of the frame tail of the LZ4 format file; Assigning the value of the source data byte count to the decompression length of the frame header of the LZ4 format file.
4. The method for converting a GZIP format file into an LZ4 format file according to claim 3, wherein The constructing a Huffman tree and generating a Huffman code table according to the Huffman tree includes: Constructing a first Huffman tree and generating a first Huffman code table according to the first Huffman tree; the first Huffman code table is used for parsing the original text and the length; Constructing a second Huffman tree and generating a second Huffman table according to the second Huffman tree; the second Huffman code table is used for parsing the displacement.
5. The method for converting a GZIP format file into an LZ4 format file according to claim 4, wherein It further includes: Parsing the block header of the data block of the GZIP format file to identify the last data block of the GZIP format file; Adding flag information indicating the last data block to the data block in the LZ4 format file corresponding to the last data block of the GZIP format file.
6. An apparatus for converting a GZIP format file into an LZ4 format file, characterized in that, including: A first parsing module, configured to parse the end of the GZIP format file to obtain the value of the target syntax element; A first encoding module, configured to encode the frame header and frame tail of the LZ4 format file, and assign the value of the target syntax element to the corresponding syntax element in the frame header and the frame tail; A constructing module, configured to construct a Huffman tree and generate a Huffman code table according to the Huffman tree; A second parsing module, configured to parse the GZIP format file according to the Huffman code table to obtain encoded data; A second encoding module, configured to encode the encoded data to obtain a sequence of the LZ4 format file; An encapsulating module, configured to encapsulate the frame header, the frame tail, and the sequence to obtain the LZ4 format file.
7. The apparatus for converting a GZIP format file into an LZ4 format file according to claim 6, wherein The first parsing module is specifically configured to: Parse the end of the GZIP format file to obtain the value of the source data checksum and the value of the source data byte count.
8. The apparatus for converting a GZIP format file into an LZ4 format file according to claim 7, characterized in that, The first encoding module is specifically configured to: Assign the value of the source data checksum to the frame data checksum of the frame tail of the LZ4 format file; Assign the value of the source data byte count to the decompression length of the frame header of the LZ4 format file.
9. A device for converting a GZIP format file into an LZ4 format file, characterized in that, including: A memory, configured to store a computer program; A processor for implementing the steps of the method for converting a GZIP format file into an LZ4 format file as described in any one of claims 1 to 5 when executing the computer program.
10. A computer-readable storage medium, characterized in that, A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the method for converting a GZIP format file into an LZ4 format file as described in any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
Three-dimensional image data compression method and system
CN102547315A
Business file storage method and device based on block chain
CN110032581A