Method, device, electronic device and readable storage medium for processing compressed format data
By extracting and reassembling the uncompressed original data and checksum information of the compressed format data, the processing performance problem when merging compressed data is solved, efficient data merging and transmission is achieved, and resource consumption is saved.
Patent Information
- Application Number
- CN202111563243.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-12-20
- Publication Date
- 2025-09-09
- Estimated Expiration
- 2041-12-20
AI Technical Summary
In the prior art, when merging compressed format data, the compression process consumes a large amount of processing resources, resulting in reduced processing performance. In particular, when dynamically appending gzip compressed format content, the efficiency is low and the performance consumption is excessive.
By obtaining the data verification information and compressed data blocks of the uncompressed original data of the compressed format data to be merged, the compressed data blocks and the data verification information are directly extracted and reassembled to construct new compressed format data, thus avoiding re-compression processing.
It saves processing resources, improves processing performance, reduces memory and CPU consumption, improves the efficiency of large-scale traffic transmission and distributed storage, and improves user experience.
Smart Images

Figure CN114282141B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of data processing technology, and specifically to the field of artificial intelligence technology such as big data and information flow. Background Art
[0002] In order to save data storage space in the terminal and speed up data transmission in the network, it is necessary to effectively compress the data into a compressed format, such as the GNUzip (gzip) compression format. In practical applications, it is often necessary to add data to already compressed data. For example, a web server needs to append content to the compressed HyperText Markup Language (HTML) data sent from upstream.
[0003] Typically, the appended data and the data of the content to be appended can be decompressed separately, and then the decompressed original data can be merged. Then, the merged and decompressed original data can be compressed to form new compressed data.
[0004] However, since the algorithm used in the compression process determines that data compression will consume a large amount of processing resources, it will greatly affect the processing performance. Summary of the Invention
[0005] The present disclosure provides a method, device, electronic device, and readable storage medium for processing compressed format data.
[0006] According to one aspect of the present disclosure, a method for processing compressed format data is provided, comprising:
[0007] Acquire first data and second data to be merged, where the first data and the second data are data in a specified compression format;
[0008] Obtaining data verification information of uncompressed original data of the first data and data verification information of uncompressed original data of the second data based on the first data and the second data respectively;
[0009] Based on the data verification information of the uncompressed original data of the first data, the data verification information of the uncompressed original data of the second data, the compressed data blocks in the first data, the compressed data blocks in the second data and the constructed data basic information of the specified compression format, third data with the specified compression format is obtained.
[0010] According to another aspect of the present disclosure, there is provided a device for processing compressed format data, comprising:
[0011] a data acquisition unit, configured to acquire first data and second data to be merged, wherein the first data and the second data are data in a specified compression format;
[0012] A data processing unit, configured to obtain data verification information of uncompressed original data of the first data and data verification information of uncompressed original data of the second data based on the first data and the second data, respectively;
[0013] a data construction unit, configured to obtain third data having the specified compression format based on data verification information of the uncompressed original data of the first data, data verification information of the uncompressed original data of the second data, compressed data blocks in the first data, compressed data blocks in the second data, and constructed basic information of data in the specified compression format.
[0014] According to another aspect of the present disclosure, there is provided an electronic device, comprising:
[0015] at least one processor; and
[0016] a memory communicatively connected to the at least one processor; wherein,
[0017] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method of any possible implementation manner and the aspects described above.
[0018] According to yet another aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause the computer to execute the method of the above-mentioned aspect and any possible implementation manner.
[0019] It can be seen from the above technical solution that the embodiment of the present disclosure obtains the first data and the second data to be merged, and the first data and the second data are data with a specified compression format. Then, based on the first data and the second data, data verification information of the uncompressed original data of the first data and data verification information of the uncompressed original data of the second data are obtained respectively, so that the third data with the specified compression format can be obtained based on the data verification information of the uncompressed original data of the first data, the data verification information of the uncompressed original data of the second data, the compressed data blocks in the first data, the compressed data blocks in the second data and the constructed basic information of the data in the specified compression format. Since the compressed data blocks in the first data and the compressed data blocks in the second data are directly extracted as the main body of the data in the merged third data during the merging process, and there is no need to compress any data, it can effectively avoid the technical problem of affecting the processing performance due to the large amount of processing resources consumed by data compression, thereby saving processing resources and improving processing performance.
[0020] In addition, compared with traditional content appending methods, the technical solution provided by the present disclosure can reduce the application of memory space required for caching data and the CPU performance consumed by compressing data, and has great performance improvements in application scenarios such as large-scale traffic transmission and distributed efficient storage.
[0021] In addition, the technical solution provided by the present disclosure can effectively improve the user experience.
[0022] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0023] To more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following briefly introduces the drawings required for use in the embodiments or prior art descriptions. Obviously, the drawings described below are some embodiments of the present disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without inventive effort. The drawings are used to better understand the present solution and do not constitute a limitation of the present disclosure. Among them:
[0024] Figure 1A is a schematic diagram according to a first embodiment of the present disclosure;
[0025] Figure 1B for Figure 1A Schematic diagram of the principle in the corresponding embodiment;
[0026] Figure 2 is a schematic diagram according to a second embodiment of the present disclosure;
[0027] Figure 3 is a schematic diagram according to a third embodiment of the present disclosure;
[0028] Figure 4 The block diagram is a block diagram of an electronic device for implementing the method for processing compressed format data according to an embodiment of the present disclosure. DETAILED DESCRIPTION
[0029] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding. These details should be considered as merely exemplary. Therefore, those skilled in the art will recognize that various changes and modifications may be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.
[0030] Obviously, the described embodiments are only part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present disclosure.
[0031] It should be noted that the terminal devices involved in the embodiments of the present disclosure may include but are not limited to mobile phones, personal digital assistants (PDAs), wireless handheld devices, tablet computers and other smart devices; display devices may include but are not limited to personal computers, televisions and other devices with display functions.
[0032] In this document, the term "and / or" simply describes a relationship between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the related objects are in an "or" relationship.
[0033] To save data storage space on terminals and speed up data transmission on networks, data must be effectively compressed into a compressed format. For example, the GNUzip (gzip) compression format is a lossless data compression format widely used for storage and on computers and the internet. It significantly reduces the amount of data required to store and transmit files, especially text files.
[0034] However, in actual applications, it is often necessary to add data to compressed data. For example, a Web server needs to append content to compressed HyperText Markup Language (HTML) data sent from an upstream server.
[0035] Typically, the appended data and the data of the content to be appended can be decompressed separately, and then the decompressed original data can be merged. Then, the merged and decompressed original data can be compressed to form new compressed data.
[0036] However, since the algorithm used in the compression process determines that data compression will consume a large amount of processing resources, it will greatly affect the processing performance.
[0037] Taking the gzip compression format as an example, the Deflate compression algorithm used in gzip is significantly more complex than the Inflate decompression algorithm. Consequently, the Deflate compression algorithm consumes significant CPU and memory resources, with 90% of CPU consumption occurring during the data compression process. Consequently, existing gzip data merging methods are significantly inefficient and suffer from excessive performance overhead when large amounts of gzip-compressed content need to be dynamically appended.
[0038] Therefore, there is an urgent need to provide a method for processing compressed format data that can effectively improve processing performance.
[0039] Figure 1A is a schematic diagram according to the first embodiment of the present disclosure, as shown in Figure 1A shown.
[0040] 101. Acquire first data and second data to be merged, where the first data and the second data are data in a specified compression format.
[0041] 102. Obtain data verification information of uncompressed original data of the first data and data verification information of uncompressed original data of the second data based on the first data and the second data, respectively.
[0042] 103. Obtain third data having the specified compression format based on the data verification information of the uncompressed original data of the first data, the data verification information of the uncompressed original data of the second data, the compressed data blocks in the first data, the compressed data blocks in the second data, and the constructed basic information of the data in the specified compression format.
[0043] At this point, compressed data blocks are extracted from the first and second data to be merged, and then reassembled together with the data verification information of the original data of these compressed data blocks and the constructed basic data information of the specified compression format to obtain third data with the specified compression format. During the entire reassembly process, the data compression process is not re-performed.
[0044] It should be noted that part or all of the execution entities of 101 to 103 may be applications located in the local terminal, or may also be functional units such as plug-ins or software development kits (SDKs) set in the applications located in the local terminal, or may also be processing engines located in the network-side server, or may also be distributed systems located on the network side, for example, a processing engine or distributed system in a compressed format data processing device on the network side, etc. This embodiment does not specifically limit this.
[0045] It is understandable that the application may be a native program (nativeApp) installed on the local terminal, or may be a webpage program (webApp) of a browser on the local terminal, which is not limited in this embodiment.
[0046] In this way, by obtaining the first data and the second data to be merged, the first data and the second data are data with a specified compression format, and then, based on the first data and the second data, the data verification information of the uncompressed original data of the first data and the data verification information of the uncompressed original data of the second data are obtained respectively, so that the third data with the specified compression format can be obtained based on the data verification information of the uncompressed original data of the first data, the data verification information of the uncompressed original data of the second data, the compressed data blocks in the first data, the compressed data blocks in the second data and the constructed data basic information of the specified compression format. Since the compressed data blocks in the first data and the compressed data blocks in the second data are directly extracted as the data body part of the third data after the merger during the merging process, and there is no need to compress any data, it can effectively avoid the technical problem of affecting the processing performance due to the large amount of processing resources consumed by data compression, thereby saving processing resources and improving processing performance.
[0047] Taking the gzip compression format as an example of the specified compression format described in this disclosure, the compression algorithm commonly used in the gzip compression format is implemented based on the deflate algorithm. The data format can be composed of a header part, a compressed body and a tail part. The specific format can be as follows.
[0048]
[0049] in,
[0050] The header portion may include, but is not limited to, basic data information, time information, compression algorithm, and operating system information. It may consist of a 10-byte fixed header and an extended header of variable length. The length of the extended header may be determined by the fixed header.
[0051] The compression body can be composed of one or more blocks. The length of each block is not fixed. It is compressed mainly by the Deflate algorithm to form a compressed data block (Deflate compressed block). The Deflate algorithm is a compression algorithm that combines the LZ77 (Lempel-Ziv-1977) code and the Huffman code.
[0052] The tail portion may include data verification information of the uncompressed original data, such as the cyclic redundancy check (CRC) of the uncompressed original data and the original size (LEN) of the uncompressed original data, and is usually 8 bytes.
[0053] If you need to merge files gzip-1 and gzip-2 into file gzip-new.
[0054] The traditional method is to first decompress the gzip-1 and gzip-2 files separately to obtain the uncompressed original data text1 of gzip-1 and the uncompressed original data text2 of gzip-2. Then, the uncompressed original data text1 and text2 are merged into the merged data text3, which is then further compressed to obtain the file gzip-new.
[0055] During this process, memory space needs to be allocated for the uncompressed original data text1 of file gzip-1, the uncompressed original data text2 of file gzip-2, and the merged data text3. At the same time, the compression processing performed on the merged data text3 will consume a lot of CPU performance.
[0056] The method provided by the present disclosure is as follows: first, compressed data blocks (Deflate compressed block) are extracted from the files gzip-1 and gzip-2 respectively, the header part (Header) and the trailer part (Trailer) are reconstructed, and the extracted compressed data blocks (Deflate compressed block) 1 and compressed data blocks (Deflate compressed block) 2 are reassembled together with the reconstructed header part (Header) and trailer part (Trailer) into a new file gzip-new. The principle is as follows: Figure 1B shown.
[0057] Compared with traditional methods, during the execution of the method provided by the present invention, there is no need to read the entire file into the memory space. Only the compressed data blocks (Deflate compressed blocks) need to be read into the memory space. A streaming cache method is adopted. When a compressed data block (Deflate compressed block) is read, the next compressed data block (Deflate compressed block) is overwritten and read until the trailing part (Trailer) is read. Therefore, only one piece of memory space needs to be applied for caching the compressed data blocks (Deflate compressed blocks). Therefore, very little memory space is occupied, which greatly saves memory consumption.
[0058] Furthermore, no data is compressed, significantly reducing CPU consumption. Compared to traditional methods, this approach has been shown to reduce CPU consumption by over 90%. Since data is not recompressed, data integrity is effectively guaranteed.
[0059] In addition, with regard to disk reading and writing, compared with the traditional method that requires reading at least three files, namely file gzip-1, file gzip-2 and the merged data text3 after decompression, and writing three files, namely the uncompressed original data text1 after decompression of file gzip-1, the uncompressed original data text2 after decompression of file gzip-2 and file gzip-new, the technical solution provided by the present invention only needs to read two files, namely file gzip-1 and file gzip-2, and write one file, namely file gzip-new, which can effectively save more than 50% of disk input and output (IO) resources.
[0060] Optionally, in a possible implementation of this embodiment, in 102, the first data and the second data can be decompressed respectively to obtain the uncompressed original data of the first data and the uncompressed original data of the second data, and then, data verification information of the uncompressed original data of the first data and data verification information of the uncompressed original data of the second data can be obtained respectively based on the uncompressed original data of the first data and the uncompressed original data of the second data.
[0061] Specifically, the inflate algorithm can be used to decompress the first data and the second data, respectively. After obtaining the uncompressed original data after the first data is decompressed and the uncompressed original data after the second data is decompressed, the data verification information of the uncompressed original data of the first data can be obtained based on the uncompressed original data after the first data is decompressed, using a pre-configured verification strategy, and the data verification information of the uncompressed original data of the second data can be obtained based on the uncompressed original data after the second data is decompressed, using a pre-configured verification strategy.
[0062] In this implementation, the uncompressed original data of the first data and the second data after decompression are read, and the data verification information of each uncompressed original data is recalculated in turn, thereby realizing the construction of data verification information.
[0063] Optionally, in a possible implementation of this embodiment, in 102, the first data and the second data can be decompressed respectively to obtain data verification information of the uncompressed original data of the first data and data verification information of the uncompressed original data of the second data.
[0064] Specifically, the inflate algorithm can be used to decompress the first data and the second data respectively. While obtaining the data verification information, the data verification information can be further used for verification processing. Compared with other methods, the security and reliability of the data will be higher.
[0065] In this implementation, by reading the data verification information after the first data and the second data are decompressed, the construction of data verification information with high security and reliability is achieved.
[0066] Optionally, in a possible implementation of this embodiment, in 102, data reading processing can be performed on the data portion at a specified position in the first data and the data portion at a specified position in the second data, respectively, to obtain data verification information of the uncompressed original data of the first data and data verification information of the uncompressed original data of the second data.
[0067] Specifically, decompression processing is no longer performed on the first data and the second data. Instead, according to the data format characteristics of the specified compression format, the specified positions in the first data and the second data are directly read to obtain data verification information of the uncompressed original data of the first data and data verification information of the uncompressed original data of the second data.
[0068] Still using the gzip compression format as an example of the specified compression format described in this disclosure, generally speaking, the last 8 bytes of gzip compressed data are data verification information. Due to the data format characteristics of the gzip compression format, the last 8KB of the gzip compressed data can be directly read to obtain the data verification information of the uncompressed original data.
[0069] In this implementation, data verification information is quickly constructed by directly reading data at a specified position in the first data and the second data.
[0070] Optionally, in a possible implementation of this embodiment, in 103, the constructed basic information of the data in the specified compression format can be used as the data header part of the third data, and the compressed data blocks in the first data and the compressed data blocks in the second data can be used as the data body part of the third data, and the data verification information of the uncompressed original data of the first data and the data verification information of the uncompressed original data of the second data after merging can be used as the data tail part of the third data.
[0071] In a specific implementation process, in 103, basic information of data in a specified compression format for this operation may be constructed based on basic information of the first data and the second data.
[0072] In another specific implementation process, in 103, the data verification information of the uncompressed original data of the first data and the data verification information of the uncompressed original data of the second data can be merged to obtain a merged processing result, and the result is used as the data tail part of the third data.
[0073] The gzip compression format is still taken as an example of the specified compression format described in this disclosure.
[0074] First, an empty file named gzip-new may be created, and a header part (Header) new may be constructed, and the constructed header part (Header) new may be written into the created empty file gzip-new.
[0075] Next, the gzip-1 file is read, skipping its header 1. The gzip-1 file is decompressed using the inflate algorithm. The decompressed data of the gzip-1 file is read in a loop. If the decompressed data is a compressed data block 1, the compressed data block 1 is sequentially written to the gzip-new file, which already has the header 1 new written to it. If the decompressed data is a trailer 1, the data checksum information (LEN1 and CRC1) of the gzip-1 file's uncompressed raw data 1 is recorded and the loop is exited.
[0076] Similarly, file gzip-2 is read, and header 2 of file gzip-2 is skipped. File gzip-2 is decompressed using the inflate algorithm. The decompressed data of file gzip-2 is read in a loop. If the decompressed data is compressed data block 2, compressed data block 2 is written sequentially to file gzip-new, which already has header new and compressed data block 1 written to it. If the data is trailer 2, the data check information (LEN2 and CRC2) of the uncompressed original data 1 of file gzip-2 is recorded and the loop is exited.
[0077] Next, after obtaining the data verification information of the uncompressed original data 1 of the file gzip-1, namely LEN1 and CRC1, and the data verification information of the uncompressed original data 1 of the file gzip-2, namely LEN2 and CRC2, the recorded LEN1 and LEN2 can be merged to obtain LENnew, and the recorded CRC1 and CRC2 can be merged to obtain CRCnew.
[0078] Specifically, LEN1 and LEN2 can be added together as the result of merging LEN1 and LEN2 to obtain LENnew of the file gzip-new. At the same time, CRC1 and CRC2 can be shifted based on LEN1 and LEN2. The shifted CRC1 can be expressed as CRC1+0...0 (the number of 0s is LEN2), and the shifted CRC2 can be expressed as 0...0+CRC2 (the number of 0s is LEN1). Then, the shifted CRC1 and CRC2 can be XORed as the result of merging CRC1 and CRC2 to obtain CRCnew of the file gzip-new.
[0079] Finally, after obtaining LENnew and CRCnew, LENnew and CRCnew are used to construct the trailer part (Trailer) new, and the constructed trailer part (Trailer) new is written into the file gzip-new in which the header part (Header) new, the compressed data block (Deflate compressed block) 1 and the compressed data block (Deflate compressed block) 2 have been written.
[0080] The technical solution provided by the present disclosure can be applied to distributed efficient storage and network transmission of web applications. In practical applications, the technical solution provided by the present disclosure can be used to dynamically append content to a data response stream with a specified compression format in a web server or a web server load without recompression processing, such as Figure 2 shown.
[0081] 201. The web server or the web server's payload obtains an HTML data response stream.
[0082] 202. The web server or the web server load determines whether the HTML data response stream needs to be appended. If so, execute 203; otherwise, execute 207.
[0083] Specifically, the Web server or the load of the Web server can determine whether the HTML data response stream needs to be supplemented with content based on the header data and / or body data of the HTML data response stream.
[0084] 203. The web server or the web server's load determines whether the HTML data response stream is a gzip data stream in gzip compression format. If so, execute 204; otherwise, execute 208.
[0085] Specifically, the Web server or the load of the Web server may judge the HTML data response stream according to the header data of the HTML data response stream to determine whether the HTML data response stream is a gzip data stream in a gzip compression format.
[0086] 204. The Web server or the load of the Web server obtains data verification information of uncompressed original data of the HTML data response stream and data verification information of uncompressed original data of the additional data according to the HTML data response stream and the additional data in gzip compression format, respectively.
[0087] 205. The web server or the load of the web server obtains merged data in gzip compression format based on the data verification information of the uncompressed original data of the HTML data response stream, the data verification information of the uncompressed original data of the appended data, the compressed data blocks in the HTML data response stream, the compressed data blocks in the appended data, and the constructed basic information of the data in gzip compression format.
[0088] 206. The web server or the payload of the web server outputs the merged data in a gzip compressed format.
[0089] 207. The web server or the web server's load outputs the obtained HTML data response stream.
[0090] 208. The web server or the load of the web server merges the uncompressed additional data and the obtained HTML data response stream and then outputs the merged data.
[0091] Compared with traditional methods, the performance and efficiency can be improved fourfold. The technical solution provided in this embodiment can dynamically append content to a data response stream in gzip compression format at the lowest performance cost. It is an efficient merging method that can merge multiple data response streams into one data response stream while ensuring data integrity.
[0092] In this embodiment, by obtaining the first data and the second data to be merged, the first data and the second data are data with a specified compression format, and then, based on the first data and the second data, the data verification information of the uncompressed original data of the first data and the data verification information of the uncompressed original data of the second data are obtained respectively, so that the third data with the specified compression format can be obtained based on the data verification information of the uncompressed original data of the first data, the data verification information of the uncompressed original data of the second data, the compressed data blocks in the first data, the compressed data blocks in the second data and the constructed basic information of the data in the specified compression format. Since the compressed data blocks in the first data and the compressed data blocks in the second data are directly extracted as the main body of the data in the merged third data during the merging process, and there is no need to compress any data, it can effectively avoid the technical problem of affecting the processing performance due to the large amount of processing resources consumed by data compression, thereby saving processing resources and improving processing performance.
[0093] In addition, compared with traditional content appending methods, the technical solution provided by the present disclosure can reduce the application of memory space required for caching data and the CPU performance consumed by compressing data, and has great performance improvements in application scenarios such as large-scale traffic transmission and distributed efficient storage.
[0094] In addition, the technical solution provided by the present disclosure can effectively improve the user experience.
[0095] It should be noted that for the aforementioned method embodiments, for simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the present disclosure is not limited by the order of the actions described, because according to the present disclosure, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present disclosure.
[0096] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0097] Figure 3 is a schematic diagram according to the third embodiment of the present disclosure, as shown in Figure 3As shown. The compressed format data processing device 300 of this embodiment may include a data acquisition unit 301, a data processing unit 302 and a data construction unit 303. The data acquisition unit 301 is used to acquire first data and second data to be merged, wherein the first data and the second data are data with a specified compression format; the data processing unit 302 is used to obtain data verification information of the uncompressed original data of the first data and data verification information of the uncompressed original data of the second data based on the first data and the second data respectively; the data construction unit 303 is used to obtain third data with the specified compression format based on the data verification information of the uncompressed original data of the first data, the data verification information of the uncompressed original data of the second data, the compressed data blocks in the first data, the compressed data blocks in the second data and the constructed basic information of the data of the specified compression format.
[0098] It should be noted that part or all of the compressed format data processing device of this embodiment may be an application located in the local terminal, or may also be a functional unit such as a plug-in or software development kit (SDK) set in the application located in the local terminal, or may also be a processing engine located in a network side server, or may also be a distributed system located on the network side, for example, a processing engine or distributed system in a compressed format data processing device on the network side, etc. This embodiment does not specifically limit this.
[0099] It is understandable that the application may be a native program (nativeApp) installed on the local terminal, or may be a webpage program (webApp) of a browser on the local terminal, which is not limited in this embodiment.
[0100] Optionally, in a possible implementation of this embodiment, the data processing unit 302 can be specifically used to decompress the first data and the second data respectively to obtain the uncompressed original data of the first data and the uncompressed original data of the second data; and obtain data verification information of the uncompressed original data of the first data and the data verification information of the uncompressed original data of the second data based on the uncompressed original data of the first data and the uncompressed original data of the second data.
[0101] Optionally, in a possible implementation of this embodiment, the data processing unit 302 can be specifically used to decompress the first data and the second data respectively to obtain data verification information of the uncompressed original data of the first data and data verification information of the uncompressed original data of the second data.
[0102] Optionally, in a possible implementation of this embodiment, the data processing unit 302 can be specifically used to perform data reading processing on the data portion at a specified position in the first data and the data portion at a specified position in the second data, respectively, to obtain data verification information of the uncompressed original data of the first data and data verification information of the uncompressed original data of the second data.
[0103] Optionally, in a possible implementation of this embodiment, the data construction unit 303 can be specifically used to use the constructed basic information of the data in the specified compression format as the data header part of the third data; use the compressed data blocks in the first data and the compressed data blocks in the second data as the data body part of the third data; and use the data verification information of the uncompressed original data of the first data and the data verification information of the uncompressed original data of the second data after merging as the data tail part of the third data.
[0104] It should be noted that Figure 1A The corresponding embodiment method, and Figure 2 The method executed by the web server or the load of the web server in the corresponding embodiment can be implemented by the compressed format data processing device provided by this embodiment. Figure 1A and Figure 2 The relevant contents in the corresponding embodiments will not be repeated here.
[0105] In this embodiment, the first data and the second data to be merged are acquired by the data acquisition unit, and the first data and the second data are data with a specified compression format. Then, the data processing unit obtains data verification information of the uncompressed original data of the first data and data verification information of the uncompressed original data of the second data according to the first data and the second data, respectively, so that the data construction unit can obtain the third data with the specified compression format according to the data verification information of the uncompressed original data of the first data, the data verification information of the uncompressed original data of the second data, the compressed data blocks in the first data, the compressed data blocks in the second data and the constructed basic information of the data in the specified compression format. Since the compressed data blocks in the first data and the compressed data blocks in the second data are directly extracted as the main body of the data in the merged third data during the merging process, and there is no need to compress any data, it can effectively avoid the technical problem of affecting the processing performance due to the large amount of processing resources consumed by data compression, thereby saving processing resources and improving processing performance.
[0106] In addition, compared with traditional content appending methods, the technical solution provided by the present disclosure can reduce the application of memory space required for caching data and the CPU performance consumed by compressing data, and has great performance improvements in application scenarios such as large-scale traffic transmission and distributed efficient storage.
[0107] In addition, the technical solution provided by the present disclosure can effectively improve the user experience.
[0108] Figure 4 A schematic block diagram of an example electronic device 400 that can be used to implement embodiments of the present disclosure is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are provided as examples only and are not intended to limit the implementation of the present disclosure described and / or claimed herein.
[0109] like Figure 4 As shown, electronic device 400 includes a computing unit 401, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 402 or a computer program loaded from a storage unit 408 into a random access memory (RAM) 403. Various programs and data required for the operation of electronic device 400 can also be stored in RAM 403. Computing unit 401, ROM 402, and RAM 403 are connected to each other via a bus 404. An input / output (I / O) interface 405 is also connected to bus 404.
[0110] Multiple components in the electronic device 400 are connected to the I / O interface 405, including an input unit 406, such as a keyboard, a mouse, etc.; an output unit 407, such as various types of displays, speakers, etc.; a storage unit 408, such as a magnetic disk, an optical disk, etc.; and a communication unit 409, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 409 allows the electronic device 400 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0111] The computing unit 401 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 401 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, digital signal processors (DSPs), and any appropriate processors, controllers, microcontrollers, etc. The computing unit 401 performs the various methods and processes described above, such as the method for processing compressed format data. For example, in some embodiments, the method for processing compressed format data can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as the storage unit 408. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 400 via the ROM 402 and / or the communication unit 409. When the computer program is loaded into the RAM 403 and executed by the computing unit 401, one or more steps of the method for processing compressed format data described above can be performed. Alternatively, in other embodiments, the computing unit 401 can be configured to perform the method for processing compressed format data by any other appropriate means (e.g., by means of firmware).
[0112] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.
[0113] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0114] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0115] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0116] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.
[0117] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, establishing a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and VPS services ("Virtual Private Servers" or simply "VPS"). The server may also be a server in a distributed system or a server integrated with blockchain.
[0118] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not a limitation herein.
[0119] The above specific embodiments do not constitute a limitation on the scope of protection of this disclosure. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the scope of protection of this disclosure.
Claims
1. A method for processing compressed format data, characterized in that: include: Acquire first data and second data to be merged, where the first data and the second data are data in a specified compression format; Obtaining data verification information of uncompressed original data of the first data and data verification information of uncompressed original data of the second data based on the first data and the second data respectively; According to the data verification information of the uncompressed original data of the first data, the data verification information of the uncompressed original data of the second data, the compressed data blocks extracted from the first data, the compressed data blocks extracted from the second data and the constructed data basic information of the specified compression format, third data with the specified compression format is obtained, including: using the constructed data basic information of the specified compression format as the data header part of the third data; using the compressed data blocks extracted from the first data and the compressed data blocks extracted from the second data as the data body part of the third data; using the data verification information of the uncompressed original data of the first data and the data verification information of the uncompressed original data of the second data after being merged as the data tail part of the third data.
2. The method according to claim 1, characterized in that The obtaining, based on the first data and the second data, data verification information of the uncompressed original data of the first data and data verification information of the uncompressed original data of the second data, respectively, includes: Decompressing the first data and the second data respectively to obtain uncompressed original data of the first data and uncompressed original data of the second data; Data verification information of the uncompressed original data of the first data and data verification information of the uncompressed original data of the second data are obtained according to the uncompressed original data of the first data and the uncompressed original data of the second data respectively.
3. The method according to claim 1, characterized in that The obtaining, based on the first data and the second data, data verification information of the uncompressed original data of the first data and data verification information of the uncompressed original data of the second data, respectively, includes: Decompression processing is performed on the first data and the second data respectively to obtain data verification information of uncompressed original data of the first data and data verification information of the uncompressed original data of the second data.
4. The method according to claim 1, wherein The obtaining, based on the first data and the second data, data verification information of the uncompressed original data of the first data and data verification information of the uncompressed original data of the second data, respectively, includes: Data reading processing is performed on the data portion at a specified position in the first data and the data portion at a specified position in the second data respectively to obtain data verification information of the uncompressed original data of the first data and data verification information of the uncompressed original data of the second data.
5. A device for processing compressed format data, characterized in that: include: a data acquisition unit, configured to acquire first data and second data to be merged, wherein the first data and the second data are data in a specified compression format; A data processing unit, configured to obtain data verification information of uncompressed original data of the first data and data verification information of uncompressed original data of the second data based on the first data and the second data, respectively; A data construction unit is used to obtain third data having the specified compression format based on data verification information of the uncompressed original data of the first data, data verification information of the uncompressed original data of the second data, compressed data blocks extracted from the first data, compressed data blocks extracted from the second data, and constructed data basic information of the specified compression format, including: using the constructed data basic information of the specified compression format as the data header part of the third data; using the compressed data blocks extracted from the first data and the compressed data blocks extracted from the second data as the data body part of the third data; and using the data verification information of the uncompressed original data of the first data and the data verification information of the uncompressed original data of the second data after being merged as the data tail part of the third data.
6. The device according to claim 5, characterized in that The data processing unit is specifically used for Decompressing the first data and the second data respectively to obtain uncompressed original data of the first data and uncompressed original data of the second data; as well as Data verification information of the uncompressed original data of the first data and data verification information of the uncompressed original data of the second data are obtained according to the uncompressed original data of the first data and the uncompressed original data of the second data respectively.
7. The device according to claim 5, characterized in that The data processing unit is specifically used for Decompression processing is performed on the first data and the second data respectively to obtain data verification information of uncompressed original data of the first data and data verification information of the uncompressed original data of the second data.
8. The device according to claim 5, characterized in that The data processing unit is specifically used for Data reading processing is performed on the data portion at a specified position in the first data and the data portion at a specified position in the second data respectively to obtain data verification information of the uncompressed original data of the first data and data verification information of the uncompressed original data of the second data.
9. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 4.
10. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to cause the computer to execute the method according to any one of claims 1-4.
Citation Information
Patent Citations
File compression?method and device, file decompression method and device, and server
CN103384884A
Method and device for making upgrade package and method and device for upgrading file
CN107391145A