Data processing method and device, equipment, storage medium and program product
By using the shared compression head to determine the target sub-compression block for decompression processing in the merge compressed block, the read operation overhead problem during merge compressed data access is solved, and data access efficiency and system resource utilization are improved.
Patent Information
- Application Number
- CN202311865226.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-29
- Publication Date
- 2025-07-01
AI Technical Summary
When existing merge compressed data access, the entire merge compressed block needs to be decompressed, resulting in an increase in read operation overhead.
By obtaining the data identification in the data access request, the merged compression block to which the target data belongs, and using a shared compression head to determine the target sub-compression block from multiple sub-compression blocks for decompression processing, avoiding decompression of the entire merged compression block.
Reduce the overhead of reading operation during compression, improve data access efficiency, and reduce system resource burden.
Smart Images

Figure CN120238133A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of storage technologies, and in particular, to a data processing method, apparatus, device, storage medium, and program product. Background Art
[0002] Data compression can reduce the volume of data, save storage resources such as hard disks and cloud storage, and improve efficiency when storing, transmitting, and processing data. Data compression has many advantages. Currently, common data compression methods are mainly divided into differential compression and merge compression.
[0003] Taking merge compression as an example, when accessing merge-compressed data, it is first necessary to decompress the merge-compressed data. For example, when the data to be accessed is located at the end of the merge-compressed data, it is necessary to decompress the entire merge-compressed data block, and only after decompression can the required data be accessed.
[0004] However, the above decompression method will increase the overhead of the read operation. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a data processing method, apparatus, device, storage medium, and program product that can reduce the read operation overhead during decompression of compressed data.
[0006] In a first aspect, this application provides a data processing method, and the method includes:
[0007] Obtain a data access request, where the data access request carries a data identifier of target data;
[0008] According to the data identifier, determine the merge compression block to which the target data belongs, where the merge compression block includes a plurality of sub-compression blocks and at least one shared compression header;
[0009] According to the data identifier and the shared compression header, determine a target sub-compression block from the plurality of sub-compression blocks, and perform decompression processing on the target sub-compression block to obtain the target data.
[0010] When receiving a data access request for target data in an embodiment of this application, a target sub-compression block is determined from the plurality of sub-compression blocks of the merge compression block to which the target data belongs. The target sub-compression block is the sub-compression block that contains the target data. Then, decompression processing is performed on the target sub-compression block to obtain the target data. Since the target sub-compression block has a smaller data volume than the merge compression block, this method reduces the read operation overhead during decompression compared to the current method of decompressing the merge compression block after determining the merge compression block to which the target data belongs. If the target data is located at the end of the merge compression block, the entire merge-compressed data block needs to be decompressed, and only after decompression can the required data be accessed.
[0011] In one embodiment, the shared compression header includes index information. Determining a target sub-compressed block from the multiple sub-compressed blocks according to the data identifier and the shared compression header includes:
[0012] Querying the index information in the at least one shared compression header according to the data identifier to determine a target shared compression header to which the index information containing the data identifier belongs;
[0013] Determining the target sub-compressed block from the sub-compressed blocks corresponding to the target shared compression header.
[0014] In this embodiment, by querying the index information in the at least one shared compression header according to the data identifier of the target data, the target sub-compressed block to which the target data belongs can be determined, and determining the target sub-compressed block according to the data identifier is realized.
[0015] In one embodiment, determining the target sub-compressed block from the sub-compressed blocks corresponding to the target shared compression header includes:
[0016] If the target shared compression header corresponds to two sub-compressed blocks located on both sides of the target shared compression header, determining a sub-compressed block identifier corresponding to the data identifier from the index information of the target shared compression header;
[0017] Determining the target sub-compressed block from the two sub-compressed blocks corresponding to the target shared compression header according to the sub-compressed block identifier.
[0018] In this embodiment, the target sub-compressed block can be determined according to the index information of the target shared compression header, without occupying other data resources, avoiding additional overhead.
[0019] In one embodiment, determining the target sub-compressed block from the sub-compressed blocks corresponding to the target shared compression header includes:
[0020] If the target shared compression header corresponds to one sub-compressed block located on one side of the target shared compression header, taking the sub-compressed block corresponding to the target shared compression header as the target sub-compressed block.
[0021] In this embodiment, when the number of sub-compressed blocks corresponding to the target shared compression header is one, it is determined as the target sub-compressed block. The above can determine the target sub-compressed block when the target shared compression header corresponds to different numbers of sub-compressed blocks, realizing compatibility with different compression methods.
[0022] In one embodiment, decompressing the target sub-compressed block to obtain the target data includes:
[0023] Obtain the target data location corresponding to the data identifier from the index information included in the target shared compression header;
[0024] Decompress the target sub-compressed block according to the target data location to obtain the target data.
[0025] In this embodiment, decompressing the target sub-compressed block according to the target data location can obtain the target data, without decompressing the entire merged compressed block. Since the size of the target sub-compressed block is smaller than that of the merged compressed block, the overhead of the read operation can be reduced.
[0026] In one embodiment, the decompressing the target sub-compressed block according to the target data location to obtain the target data includes:
[0027] Decompress the data in the target sub-compressed block located at and before the target data location to obtain the target data.
[0028] In this embodiment, decompressing the target sub-compressed block according to the target data location stops decompressing when reaching the target data location, without decompressing the entire target sub-compressed block, further reducing the overhead of the read operation.
[0029] In one embodiment, the decompressing the target sub-compressed block according to the target data location to obtain the target data includes:
[0030] Decompress the target sub-compressed block in the direction pointed to by the target shared compression header to the target data location to obtain the target data.
[0031] In this embodiment, decompressing the target sub-compressed block in the direction pointed to by the target shared compression header to the target data location does not require decompressing from the starting position of the merged compressed block, further reducing the overhead of the read operation.
[0032] In one embodiment, the method further includes:
[0033] Obtain the data to be compressed, divide the data to be compressed into multiple data blocks, and perform compression processing on each data block respectively to obtain multiple sub-compressed blocks;
[0034] Generate at least one shared compression header according to the multiple sub-compressed blocks;
[0035] Generate the merged compressed block based on the multiple sub-compressed blocks and the at least one shared compression header.
[0036] In this embodiment, by dividing the data to be compressed into multiple data blocks and then performing combined compression processing on each of them respectively to obtain combined compression blocks, the decompression length can be reduced during data access, thereby reducing the overhead of read operations.
[0037] In one embodiment, generating at least one shared compression header according to the multiple sub-compression blocks includes:
[0038] If the number of the multiple sub-compression blocks is even, determining at least one pair of first sub-compression blocks from the multiple sub-compression blocks, where the pair of first sub-compression blocks includes two sub-compression blocks;
[0039] Generating one shared compression header for each pair of first sub-compression blocks respectively.
[0040] In this embodiment, when the number of multiple sub-compression blocks is even, the multiple sub-compression blocks are divided into multiple pairs of first sub-compression blocks. One shared compression header is generated during compression for each pair of first sub-compression blocks. One shared compression header includes index information of two sub-compression blocks, which can save resources.
[0041] In one embodiment, generating at least one shared compression header according to the multiple sub-compression blocks includes:
[0042] If the number of the multiple sub-compression blocks is odd, determining one first sub-compression block from the multiple sub-compression blocks, determining multiple second sub-compression blocks except the first sub-compression block, and determining at least one pair of second sub-compression blocks from the multiple second sub-compression blocks, where the pair of second sub-compression blocks includes two second sub-compression blocks;
[0043] Generating one shared compression header for the first sub-compression block and generating one shared compression header for each pair of second sub-compression blocks respectively.
[0044] In this embodiment, when the number of multiple sub-compression blocks is odd, the multiple sub-compression blocks are divided into multiple pairs of second sub-compression blocks and one first sub-compression block. One shared compression header is generated for each pair of second sub-compression blocks and one shared compression header is generated for the first sub-compression block. The data processing method can be compatible with different situations of sub-compression blocks.
[0045] In one embodiment, generating the combined compression block based on the multiple sub-compression blocks and the at least one shared compression header includes:
[0046] For each pair of first sub-compression blocks, placing the shared compression header corresponding to the pair of first sub-compression blocks between the two sub-compression blocks of the pair of first sub-compression blocks.
[0047] Second aspect, the present application also provides a data processing method, the method comprising:
[0048] Obtain data to be compressed, divide the data to be compressed into a plurality of data blocks, perform compression processing on each of the data blocks respectively to obtain a plurality of sub-compressed blocks;
[0049] Generate at least one shared compression header according to the plurality of sub-compressed blocks;
[0050] Generate the merged compressed block based on the plurality of sub-compressed blocks and the at least one shared compression header.
[0051] Third aspect, the present application also provides a data transmission method, the method comprising:
[0052] Receive the merged compressed block sent by the sending end;
[0053] Obtain a data access request, the data access request carrying a data identifier of target data;
[0054] According to the data identifier, determine the merged compressed block to which the target data belongs, the merged compressed block comprising a plurality of sub-compressed blocks and at least one shared compression header;
[0055] According to the data identifier and the shared compression header, determine a target sub-compressed block from the plurality of sub-compressed blocks, and perform decompression processing on the target sub-compressed block to obtain the target data.
[0056] Fourth aspect, the present application also provides a data transmission method, the method comprising:
[0057] Obtain data to be compressed, divide the data to be compressed into a plurality of data blocks, perform compression processing on each of the data blocks respectively to obtain a plurality of sub-compressed blocks;
[0058] Generate at least one shared compression header according to the plurality of sub-compressed blocks;
[0059] Generate the merged compressed block based on the plurality of sub-compressed blocks and the at least one shared compression header;
[0060] Send the merged compressed block to the receiving end.
[0061] Fifth aspect, the present application also provides a data processing device, the device comprising:
[0062] A first acquisition module, configured to acquire a data access request, the data access request carrying a data identifier of target data;
[0063] A first determination module, configured to determine, according to the data identifier, a merged compression block to which the target data belongs, where the merged compression block includes a plurality of sub-compression blocks and at least one shared compression header;
[0064] A second determination module, configured to determine, according to the data identifier and the shared compression header, a target sub-compression block from the plurality of sub-compression blocks, and perform decompression processing on the target sub-compression block to obtain the target data.
[0065] In a sixth aspect, the present application further provides a data processing device, where the device includes:
[0066] A second acquisition module, configured to acquire data to be compressed, divide the data to be compressed into a plurality of data blocks, and perform compression processing on each of the data blocks respectively to obtain a plurality of sub-compression blocks;
[0067] A first generation module, configured to generate at least one shared compression header according to the plurality of sub-compression blocks;
[0068] A second generation module, configured to generate the merged compression block based on the plurality of sub-compression blocks and the at least one shared compression header.
[0069] In a seventh aspect, the present application further provides a data transmission device, where the device includes:
[0070] A receiving module, configured to receive a merged compression block sent by a sending end;
[0071] A third acquisition module, configured to acquire a data access request, where the data access request carries a data identifier of target data;
[0072] A third determination module, configured to determine, according to the data identifier, a merged compression block to which the target data belongs, where the merged compression block includes a plurality of sub-compression blocks and at least one shared compression header;
[0073] A fourth determination module, configured to determine, according to the data identifier and the shared compression header, a target sub-compression block from the plurality of sub-compression blocks, and perform decompression processing on the target sub-compression block to obtain the target data.
[0074] In an eighth aspect, the present application further provides a data transmission device, where the device includes:
[0075] A fourth acquisition module, configured to acquire data to be compressed, divide the data to be compressed into a plurality of data blocks, and perform compression processing on each of the data blocks respectively to obtain a plurality of sub-compression blocks;
[0076] A third generation module, configured to generate at least one shared compression header according to the plurality of sub-compression blocks;
[0077] A fourth generation module, configured to generate the merged compression block based on the multiple sub-compression blocks and the at least one shared compression header;
[0078] A sending module, configured to send the merged compression block to a receiving end.
[0079] In a ninth aspect, the present application further provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, the steps of the method according to any one of the first aspect, the second aspect, the third aspect, or the fourth aspect are implemented.
[0080] In a tenth aspect, the present application further provides a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, the steps of the method according to any one of the first aspect, the second aspect, the third aspect, or the fourth aspect are implemented.
[0081] In an eleventh aspect, the present application further provides a computer program product, including a computer program, and when the computer program is executed by a processor, the steps of the method according to any one of the first aspect, the second aspect, the third aspect, or the fourth aspect are implemented.
[0082] For the above data processing method, apparatus, device, storage medium, and program product, by obtaining a data access request carrying a data identifier of target data, and then, according to the data identifier, determining the merged compression block to which the target data belongs, where the merged compression block includes multiple sub-compression blocks and at least one shared compression header, and then, according to the data identifier and the shared compression header, determining a target sub-compression block from the multiple sub-compression blocks, and performing decompression processing on the target sub-compression block to obtain the target data. In this way, when a data access request for the target data is received, a target sub-compression block is determined among the multiple sub-compression blocks of the merged compression block to which the target data belongs, and the target sub-compression block is the sub-compression block containing the target data, and then, by performing decompression processing on the target sub-compression block, the target data is obtained. Since the target sub-compression block has a smaller data volume than the merged compression block, this method reduces the overhead of the read operation during decompression compared to the current method of performing decompression processing on the merged compression block after determining the merged compression block to which the target data belongs. If the target data is located at the end position of the merged compression block, the entire merged compression data block needs to be decompressed before the required data can be accessed. Description of the Drawings
[0083] In order to more clearly illustrate the technical solutions in the embodiments of the present application or related technologies, the following will briefly introduce the drawings required for use in the description of the embodiments or related technologies. Obviously, the drawings in the following description are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.
[0084] Figure 1 It is a schematic diagram of the merge compression process;
[0085] Figure 2 It is a schematic diagram of the read amplification phenomenon of merge compression;
[0086] Figure 3 It is an application environment diagram of the data processing method in one embodiment;
[0087] Figure 4 It is a schematic flowchart of the data processing method in one embodiment;
[0088] Figure 5 It is a schematic flowchart of the data processing method in one embodiment;
[0089] Figure 6 It is a schematic flowchart of the data processing method in another embodiment;
[0090] Figure 7 It is a schematic flowchart of the data processing method in another embodiment;
[0091] Figure 8 It is a schematic flowchart of the data processing method in another embodiment;
[0092] Figure 9 It is a schematic flowchart of the data processing method in another embodiment;
[0093] Figure 10 It is a schematic flowchart of the data processing method in another embodiment;
[0094] Figure 11 It is a schematic flowchart of the data processing method in another embodiment;
[0095] Figure 12 It is a flowchart of merge compression in another embodiment;
[0096] Figure 13 It is a schematic diagram of merge compression in another embodiment;
[0097] Figure 14 It is a schematic diagram of merge compression data access in another embodiment;
[0098] Figure 15 It is a schematic diagram of merge compression data access in another embodiment;
[0099] Figure 16 It is a schematic flowchart of the data layout method in another embodiment;
[0100] Figure 17 It is a schematic flowchart of the data layout method in another embodiment;
[0101] Figure 18 Schematic diagram of the result of the data layout method in another embodiment;
[0102] Figure 19 Flow diagram of the data processing method in another embodiment;
[0103] Figure 20 Application environment diagram of the data transmission method in another embodiment;
[0104] Figure 21 Flow diagram of the data transmission method in another embodiment;
[0105] Figure 22 Flow diagram of the data transmission method in another embodiment;
[0106] Figure 23 Structural block diagram of the data processing device in one embodiment. Detailed implementation manners
[0107] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0108] With the continuous development of storage technology, the amount of data that users need to store is increasing continuously. According to different application scenarios, it is very important to select a suitable data compression algorithm to compress the data. How to select a suitable compression method depends on the data type, performance requirements and available resources. Currently, the commonly used compression algorithms include differential compression and merge compression. By reducing the volume of the data, storage space or transmission bandwidth can be saved. Both differential compression and merge compression identify and process redundant information in the data to reduce the redundancy of the data. Among them, differential compression focuses on the changes or differences in the data and compresses the data by comparing the differences between two data sets. Differential compression is currently applied to scenarios such as processing historical records, version control, and incremental backups. Merge compression is to identify and eliminate redundant information in the data, merge duplicate data into a single instance, so as to reduce the data storage and transmission costs. Merge compression can be used in compression scenarios for most data types, such as compression of text, images, audio, and video, etc.
[0109] However, when accessing compressed data, it is necessary to decompress the compressed data first. Since compressed data is usually not stored in the logical order of the original data, but is organized according to its internal structure during the merging compression process to improve the compression efficiency. Therefore, during the decompression process, in order to obtain the required data, it may be necessary to decompress more content than the required data, which increases the system's reading time and computing resources, that is, it increases the overhead of the read operation, which is the read amplification phenomenon. Currently, there are some corresponding problems with the solutions to read amplification. For example, in the prefetching method, by predicting the data that may need to be accessed and decompressing it in advance when needed, this method can reduce the latency during reading. However, if the prediction accuracy is not high, it may prefetch unnecessary data incorrectly, which will also waste resources. At the same time, the overhead of prefetching also needs to be considered. If the method of using indexes and metadata is adopted, since creating and maintaining indexes requires additional storage space, and if the index is damaged or incorrect, it may lead to data inconsistency or inaccessibility. At the same time, the query speed of the index may be affected by the size of the index and the data distribution. If the method of data chunking is adopted, by chunking the data to reduce the overhead during reading, because only the required data chunks need to be decompressed. This method can be applied to large datasets, but if the data chunking design is improper, it may lead to excessive chunk management overhead. At the same time, the chunking of data may introduce the problem of data fragmentation. Therefore, any of the above methods is not very efficient in reducing the read amplification phenomenon during the data access process of the compressed data chunks.
[0110] For example, as Figure 1 shown, during the merging compression, multiple data chunks are first taken out from the database to form a merged chunk, and then the data in the merged chunk is analyzed through the merging compression algorithm. During the process, the identification and replacement of duplicate data may be involved, so as to effectively compress the data and reduce the volume of the data. During the compression process, a data index is generated at the same time. When it is necessary to access the original data, only by reading the index can one know the position of the target data in the merged compressed data, and then decompress it to the corresponding position. However, although the establishment of the index reduces the read amplification phenomenon to a certain extent, if the accessed data is at the end of the compressed chunk, a large amount of resource waste will still be caused. As Figure 2 shown, if it is necessary to access the data at the end of the original merged chunk, it is necessary to decompress the entire merged compressed data chunk, which will cause serious resource waste. If the data in the merged chunk is randomly distributed, there will be a 50% probability that the target data to be accessed in any user access request will be in the second half of the merged compressed data. In this way, the system will consume a large amount of resources to decompress the unnecessary data, resulting in a serious read amplification phenomenon.
[0111] In view of this, in the embodiment of the present application, a data access request carrying a data identifier of target data is obtained, and then, according to the data identifier, a merged compression block to which the target data belongs is determined, where the merged compression block includes a plurality of sub-compression blocks and a shared compression header. Then, according to the data identifier and the shared compression header, a target sub-compression block is determined from the plurality of sub-compression blocks, and the target sub-compression block is decompressed to obtain the target data. In this way, when a data access request for the target data is received, the target sub-compression block is determined among the plurality of sub-compression blocks of the merged compression block to which the target data belongs, and the target sub-compression block is the sub-compression block containing the target data. Then, decompressing the target sub-compression block obtains the target data. Since the amount of data in the target sub-compression block is smaller than that of the merged compression block, this method reduces the overhead of the read operation during decompression compared to the current method of decompressing the merged compression block after determining the merged compression block to which the target data belongs. If the target data is located at the end position of the merged compression block, the entire merged compression data block needs to be decompressed before the required data can be accessed.
[0112] The data processing method provided in this embodiment can be applied to a computer device, which can be a server, and its internal structure diagram can be as Figure 3 shown. The computer device includes a processor, a memory, an input / output interface (Input / Output, abbreviated as I / O), and a communication interface. Among them, the processor, the memory, and the input / output interface are connected through a system bus, and the communication interface is connected to the system bus through the input / output interface. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system, a computer program, and a database. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The database of the computer device is used to store data to be compressed. The input / output interface of the computer device is used to exchange information between the processor and external devices. The communication interface of the computer device is used to communicate with an external terminal through a network connection. When the computer program is executed by the processor, a data processing method is implemented.
[0113] Those skilled in the art can understand that Figure 3 the structure shown in
[0114] is only a block diagram of a part of the structure related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0114] In an exemplary embodiment, as Figure 4 shown, a data processing method is provided, and this method is applied to Figure 3Taking the computer device in [as an example for illustration, including the following steps 401 to 403. Among them:
[0115] Step 401, obtain a data access request.
[0116] Among them, the data access request carries the data identifier of the target data. The target data is also the data in the merged and compressed block that the user needs to access. Since the merged and compressed block has undergone the merging and compression process, when the user accesses the target data, the merged and compressed block needs to be decompressed to obtain the target data in the merged and compressed block.
[0117] Step 402, determine the merged and compressed block to which the target data belongs according to the data identifier.
[0118] Among them, the merged and compressed block includes multiple sub-compressed blocks and at least one shared compression header. Optionally, the storage system may include multiple merged and compressed blocks. When obtaining a data access request, it is first necessary to determine the merged and compressed block where the target data is located. By parsing the metadata of each merged and compressed block, that is, the original data, the data range and location included in each merged and compressed block can be determined, so that the merged and compressed block to which the target data belongs can be determined according to the data range included in each merged and compressed block for the data identifier of the target data.
[0119] Step 403, determine the target sub-compressed block from multiple sub-compressed blocks according to the data identifier and the shared compression header, and perform decompression processing on the target sub-compressed block to obtain the target data.
[0120] After determining the merged and compressed block where the target data is located, the merged and compressed block includes a shared compression header. The shared compression header includes the data information and the location information of the data in each sub-compressed block. According to the data identifier of the target data and the shared compression header, the target sub-compressed block is determined. The target sub-compressed block is also the sub-compressed block that contains the target data. Then, by decompressing the target sub-compressed block, instead of decompressing the entire merged and compressed block, the target data can be obtained.
[0121] In the above embodiments, a data access request carrying a data identifier of target data is obtained, and then, according to the data identifier, a merged compression block to which the target data belongs is determined, where the merged compression block includes a plurality of sub-compression blocks and at least one shared compression header. Then, according to the data identifier and the shared compression header, a target sub-compression block is determined from the plurality of sub-compression blocks, and the target sub-compression block is decompressed to obtain the target data. In this way, when a data access request for the target data is received, the target sub-compression block is determined among the plurality of sub-compression blocks of the merged compression block to which the target data belongs. The target sub-compression block is the sub-compression block containing the target data. Then, by decompressing the target sub-compression block, the target data is obtained. Since the amount of data in the target sub-compression block is smaller than that of the merged compression block, this method reduces the overhead of the read operation during decompression compared with the current method of decompressing the merged compression block after determining the merged compression block to which the target data belongs. If the target data is located at the end position of the merged compression block, the entire merged compression data block needs to be decompressed before the required data can be accessed.
[0122] In one embodiment, the shared compression header includes index information. Based on Figure 4 the embodiments shown, the process of determining the target sub-compression block from the plurality of sub-compression blocks according to the data identifier and the shared compression header, step 402 includes Figure 5 the steps 501 and 502 shown as follows:
[0123] Step 501, query the index information in at least one shared compression header according to the data identifier to determine the target shared compression header to which the index information containing the data identifier belongs.
[0124] Among them, the shared compression header includes index information, and the index information may include each data identifier included in the merged compression block, as well as the start position, size, and compression method of each data block, etc. According to the data identifier of the target data, a match is made in the index information of at least one shared compression header, and the shared compression header including the index information containing the data identifier is determined as the target shared compression header.
[0125] Step 502, determine the target sub-compression block from the sub-compression blocks corresponding to the target shared compression header.
[0126] After determining the target shared compression header, the sub-compression blocks corresponding to the target shared compression header can be determined, and further the target sub-compression block to which the target data belongs can be determined.
[0127] In this embodiment, by querying the index information in at least one shared compression header according to the data identifier of the target data, the target sub-compression block to which the target data belongs can be determined, realizing the determination of the target sub-compression block according to the data identifier.
[0128] Optionally, the sub-compression blocks corresponding to the target shared compression header can be two sub-compression blocks or one sub-compression block. When the sub-compression blocks corresponding to the target shared compression header are two sub-compression blocks, the steps for determining the target sub-compression block from the sub-compression blocks corresponding to the target shared compression header are as Figure 6 shown, including step 601 and step 602:
[0129] Step 601, if the target shared compression header corresponds to two sub-compression blocks located on both sides of the target shared compression header, determine the sub-compression block identifier corresponding to the data identifier from the index information of the target shared compression header.
[0130] Optionally, when the target shared compression header corresponds to two sub-compression blocks, that is, the case where two sub-compression blocks share a shared compression header after compression, the sub-compression block identifier corresponding to the data identifier of the target data can be determined according to the index information of the target shared compression header.
[0131] Optionally, the target data position of the target data can also be determined according to the index information. Since the position ranges of the data included in multiple sub-compression blocks are different, by comparing the target data position with the position ranges of the data included in each sub-compression block, the sub-compression block where the target data is located can be determined, which is the target sub-compression block. For example, taking the case where the target shared compression header corresponds to two sub-compression blocks as an example, the two sub-compression blocks include the first sub-compression block and the second sub-compression block. The data positions of the first sub-compression block and the second sub-compression block in the merged compression block are on both sides of the data position of the shared compression header. Among them, the first sub-compression block and the second sub-compression block share the shared compression header, and the data positions of the first sub-compression block and the second sub-compression block in the merged compression block are on both sides of the data position of the shared compression header. That is, the middle data position of the merged compression block can be the position of the shared compression header. The part from the starting data position of the merged compression block to the position of the shared compression header is the first sub-compression block, and the part from the position of the shared compression header of the merged compression block to the ending data position is the second sub-compression block. Compare the target data position with the position of the shared compression header. When it is less than the position of the shared compression header, that is, between the starting data position and the middle data position of the merged compression block, the target sub-compression block is the first sub-compression block. Compare the target data position with the position of the shared compression header. When it is greater than the position of the shared compression header, that is, between the middle data position and the ending data position of the merged compression block, the target sub-compression block is the second sub-compression block.
[0132] Step 602, determine the target sub-compression block from the two sub-compression blocks corresponding to the target shared compression header according to the sub-compression block identifier.
[0133] According to the sub-compression block identifier, the sub-compression block where the target data is located can be determined, that is, the target sub-compression block.
[0134] In this embodiment, the target sub-compression block can be determined according to the index information of the target shared compression header, without occupying other data resources and avoiding additional overhead.
[0135] When the target shared compression header corresponds to one sub-compression block, step 502 may include step A1:
[0136] Step A1: If the target shared compression header corresponds to one sub-compression block located on one side of the target shared compression header, then use the sub-compression block corresponding to the target shared compression header as the target sub-compression block.
[0137] If the target shared compression header corresponds to one sub-compression, then use the sub-compression block corresponding to the target shared compression header as the target sub-compression block where the target data is located.
[0138] In this embodiment, when there is one sub-compression block corresponding to the target shared compression header, it is determined as the target sub-compression block. The above can determine the target sub-compression block when the target shared compression header corresponds to different numbers of sub-compression blocks, achieving compatibility with different compression methods.
[0139] In one embodiment, based on the Figure 7 illustrated embodiment, the steps of decompressing the target sub-compression block may include the following steps 701 and 702:
[0140] Step 701, obtain the target data position corresponding to the data identifier from the index information included in the target shared compression header.
[0141] After determining the target shared compression header, the target data position of the target data can be determined according to the data identifier in the index information.
[0142] Step 702, perform decompression processing on the target sub-compression block according to the target data position to obtain the target data.
[0143] After determining the target sub-compression block, according to the target data position, decompress the target sub-compression block. Optionally, decompress the data located at and before the target data position in the target sub-compression block to obtain the target data. That is, it is not necessary to decompress the entire target sub-compression block, only the data located at and before the target data position needs to be decompressed, and the decompression process stops after obtaining the target data.
[0144] In this embodiment, the target sub-compression block is decompressed according to the target data position, and the decompression stops when reaching the target data position, without decompressing the entire target sub-compression block, further reducing the overhead of the read operation.
[0145] Optionally, decompress the target sub-compression block in the direction pointed by the target shared compression header to the target data position to obtain the target data.
[0146] Optionally, taking the sub-compression blocks corresponding to the above-mentioned shared compression header including the first sub-compression block and the second compression block as an example, the decompression of the target sub-compression block includes the following two cases:
[0147] In the first case, if the target sub-compression block is the first sub-compression block, optionally, if the first sub-compression block is located on the left side of the shared compression header, during decompression, start decompressing from the position of the shared compression header to the left direction, that is, when the target sub-compression block is the first sub-compression block, start decompressing from the position of the shared compression header to the left until reaching the target data position and stop.
[0148] In the second case, if the target sub-compression block is the second sub-compression block, if the first sub-compression block is located on the right side of the shared compression header, start decompressing from the position of the shared compression header to the right until reaching the target data position and stop.
[0149] In this embodiment, decompressing the target sub-compression block along the direction from the target shared compression header to the target data position does not require decompressing from the starting position of the merged compression block, further reducing the overhead of the read operation. Further, during multiple accesses, the average access time of the system can be reduced, and the overall access speed of the system is faster.
[0150] In one embodiment, as Figure 8 shown, the steps of obtaining the merged compression block further include steps 801 to 803:
[0151] Step 801, obtain the data to be compressed, divide the data to be compressed into multiple data blocks, and perform compression processing on each data block respectively to obtain multiple sub-compression blocks.
[0152] Among them, the data to be compressed includes the target data. The data to be compressed may include multiple data. Divide the data to be compressed into multiple data blocks, and then perform compression processing on each data block respectively to obtain multiple sub-compression blocks.
[0153] Step 802, generate at least one shared compression header according to the multiple sub-compression blocks.
[0154] During the compression process, at least one shared compression header is generated from the multiple sub-compression blocks. Optionally, the multiple sub-compression blocks may be an even number or an odd number.
[0155] Step 803, generate the merged compression block based on the multiple sub-compression blocks and the at least one shared compression header.
[0156] In this embodiment, the data to be compressed is divided into multiple data blocks, and then merged and compressed respectively to obtain a merged compressed block, so that the decompression length can be reduced during data access, thereby reducing the overhead of read operations.
[0157] In one embodiment, when the number of multiple sub-compressed blocks is even, step 802 may further include step 901 and step 902:
[0158] Step 901, if the number of multiple sub-compressed blocks is even, determine at least one pair of first sub-compressed blocks from the multiple sub-compressed blocks.
[0159] Wherein, a pair of first sub-compressed blocks includes two sub-compressed blocks. Since the number of multiple sub-compressed blocks is even, that is, it can be divided into multiple pairs of first sub-compressed blocks.
[0160] Step 902, generate a shared compression header for each pair of first sub-compressed blocks respectively.
[0161] For the two sub-compressed blocks in each pair of first sub-compressed blocks, a shared compression header is generated during the compression process. The two sub-compressed blocks are compressed from the division position towards both ends. The shared compression header includes the index information of the two sub-compressed blocks. The shared compression header may also include information such as the size of each data block in the merged compressed block and the compression method.
[0162] For each pair of first sub-compressed blocks, place the shared compression header corresponding to the pair of first sub-compressed blocks between the two sub-compressed blocks of the pair of first sub-compressed blocks.
[0163] In this embodiment, when the number of multiple sub-compressed blocks is even, the multiple sub-compressed blocks are divided into multiple pairs of first sub-compressed blocks. Each pair of first sub-compressed blocks generates a shared compression header during compression. A shared compression header includes the index information of two sub-compressed blocks, which can save resources.
[0164] In one embodiment, when the number of multiple sub-compressed blocks is odd, step 802 may further include step 1001 and step 1002:
[0165] Step 1001, if the number of multiple sub-compressed blocks is odd, determine a first sub-compressed block from the multiple sub-compressed blocks, and determine multiple second sub-compressed blocks other than the first sub-compressed block, and determine at least one pair of second sub-compressed blocks from the multiple second sub-compressed blocks.
[0166] Wherein, a pair of second sub-compressed blocks includes two second sub-compressed blocks. If the number of multiple sub-compressed blocks is odd, the multiple sub-compressed blocks can be divided into multiple pairs of second sub-compressed blocks and a first sub-compressed block.
[0167] Step 1002: Generate a shared compression header for the first sub-compression block, and generate a shared compression header for each pair of second sub-compression blocks.
[0168] Then, during the compression process, each of the two sub-compression blocks in each pair of second sub-compression blocks generates a shared compression header, and the first sub-compression block generates a shared compression header during the compression process.
[0169] For each pair of second sub-compression blocks, place the shared compression header corresponding to the pair of second sub-compression blocks between the two sub-compression blocks of the pair of second sub-compression blocks.
[0170] In this embodiment, when the number of multiple sub-compression blocks is odd, divide the multiple sub-compression blocks into multiple pairs of second sub-compression blocks and one first sub-compression block. Each pair of second sub-compression blocks generates a shared compression header, and one first sub-compression block generates a shared compression header. The data processing method can be compatible with different situations of sub-compression blocks.
[0171] In the embodiments of the present application, please refer to Figure 11 , which shows a flowchart of the data processing method provided by the embodiments of the present application. The data processing method includes the following steps:
[0172] Step 1101: Obtain the data to be compressed, divide the data to be compressed into multiple data blocks, and perform compression processing on each data block respectively to obtain multiple sub-compression blocks.
[0173] Step 1102: Generate at least one shared compression header according to the multiple sub-compression blocks.
[0174] Step 1103: Generate a combined compression block based on the multiple sub-compression blocks and at least one shared compression header.
[0175] Step 1104: Obtain a data access request.
[0176] Step 1105: Determine the combined compression block to which the target data belongs according to the data identifier.
[0177] Step 1106: Query the index information in at least one shared compression header according to the data identifier to determine the target shared compression header to which the index information containing the data identifier belongs.
[0178] Step 1107: Determine the target sub-compression block from the sub-compression blocks corresponding to the target shared compression header.
[0179] Step 1108: And perform decompression processing on the target sub-compression block to obtain the target data.
[0180] In the embodiments of the present application, an example is given for the process of combined compression of the data processing method of the present application. Taking 2 sub-compression blocks and 1 shared compression header as an example, asFigure 12 As shown, at the beginning of the compression process, key parameters are initialized, including the size m of the unit data block for combined compression and the number n of unit data blocks in a combined block, etc. The setting of the parameters can be adjusted according to different usage scenarios and compression effects. Then, n unit data blocks are extracted from the original database. The n unit data blocks are formed into a combined block with a size of n to ensure that each unit data block can be compressed fully and effectively. Next, the n unit data blocks are divided into two parts, the L block and the R block. Among them, the L block contains the data blocks in the first half of the combined block, and the R block contains the data blocks in the second half of the combined block. It can be understood that the L block is also the first sub-compression block in the above embodiment, and the R block is also the second sub-compression block in the above embodiment. Finally, combined compression is performed separately from the middle data position towards both ends to ensure that each unit data block is correctly processed and compressed. After the combined compression is completed, the compressed data of the L block and the R block share a compression header, and the shared compression header can improve the efficiency of the decompression process and ensure the accuracy of the decompression process. The combined compression process of the L block and the R block is as Figure 13 shown. The compression ratios of different unit data blocks are different. As Figure 11 shown, in the combined compression block after combined compression, the black positions share the position of the compression header.
[0181] After the combined block is compressed to generate the combined compression block, the access process of the combined compression block can be as Figure 14 shown. When the user needs to access the combined compression block, first, it is necessary to determine the combined compression block where the target data to be accessed is located. Then, it is judged whether the target data is in the L block or the R block of the combined compression block. This can be achieved by comparing the position of the target data with the middle position of the combined compression block. Next, the index information is read to obtain the target data position of the target data. Finally, the L block or the R block where the target data is located is decompressed to the target data position determined by the index information, and the target data is read and returned to the user. In this way, the user can access and use specific partial data in the combined compression block, and then end this access process and wait for the next access request.
[0182] Among them, the access schematic diagram of the compressed data is as Figure 15 shown. When the user needs to access the combined compression block, the data can be decompressed from the middle towards both ends until the target data position is decompressed, so as to obtain the target data. Compared with the conventional way of decompressing from the front to the back, when the target data is located in the second half of the combined compression block, the two-way decompression mode can effectively shorten the decompression length. Therefore, it can at least reduce half of the read amplification overhead, and the system resources can be utilized more effectively, thus accelerating the access speed of the compressed data. Therefore, this two-way decompression method has high practicability.
[0183] Optionally, before processing the data to be compressed, in order to further improve the access efficiency of the compressed data, the data to be compressed can be first divided into multiple data blocks, then the data layout of the multiple data blocks is performed, and then the compression process is carried out. Among them, the method of data layout can be as follows.
[0184] In one embodiment, please refer to Figure 16 , the steps of performing data layout on the data to be compressed include step 1601 and step 1602:
[0185] Step 1601, obtain multiple data blocks to be compressed, and perform data sorting processing on each data block to obtain multiple sorted data blocks.
[0186] Optionally, the computer device can obtain multiple data blocks to be compressed from the database to be compressed according to the application scenario requirements, and then perform data sorting processing on the multiple data blocks. Optionally, the multiple data blocks can be sorted according to the importance of the data blocks or the usage frequency of the data blocks, etc. In order to improve the data access efficiency, the data blocks with higher importance or higher usage frequency of the data blocks have higher corresponding sorting levels. Optionally, the multiple data blocks can also be sorted according to the degree of attention of the data blocks.
[0187] Step 1602, perform compression processing on the multiple sorted data blocks.
[0188] After performing data sorting on each data block to obtain multiple sorted data blocks, compression processing can be performed on the multiple sorted data blocks. Optionally, compression processing can be performed by merge compression, or other compression methods can be used for compression processing. Taking merge compression as an example for illustration, merge compression reduces the space or bandwidth required for data storage and transmission by identifying and eliminating redundant information in the data blocks and merging duplicate data blocks into a single instance. Merge compression is usually used for lossless compression, that is, compressing and decompressing data will not cause any information loss. Merge compression can be applied to various data types, including text, images, audio, and video. The merge compression method can significantly reduce the system storage space requirements in the case of storing a large amount of duplicate data.
[0189] In the above embodiments, by obtaining a plurality of data blocks to be compressed and performing data sorting processing on each data block to obtain a plurality of sorted data blocks, and then performing compression processing on the plurality of sorted data blocks. In this way, before performing compression processing on the plurality of data blocks to be compressed, the present application embodiment first performs data sorting processing on each data block. It can be understood that during the sorting process, the data with higher attention can be arranged in the front positions of the plurality of data blocks. In this way, during the data access process, since the data with higher attention is arranged in the more front positions, the amount of data that needs to be decompressed during data access is smaller. Therefore, the efficiency is higher than that of the data access process in the random distribution method, and the waste of computing resources and time is less.
[0190] In one embodiment, based on Figure 16 the embodiments shown, the steps of performing data sorting processing on each of the data blocks to obtain a plurality of sorted data blocks include A1:
[0191] Step A1, perform data sorting processing on each data block according to the first feature information of each data block to obtain a plurality of sorted data blocks.
[0192] Among them, the first feature information is used to characterize the attention degree of the data block. The attention degree is also the degree of attention of the user to the data block. Perform data sorting processing on each data block according to the attention degree. Optionally, the level of attention degree is positively correlated with the sorting priority of the data block, that is, the data block with high attention degree is arranged in the front position of the plurality of data blocks, and the data block with low attention degree is arranged in the rear position of the plurality of data blocks.
[0193] In this embodiment, perform data sorting on each data block according to the first feature information. Since the first feature information characterizes the attention degree of the data block and arranges the data blocks according to the attention degree, it can improve the efficiency of accessing compressed data compared with the random distribution method.
[0194] Optionally, as Figure 17 shown, step A1 may include step 1701 and step 1702:
[0195] Step 1701, classify the plurality of data blocks according to the first feature information of each data block to determine the data type of the data block.
[0196] Classify the plurality of data blocks according to the attention degree of each data block. Optionally, the data blocks can be classified according to the level of attention degree. For example, the data blocks can be divided into data types with high attention degree and data types with low attention degree, or the data blocks can be divided into data types with high attention degree, data types with normal attention degree, and data types with low attention degree.
[0197] Step 1702: Perform data sorting on each data block according to each data type to obtain multiple sorted data blocks.
[0198] It can be understood that, in order to improve the efficiency and flexibility of data access, the data blocks corresponding to the data types with higher data attention, that is, the data blocks corresponding to the high-attention data types, have higher sorting priorities, and the data blocks corresponding to the data types with lower data attention, that is, the data blocks corresponding to the low-attention data types, have lower sorting priorities.
[0199] In this embodiment, according to the first feature information, the data type is determined, and then the data of each data block is sorted according to the data type to obtain multiple sorted data blocks. The data layout of the data blocks is realized through sorting.
[0200] In a possible embodiment of step 1701, the first feature information includes a data feature value, and step 1701 may be the following step B1:
[0201] Step B1: Classify multiple data blocks according to the data feature value to determine the data type of each data block.
[0202] Among them, the data feature value may be a numerical value. Optionally, the data feature value may be a data heat value. The higher the attention of the data, the higher the data heat value, and the lower the attention of the data, the lower the data heat value. Optionally, the data type may include at least two types. Multiple data blocks are classified according to the feature values of each data block. For example, each data block is divided into a hot data type and a cold data type, or divided into a hot data type, a normal data type, and a cold data type. The embodiments of the present application do not limit this. Among them, the data blocks with high attention can be of the hot data type, and the data blocks with low attention can be of the cold data type. Optionally, multiple data blocks can be classified according to the comparison between the data heat value and different thresholds. For example, as follows.
[0203] If the data heat value is greater than or equal to the first heat threshold, it is determined that the data heat type is the hot data type.
[0204] If the data heat value is less than or equal to the second heat threshold, it is determined that the data heat type is the cold data type, and the first heat threshold is greater than the second heat threshold.
[0205] If the data heat value is less than the first heat threshold and greater than the second heat threshold, it is determined that the data heat type is the normal data type.
[0206] Among them, the first heat threshold and the second heat threshold can be divided according to different operating scenarios. Taking the data characteristic value including the access frequency as an example, the first heat threshold and the second heat threshold can be determined according to the total access volume of the system and the final data layout effect. Optionally, the first heat threshold and the second heat threshold can be determined by multiplying the total access volume by different preset ratios, or can be determined according to the final required data layout effect. For example, if the final data layout effect is that 20% of the data blocks are of the hot data type, 60% of the data blocks are of the normal data type, and 20% of the data blocks are of the cold data type, the specific values of the first heat threshold and the second heat threshold can be determined according to the ratio of each data type and the total access volume. For example, the first heat threshold is set to 100, that is, the data blocks with an access frequency greater than or equal to 100 within a specified period of time are classified as the hot data type, and the second heat threshold is set to 10, that is, the data blocks with an access frequency less than or equal to 10 within a specified period of time are classified as the cold data type, and the data blocks with the remaining access frequencies are classified as the normal data type. By setting the first heat threshold and the second heat threshold, the division of the data heat type of the database is realized. The division method takes into account the total access volume of the system and the final data layout effect, and the division method is more flexible.
[0207] In this embodiment, multiple data blocks are classified according to the data characteristic value of the data block. The data characteristic value can be a numerical value. Classifying the data blocks according to the data characteristic value means that it can be judged and classified according to the data characteristic value and the thresholds of different categories. The classification method of the data blocks is convenient and efficient.
[0208] Optionally, the sorting order of multiple data blocks is positively correlated with the data attention corresponding to the data type to which each data block belongs.
[0209] The higher the data attention corresponding to the data type to which the data block belongs, the higher the sorting priority of the data block, that is, it is arranged in the front position of multiple data blocks. The lower the data attention corresponding to the data type to which the data block belongs, the lower the sorting priority of the data block, that is, it is arranged in the back position of multiple data blocks.
[0210] After determining the data heat type of each data block, all the data blocks to be stored are sorted in the order of the hot data type, the normal data type, and the cold data type. That is, the data blocks classified as the hot data type are located at the beginning position of multiple data blocks, the data blocks classified as the normal data type are located in the middle position of multiple data blocks, and the data blocks classified as the cold data type are located at the end position of multiple data blocks, so as to obtain multiple sorted data blocks. The data blocks within each data type can be randomly distributed. For example Figure 18 As shown, a new data layout is obtained after sorting the randomly distributed data blocks, and 8K is the size of each data block.
[0211] In this embodiment, the higher the data attention corresponding to the data type to which the data block belongs, the higher the sorting priority of the data block; the lower the data attention corresponding to the data type to which the data block belongs, the lower the sorting priority of the data block. Thus, data blocks with high data attention can be sorted to the front positions among multiple data blocks.
[0212] In one embodiment, step 1701 may further include step C1:
[0213] Step C1: Perform data sorting processing on each data block in the order from high to low of the data attention characterized by the first feature information of each data block, so as to obtain multiple sorted data blocks.
[0214] Optionally, when the first feature information of each data block is obtained, it may not be necessary to divide the data types of each data block. Sort each data block according to the order from high to low of the data attention characterized by the first feature information of each data block to obtain multiple sorted data blocks.
[0215] In this embodiment, directly sort each data block in the order from high to low of the data attention characterized by the first feature information. Data with high data attention is closer to the compression header of the compressed data, so that the decompression length can be reduced during the data access process, the waste of decompression resources can be reduced, and the read amplification phenomenon of the compressed data can be effectively reduced.
[0216] In the embodiments of the present application, for each data block, determine the first feature information of each data block according to at least one of the access frequency and access time of each data block.
[0217] Among them, the access frequency is the number of times a user accesses a data block, and the access time may be the duration of the user's access to the data block during one access process.
[0218] According to the access frequency and access duration, the first feature information of the data block can be determined. When the first feature information is a data feature value, optionally, each time a data block is accessed during the operation of the system, the access frequency of the data block is incremented by 1. The higher the access frequency of a data block, the higher the corresponding data feature value of the data block. Optionally, during the operation of the system, the access duration of each data block is statistically counted. The longer the access duration, the higher the corresponding data feature value of the data block. Optionally, the first feature information can also be determined by other means, such as the importance degree of the data block, the data volume of the data block, etc. The embodiments of the present application do not limit this.
[0219] In this embodiment, the access information of the data block can reflect the degree of attention of the data block, and the first feature information of the data block is determined according to at least one of the access frequency and access time of the data block, so that the first feature information can characterize the degree of attention of the data block.
[0220] After the above data layout processing, then perform compression processing on multiple data blocks respectively. Among the obtained multiple sub-compressed blocks, the data with high data attention degree is located at the front position of each sub-compression. In this way, the decompression length can be further reduced during decompression, thereby reducing the overhead of the read operation.
[0221] In one embodiment, the present application also provides a data processing method, as Figure 19 shown, the method includes:
[0222] Step 1901, obtain the data to be compressed, divide the data to be compressed into multiple data blocks, and perform compression processing on each data block respectively to obtain multiple sub-compressed blocks.
[0223] Step 1902, generate at least one shared compression header according to the multiple sub-compressed blocks.
[0224] Step 1903, generate a combined compression block based on the multiple sub-compressed blocks and at least one shared compression header.
[0225] Among them, the specific steps are as described in the above method embodiment, and will not be elaborated here.
[0226] In the embodiment of the present application, a data transmission system is also provided, as Figure 20 shown, the data transmission system includes a sending end and a receiving end. Among them, both the sending end and the receiving end can be a kind of computer device, including a memory and a processor. There is a communication connection between the sending end and the receiving end, and data transmission can be carried out.
[0227] In one embodiment, a data transmission method is also provided. Taking the method applied to the Figure 20 receiving end in the data transmission system shown as an example, as Figure 21 shown, the method includes:
[0228] Step 2101, receive the combined compression block sent by the sending end.
[0229] After the sending end generates the combined compression block, send the combined compression block to the receiving end that needs to access the target data.
[0230] Step 2102, obtain a data access request, and the data access request carries the data identifier of the target data.
[0231] Step 2103: Determine the merged compression block to which the target data belongs according to the data identifier. The merged compression block includes multiple sub-compression blocks and at least one shared compression header.
[0232] Step 2104: Determine the target sub-compression block from the multiple sub-compression blocks according to the data identifier and the shared compression header, and perform decompression processing on the target sub-compression block to obtain the target data.
[0233] The receiving end decompresses the merged compression block according to the above decompression method, so as to obtain the target data.
[0234] In this embodiment, before data transmission, the data is compressed and then transmitted, which can reduce the amount of data transmitted, thereby improving the efficiency of data transmission.
[0235] In one embodiment, a data transmission method is further provided. Taking the sending end in the data transmission system shown as an example, as Figure 20 shown, the method includes: Figure 22 shown, the method includes:
[0236] Step 2201: Obtain the data to be compressed, divide the data to be compressed into multiple data blocks, and perform compression processing on each data block respectively to obtain multiple sub-compression blocks.
[0237] Step 2202: Generate at least one shared compression header according to the multiple sub-compression blocks.
[0238] Step 2203: Generate a merged compression block based on the multiple sub-compression blocks and at least one shared compression header.
[0239] Step 2204: Send the merged compression block to the receiving end.
[0240] In this embodiment, before data transmission, the data is compressed and then transmitted, which can reduce the amount of data transmitted, thereby improving the efficiency of data transmission.
[0241] It should be understood that although the steps in the flowcharts involved in the above-mentioned embodiments are shown in sequence according to the arrows, these steps do not necessarily need to be executed in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-mentioned embodiments may include multiple steps or multiple stages. These steps or stages do not necessarily need to be executed at the same time, but can be executed at different times. The execution order of these steps or stages does not necessarily need to be sequential, but can be executed alternately or alternately with at least a part of other steps or steps in other steps.
[0242] Based on the same inventive concept, an embodiment of the present application further provides a data processing apparatus for implementing the data processing method involved above. The solution provided by this apparatus for solving problems is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the data processing apparatus provided below can refer to the limitations on the data processing method in the foregoing, and will not be elaborated here.
[0243] In an exemplary embodiment, as Figure 23 shown, a data processing apparatus is provided, including:
[0244] A first acquisition module 2301, configured to acquire a data access request, where the data access request carries a data identifier of target data;
[0245] A first determination module 2302, configured to determine, according to the data identifier, a merged compression block to which the target data belongs, where the merged compression block includes a plurality of sub-compression blocks and at least one shared compression header;
[0246] A second determination module 2303, configured to determine a target sub-compression block from the plurality of sub-compression blocks according to the data identifier and the shared compression header, and perform decompression processing on the target sub-compression block to obtain the target data.
[0247] In one embodiment, the shared compression header includes index information. Specifically, the second determination module 2302 is configured to query the index information in the at least one shared compression header according to the data identifier to determine a target shared compression header to which the index information including the data identifier belongs; and determine the target sub-compression block from the sub-compression blocks corresponding to the target shared compression header.
[0248] In one embodiment, specifically, if the target shared compression header corresponds to two sub-compression blocks located on both sides of the target shared compression header, the second determination module 2303 is configured to determine a sub-compression block identifier corresponding to the data identifier from the index information of the target shared compression header; and determine the target sub-compression block from the two sub-compression blocks corresponding to the target shared compression header according to the sub-compression block identifier.
[0249] In one embodiment, specifically, if the target shared compression header corresponds to one sub-compression block located on one side of the target shared compression header, the second determination module 2303 is configured to use the sub-compression block corresponding to the target shared compression header as the target sub-compression block.
[0250] In one embodiment, the second determination module 2303 is specifically configured to obtain a target data position corresponding to the data identifier from the index information included in the target shared compression header; perform decompression processing on the target sub-compression block according to the target data position to obtain the target data.
[0251] In one embodiment, the second determination module 2303 is specifically configured to decompress the data located at and before the target data position in the target sub-compression block to obtain the target data.
[0252] In one embodiment, the second determination module 2303 is specifically configured to perform decompression processing on the target sub-compression block along the direction pointed to by the target shared compression header to the target data position to obtain the target data.
[0253] In one embodiment, the device further includes a compression module, which is configured to obtain data to be compressed, divide the data to be compressed into multiple data blocks, perform compression processing on each of the data blocks respectively to obtain multiple sub-compression blocks; generate at least one shared compression header according to the multiple sub-compression blocks; generate the merged compression block based on the multiple sub-compression blocks and the at least one shared compression header.
[0254] In one embodiment, the compression module is specifically configured to, if the number of the multiple sub-compression blocks is even, determine at least one pair of first sub-compression blocks from the multiple sub-compression blocks, and the pair of first sub-compression blocks includes two sub-compression blocks; generate one of the shared compression headers for each of the pairs of first sub-compression blocks.
[0255] In one embodiment, the compression module is specifically configured to, if the number of the multiple sub-compression blocks is odd, determine one first sub-compression block from the multiple sub-compression blocks, and determine multiple second sub-compression blocks other than the first sub-compression block, determine at least one pair of second sub-compression blocks from the multiple second sub-compression blocks, and the pair of second sub-compression blocks includes two second sub-compression blocks; generate one of the shared compression headers for the first sub-compression block, and generate one of the shared compression headers for each of the pairs of second sub-compression blocks.
[0256] In one embodiment, the compression module is specifically configured to, for each of the pairs of first sub-compression blocks, place the shared compression header corresponding to the pair of first sub-compression blocks between the two sub-compression blocks of the pair of first sub-compression blocks.
[0257] In an exemplary embodiment, the present application further provides a data processing device, and the device includes:
[0258] A second acquisition module, configured to acquire data to be compressed, divide the data to be compressed into multiple data blocks, perform compression processing on each of the data blocks respectively, and obtain multiple sub-compressed blocks;
[0259] A first generation module, configured to generate at least one shared compression header according to the multiple sub-compressed blocks;
[0260] A second generation module, configured to generate the combined compressed block based on the multiple sub-compressed blocks and the at least one shared compression header.
[0261] In an exemplary embodiment, the present application further provides a data transmission device, and the device includes:
[0262] A receiving module, configured to receive the combined compressed block sent by a sending end;
[0263] A third acquisition module, configured to acquire a data access request, where the data access request carries a data identifier of target data;
[0264] A third determination module, configured to determine, according to the data identifier, the combined compressed block to which the target data belongs, where the combined compressed block includes multiple sub-compressed blocks and at least one shared compression header;
[0265] A fourth determination module, configured to determine a target sub-compressed block from the multiple sub-compressed blocks according to the data identifier and the shared compression header, and perform decompression processing on the target sub-compressed block to obtain the target data.
[0266] In an exemplary embodiment, the present application further provides a data transmission device, and the device includes:
[0267] A fourth acquisition module, configured to acquire data to be compressed, divide the data to be compressed into multiple data blocks, perform compression processing on each of the data blocks respectively, and obtain multiple sub-compressed blocks;
[0268] A third generation module, configured to generate at least one shared compression header according to the multiple sub-compressed blocks;
[0269] A fourth generation module, configured to generate the combined compressed block based on the multiple sub-compressed blocks and the at least one shared compression header;
[0270] A sending module, configured to send the combined compressed block to a receiving end.
[0271] Each module in the above data processing device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in a processor in a computer device in a hardware form or be independent of the processor, or can be stored in a memory in the computer device in a software form, so that the processor can call and execute operations corresponding to the above respective modules.
[0272] In an exemplary embodiment, a computer device is provided, which includes a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps in the above method embodiments are implemented.
[0273] In an embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0274] In an embodiment, a computer program product is provided, which includes a computer program. When the computer program is executed by a processor, the steps in the above method embodiments are implemented.
[0275] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data need to comply with relevant regulations.
[0276] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, database, or other medium used in the embodiments provided in the present application can include at least one of non-volatile and volatile memories. Non-volatile memories can include read-only memory (ROM), magnetic tapes, floppy disks, flash memories, optical memories, high-density embedded non-volatile memories, resistive random access memories (ReRAM), magnetoresistive random access memories (MRAM), ferroelectric random access memories (FRAM), phase change memories (PCM), graphene memories, etc. Volatile memories can include random access memory (RAM) or external cache memories, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in the present application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., without limitation. The processors involved in the embodiments provided in the present application can be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., without limitation.
[0277] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.
[0278] The above-described embodiments merely represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the appended claims.
Claims
1. A data processing method, characterized in that The method includes: Obtaining a data access request, where the data access request carries a data identifier of target data; Determining, according to the data identifier, a merged compression block to which the target data belongs, where the merged compression block includes a plurality of sub-compression blocks and at least one shared compression header; Determining a target sub-compression block from the plurality of sub-compression blocks according to the data identifier and the shared compression header, and performing decompression processing on the target sub-compression block to obtain the target data.
2. The method according to claim 1, wherein The shared compression header includes index information. The determining a target sub-compression block from the plurality of sub-compression blocks according to the data identifier and the shared compression header includes: Querying the index information in the at least one shared compression header according to the data identifier to determine a target shared compression header to which the index information containing the data identifier belongs; Determining the target sub-compression block from the sub-compression blocks corresponding to the target shared compression header.
3. The method according to claim 2, wherein The determining the target sub-compression block from the sub-compression blocks corresponding to the target shared compression header includes: If the target shared compression header corresponds to two sub-compression blocks located on both sides of the target shared compression header, determining a sub-compression block identifier corresponding to the data identifier from the index information of the target shared compression header; Determining the target sub-compression block from the two sub-compression blocks corresponding to the target shared compression header according to the sub-compression block identifier.
4. The method according to claim 2, characterized in that, The determining the target sub-compression block from the sub-compression blocks corresponding to the target shared compression header includes: If the target shared compression header corresponds to one sub-compression block located on one side of the target shared compression header, using the sub-compression block corresponding to the target shared compression header as the target sub-compression block.
5. The method according to claim 2, wherein The performing decompression processing on the target sub-compression block to obtain the target data includes: Obtaining a target data position corresponding to the data identifier from the index information included in the target shared compression header; Performing decompression processing on the target sub-compression block according to the target data position to obtain the target data.
6. The method according to claim 5, characterized in that, The performing decompression processing on the target sub-compression block according to the target data position to obtain the target data includes: Decompressing the data in the target sub-compression block located at the target data position and before the target data position to obtain the target data.
7. The method according to claim 6, wherein The performing decompression processing on the target sub-compression block according to the target data position to obtain the target data includes: Performing decompression processing on the target sub-compression block along the direction pointed by the target shared compression header to the target data position to obtain the target data.
8. The method according to any one of claims 1 to 7, characterized in that The method further includes: Obtaining data to be compressed, dividing the data to be compressed into a plurality of data blocks, and respectively performing compression processing on each data block to obtain a plurality of sub-compression blocks; Generating at least one shared compression header according to the plurality of sub-compression blocks; Generating the merged compression block based on the plurality of sub-compression blocks and the at least one shared compression header.
9. The method according to claim 8, wherein The generating at least one shared compression header according to the plurality of sub-compression blocks includes: If the number of the multiple sub-compressed blocks is even, determine at least one first sub-compressed block pair from the multiple sub-compressed blocks, where the first sub-compressed block pair includes two sub-compressed blocks; Generate one such shared compression header for each of the first sub-compressed block pairs.
10. The method according to claim 8, wherein The generating at least one shared compression header according to the multiple sub-compressed blocks includes: If the number of the multiple sub-compressed blocks is odd, determine a first sub-compressed block from the multiple sub-compressed blocks, determine multiple second sub-compressed blocks other than the first sub-compressed block, and determine at least one second sub-compressed block pair from the multiple second sub-compressed blocks, where the second sub-compressed block pair includes two second sub-compressed blocks; Generate one such shared compression header for the first sub-compressed block, and generate one such shared compression header for each of the second sub-compressed block pairs.
11. The method according to claim 9, wherein The generating the combined compressed block based on the multiple sub-compressed blocks and the at least one shared compression header includes: For each of the first sub-compressed block pairs, place the shared compression header corresponding to the first sub-compressed block pair between the two sub-compressed blocks of the first sub-compressed block pair.
12. A data processing method, characterized in that, The method includes: Obtain data to be compressed, divide the data to be compressed into multiple data blocks, perform compression processing on each of the data blocks respectively to obtain multiple sub-compressed blocks; Generate at least one shared compression header according to the multiple sub-compressed blocks; Generate a combined compressed block based on the multiple sub-compressed blocks and the at least one shared compression header.
13. A data transmission method, characterized in that, The method includes: Receive a combined compressed block sent by a sending end; Obtain a data access request, where the data access request carries a data identifier of target data; According to the data identifier, determine the combined compressed block to which the target data belongs, where the combined compressed block includes multiple sub-compressed blocks and at least one shared compression header; According to the data identifier and the shared compression header, determine a target sub-compressed block from the multiple sub-compressed blocks, and perform decompression processing on the target sub-compressed block to obtain the target data.
14. A data transmission method, characterized in that, The method includes: Obtain data to be compressed, divide the data to be compressed into multiple data blocks, perform compression processing on each of the data blocks respectively to obtain multiple sub-compressed blocks; Generate at least one shared compression header according to the multiple sub-compressed blocks; Generate a combined compressed block based on the multiple sub-compressed blocks and the at least one shared compression header; Send the combined compressed block to a receiving end.
15. A data processing device, characterized in that, The apparatus includes: A first obtaining module, configured to obtain a data access request, where the data access request carries a data identifier of target data; A first determining module, configured to determine, according to the data identifier, the combined compressed block to which the target data belongs, where the combined compressed block includes multiple sub-compressed blocks and at least one shared compression header; A second determining module, configured to determine a target sub-compressed block from the multiple sub-compressed blocks according to the data identifier and the shared compression header, and perform decompression processing on the target sub-compressed block to obtain the target data.
16. A data processing device, characterized in that, The apparatus includes: A second acquisition module, configured to acquire data to be compressed, divide the data to be compressed into a plurality of data blocks, and perform compression processing on each of the data blocks respectively to obtain a plurality of sub-compressed blocks; A first generation module, configured to generate at least one shared compression header according to the plurality of sub-compressed blocks; A second generation module, configured to generate the combined compressed block based on the plurality of sub-compressed blocks and the at least one shared compression header.
17. A data transmission device, characterized in that, The apparatus includes: A receiving module, configured to receive the combined compressed block sent by a sending end; A third acquisition module, configured to acquire a data access request, where the data access request carries a data identifier of target data; A third determination module, configured to determine, according to the data identifier, the combined compressed block to which the target data belongs, where the combined compressed block includes a plurality of sub-compressed blocks and at least one shared compression header; A fourth determination module, configured to determine a target sub-compressed block from the plurality of sub-compressed blocks according to the data identifier and the shared compression header, and perform decompression processing on the target sub-compressed block to obtain the target data.
18. A data transmission device, characterized in that, The apparatus includes: A fourth acquisition module, configured to acquire data to be compressed, divide the data to be compressed into a plurality of data blocks, and perform compression processing on each of the data blocks respectively to obtain a plurality of sub-compressed blocks; A third generation module, configured to generate at least one shared compression header according to the plurality of sub-compressed blocks; A fourth generation module, configured to generate the combined compressed block based on the plurality of sub-compressed blocks and the at least one shared compression header; A sending module, configured to send the combined compressed block to a receiving end.
19. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 14 are implemented.
20. A data transmission system, including a receiving end and a sending end, where the receiving end is configured to execute the steps of the method according to claim 13; The sending end is configured to execute the steps of the method according to claim 14.
21. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 14 are implemented.
22. A computer program product comprising a computer program, characterized in that, When the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 14 are implemented.