Data processing method and apparatus, and device and computer-readable storage medium

By determining clear decompression boundaries in cloud computing and accurately decompressing data, the problem of waste of resources and performance instability caused by the increase in data storage and computing requirements in cloud computing is solved, and efficient and stable data decompression and calculation are achieved.

WO2025108094A1PCT designated stage expired Publication Date: 2025-05-30HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/130491
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-19
Filing Date
2024-11-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In cloud computing, the demand for data storage and computing has increased, and it is difficult for the existing technology to achieve efficient data decompression and computing, resulting in waste of resources and unstable performance.

Method used

By obtaining data processing instructions, a clear decompression boundary is determined to achieve accurate decompression. The specific method includes obtaining the first compressed data, determining the starting position of the second compressed data based on the starting position and the reference distance, and decompressing the second compressed data according to the determined boundary to obtain the target data.

Benefits of technology

Improves compression recognition efficiency and accuracy, reduces resource waste, and stabilizes compression understanding performance, and is suitable for various data processing scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024130491_30052025_PF_FP_ABST
    Figure CN2024130491_30052025_PF_FP_ABST
Patent Text Reader

Abstract

The present application belongs to the technical field of cloud computing. Disclosed are a data processing method and apparatus, and a device and a computer-readable storage medium. The method in the present application comprises: acquiring a data processing instruction, which instructs decompression from first compressed data to obtain target data, wherein the first compressed data is obtained by compressing a data group to which the target data belongs, and the data processing instruction comprises a first starting position of the target data in the data group and the length of the target data; on the basis of the first starting position and a reference distance, determining in the first compressed data a second starting position of second compressed data used for decompression to obtain the target data, wherein the reference distance indicates the maximum value of the distance between any data segment in the target data and a data segment referred to thereby; and according to the second starting position and an ending position of the second compressed data, decompressing the second compressed data to obtain the target data, wherein the ending position is determined on the basis of the first starting position and the length of the target data. The present application can provide a clear decompression boundary before decompression, and improve the decompression accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Data processing method, device, equipment and computer-readable storage medium

[0001] This application claims priority to Chinese patent application number 202311559827.1 filed on November 20, 2023, with invention name “Data processing method, device, equipment and computer-readable storage medium”. This application claims priority to Chinese patent application number 202410315918.9 filed on March 19, 2024, with invention name “Data processing method, device, equipment and computer-readable storage medium”, the entire contents of which are incorporated into this application by reference. Technical Field

[0002] The present application relates to the field of cloud computing technology, and in particular to data processing methods, devices, equipment, and computer-readable storage media. Background Art

[0003] With the development of cloud computing technology, computing power has gradually increased. The amount of data required for computation has increased, and the amount of data that needs to be stored has also increased. Therefore, during data storage, data compression can be used to store compressed data, reducing storage costs. When computing on stored data, the compressed data must first be decompressed to obtain the pre-compressed data. Computation can then be performed on the pre-compressed data to meet the computational needs.

[0004] Summary of the Invention

[0005] This application provides a data processing method, apparatus, device, and computer-readable storage medium that can provide clear decompression boundaries and achieve precise decompression. The technical solution is as follows:

[0006] In a first aspect, a data processing method is provided, the method comprising: obtaining a data processing instruction, the data processing instruction instructing to decompress target data from first compressed data, the first compressed data being obtained by compressing a data group to which the target data belongs, the data processing instruction comprising a first starting position of the target data in the data group and a length of the target data; determining a second starting position of second compressed data in the first compressed data based on the first starting position and a reference distance, the second compressed data being used to decompress and obtain the target data, the reference distance indicating a maximum value of a distance between any data segment in the target data and a data segment referenced by any data segment; decompressing the second compressed data according to the second starting position and an end position of the second compressed data to obtain the target data, the end position of the second compressed data being determined based on the first starting position and the length of the target data.

[0007] Since the reference distance indicates the maximum value of the distance between any data segment in the target data and the data segment referenced by any data segment, the data segment referenced by each data segment in the target data may be located within the target data, or within a portion of the data whose distance from the first starting position is less than or equal to the reference distance. Based on the first starting position and the reference distance, the second starting position of the second compressed data used for decompression to obtain the target data can be determined, and it is ensured that the second compressed data can be decompressed from the second starting position to obtain the target data. In addition, based on the first starting position and the length of the target data, the end position of the second compressed data can be accurately determined, and decompression can be stopped at the end position of the second compressed data. Decompression is performed using the determined second starting position and the end position of the second compressed data as the decompression boundary, which not only ensures that the decompressed data includes the target data, but also reduces the decompression of compressed data that does not need to be decompressed, improves decompression efficiency and accuracy, and reduces resource waste.

[0008] In one possible implementation, the first compressed data includes multiple compressed segments, the data group includes multiple data segments, and each compressed segment is compressed by the corresponding data segment; based on the first starting position and the reference distance, the second starting position of the second compressed data in the first compressed data is determined, including: obtaining the starting position of each data segment in the data group; according to the starting position of each data segment, determining the first data segment to which the first position belongs in the multiple data segments, the first position is determined based on the difference between the first starting position and the reference distance; according to the starting position of the first data segment, determining the starting position of the first compressed segment corresponding to the first data segment; and determining the starting position of the first compressed segment as the second starting position.

[0009] Due to the difference between the first starting position and the reference distance, it can indicate the first position that is before the first starting position and whose distance from the first starting position is the reference distance, and since the distance between any data segment in the target data and the data segment referenced by any data segment is less than or equal to the reference distance, that is, the data segment referenced by any data segment in the target data segment is located at the first position or after the first position, and therefore decompression starts from the starting position of the first compressed segment corresponding to the first data segment belonging to the first position, it can be ensured that the target data is obtained by decompression.

[0010] In one possible implementation, obtaining the starting position of each data segment in a data group includes: obtaining the length of each data segment; determining the starting position of each data segment in the data group based on the length of each data segment and the order of each compressed segment in the first compressed data, wherein the order of any data segment in the data group is the same as the order of the compressed segments obtained by compressing any data segment in the first compressed data.

[0011] Since the order of any data segment in the data group is the same as the order of the compressed segments obtained by compressing any data segment in the first compressed data, the order of each data segment in the data group can be determined according to the order of each compressed segment in the first compressed data, and then according to the length of each data segment and the order of each data segment in the data group, the initial position of each data segment in the data group can be accurately determined.

[0012] In one possible implementation, determining the starting position of a first compressed segment corresponding to the first data segment based on the starting position of the first data segment includes: obtaining mapping information, the mapping information being used to indicate a mapping relationship between the starting positions of multiple data segments and the starting positions of multiple compressed segments; and determining the starting position of the first compressed segment based on the mapping information and the starting position of the first data segment. Because the mapping information can indicate the mapping relationship between the starting positions of multiple data segments and the starting positions of multiple compressed segments, the first data segment belongs to multiple data segments, and thus the starting position of the first compressed segment corresponding to the first data segment can be accurately determined based on the mapping information.

[0013] In one possible implementation, mapping information includes the starting position of at least one reference data segment among multiple data segments and the starting position of at least one reference compressed segment compressed from the at least one reference data segment; determining the starting position of the first compressed segment based on the mapping information and the starting position of the first data segment includes: determining a first reference data segment in the at least one reference data segment based on the starting position of the first data segment and the starting position of the at least one reference data segment, the first reference data segment being the reference data segment that precedes the first data segment and is closest to the first data segment; and determining the starting position of the first reference compressed segment compressed from the first reference data segment as the starting position of the first compressed segment. Storing the starting position of at least one reference data segment and the starting position of at least one reference compressed segment in the mapping information can reduce the memory space required for the mapping information and reduce resource waste.

[0014] In one possible implementation, before decompressing the second compressed data according to the second starting position and the ending position of the second compressed data to obtain the target data, it also includes: determining the second data segment to which the second position belongs in multiple data segments according to the starting positions of each data segment, the second position being the sum of the first starting position and the length of the target data; determining the starting position of a third data segment adjacent to the second data segment in the starting positions of each data segment, the third data segment being after the second data segment; determining the starting position of the second compressed segment corresponding to the third data segment according to the starting position of the third data segment; and determining the starting position of the second compressed segment as the ending position of the second compressed data.

[0015] Since the sum of the first starting position and the length of the target data can indicate the end position of the target data in the data group, that is, the second position can indicate the end position of the target data in the data group, the second data segment to which the second position belongs is the data segment to which the end position of the target data belongs. Therefore, the starting position of the third data segment adjacent to the second data segment and after the second data segment is located after the end position of the target data, and decompression is stopped from the starting position of the second compressed segment corresponding to the third data segment, which can ensure that the target data is completely decompressed.

[0016] In one possible implementation, decompressing the second compressed data according to the second starting position and the ending position of the second compressed data to obtain target data includes: decompressing the second compressed data according to the second starting position and the ending position of the second compressed data to obtain intermediate data; determining the ending position of the target data according to the first starting position and the length of the target data; and intercepting the target data from the intermediate data according to the first starting position and the ending position of the target data. Intercepting the target data from the intermediate data according to the first starting position and the ending position of the target data can ensure that the intercepted data is completely consistent with the target data, thereby increasing the accuracy of the decompressed target data.

[0017] In one possible implementation, a data group is stored in a heap-organized table (HOT), the data group includes multiple data within a storage block of the HOT, and the target data includes at least one data among the multiple data. The target data is at least one data among the multiple data, that is, the target data is part of the data in the data group. Therefore, according to the data processing method provided in the present application, it is possible to partially decompress the compressed data obtained by compressing the multiple data to obtain at least one data, thereby ensuring the accuracy of the decompression.

[0018] In one possible implementation, a data group is stored in a baseline database of a log structure merge tree (LSM-Tree). The data group includes data within multiple storage blocks in the baseline database, and the target data includes data within at least one of the multiple storage blocks. The target data is data within at least one of the multiple storage blocks, that is, the target data is a portion of the data in the data group. Therefore, according to the data processing method provided in this application, compressed data obtained by compressing the data within the multiple storage blocks can be partially decompressed to obtain data within the at least one storage block, thereby ensuring the accuracy of the decompression.

[0019] In a second aspect, a data processing method is provided, the method comprising: obtaining a data group, the data group comprising multiple bytes; dividing the multiple bytes into multiple data segments, the multiple data segments comprising at least one of a matching data segment or a non-matching data segment, the bytes in the non-matching data segment being different from the bytes in each data segment preceding the non-matching data segment, the bytes in the matching data segment being the same as the bytes in a matched data segment preceding the matching data segment, and the distance between the matching data segment and the matched data segment being less than or equal to a reference distance; compressing the multiple data segments according to the data segments referenced by the multiple data segments to obtain first compressed data comprising multiple compressed segments, each compressed segment being compressed by the corresponding data segment, the data segment referenced by any matching data segment being the matched data segment having the same bytes as those in any matching data segment, and the data segment referenced by any non-matching data segment being any non-matching data segment.

[0020] In the present application, by controlling the distance between the matching data segment and the matched data segment in the data group to be less than or equal to the reference distance, during decompression, the decompression of each compressed segment that needs to be decompressed can be achieved based on the compressed segment corresponding to the data segment whose distance between each data segment does not exceed the reference distance, thereby enabling the clear boundary of the compressed data that needs to be decompressed to be determined according to the reference distance before decompression, thereby achieving precise decompression.

[0021] In one possible implementation, multiple bytes are divided into multiple data segments, including: dividing the multiple bytes into multiple alternative data segments, the multiple alternative data segments include at least one of alternative matching data segments or non-matching data segments, and the bytes in the alternative matching data segment are the same as the bytes in at least one alternative matched data segment before the alternative matching data segment; obtaining a position array of the alternative matching data segment, the position array including the position of the alternative matching data segment and the position of at least one alternative matched data segment corresponding to the alternative matching data segment; in a case where there is an alternative distance less than or equal to a reference distance in at least one alternative distance, determining the alternative matching data segment as a matching data segment, the at least one alternative distance being determined based on the position of the alternative matching data segment and the position of the at least one alternative matched data segment; in a case where at least one alternative distance is greater than the reference distance, determining the alternative matching data segment as a non-matching data segment.

[0022] If at least one candidate distance corresponding to a candidate matching data segment is greater than the reference distance, then compressing the candidate matching data segment by referencing the candidate matched data segment corresponding to the candidate matching data segment will not allow decompression of the candidate matching data segment based on the compressed segments corresponding to the data segments whose distances from the candidate matching data segment do not exceed the reference distance. Therefore, if at least one candidate distance corresponding to the candidate matching data segment includes a candidate distance less than or equal to the reference distance, the candidate matching data segment is determined to be a matching data segment, and then the matching data segment is compressed by referencing the matched data segments whose distances from the matching data segments do not exceed the reference distance. This ensures that the compressed segments corresponding to the data segments whose distances from the matching data segments do not exceed the reference distance are decompressed to obtain the matching data segment, thereby ensuring decompression accuracy.

[0023] In one possible implementation, after compressing the multiple data segments based on the data segments referenced by the multiple data segments, the method further includes generating mapping information, the mapping information being used to indicate a mapping relationship between the starting positions of the multiple data segments and the starting positions of the multiple compressed segments. Because the mapping information can indicate the mapping relationship between the starting positions of the multiple data segments and the starting positions of the multiple compressed segments, during the decompression process, the starting position of the compressed segment corresponding to the starting position of any data segment can be efficiently determined based on the mapping information, or the starting position of the data segment corresponding to the starting position of any compressed segment can be efficiently determined.

[0024] In one possible implementation, a data group is stored in a HOT, and the data group includes multiple data in a storage block of the HOT. Through the data processing method provided by the present application, the multiple data in the storage block are compressed, and the distance between the matching data segment and the matched data segment in the multiple data in the storage block can be controlled to be less than or equal to a reference distance. When the compressed data obtained by compressing the multiple data is decompressed, a clear decompression boundary can be determined based on the reference distance to ensure the accuracy of the decompression.

[0025] In one possible implementation, a data group is stored in a baseline database of an LSM-Tree, and the data group includes data within multiple storage blocks in the baseline database. The data processing method provided in this application compresses the data within the multiple storage blocks, and the distance between matching data segments and matched data segments in the data within the multiple storage blocks can be controlled to be less than or equal to a reference distance. This allows, when decompressing compressed data obtained by compressing the data within the multiple storage blocks, to determine a clear decompression boundary based on the reference distance, thereby ensuring decompression accuracy.

[0026] According to a third aspect, a data processing device is provided, which includes: an acquisition module for acquiring a data processing instruction, the data processing instruction instructing to decompress target data from first compressed data, the first compressed data being obtained by compressing a data group to which the target data belongs, the data processing instruction including a first starting position of the target data in the data group and a length of the target data; a determination module for determining a second starting position of the second compressed data in the first compressed data based on the first starting position and a reference distance, the second compressed data being used to decompress and obtain the target data, the reference distance indicating a maximum value of a distance between any data segment in the target data and a data segment referenced by any data segment; a decompression module for decompressing the second compressed data according to the second starting position and an end position of the second compressed data to obtain the target data, the end position of the second compressed data being determined based on the first starting position and the length of the target data.

[0027] In one possible implementation, the first compressed data includes multiple compressed segments, the data group includes multiple data segments, and each compressed segment is compressed by the corresponding data segment; a determination module is used to obtain the starting position of each data segment in the data group; based on the starting position of each data segment, the first data segment to which the first position belongs is determined in the multiple data segments, and the first position is determined based on the difference between the first starting position and the reference distance; based on the starting position of the first data segment, the starting position of the first compressed segment corresponding to the first data segment is determined; and the starting position of the first compressed segment is determined as the second starting position.

[0028] In one possible implementation, a determination module is used to obtain the length of each data segment; based on the length of each data segment and the order of each compressed segment in the first compressed data, the starting position of each data segment in the data group is determined, and the order of any data segment in the data group is the same as the order of the compressed segments obtained by compressing any data segment in the first compressed data.

[0029] In one possible implementation, a determination module is used to obtain mapping information, where the mapping information is used to indicate a mapping relationship between the starting positions of multiple data segments and the starting positions of multiple compressed segments; and determine the starting position of the first compressed segment based on the mapping information and the starting position of the first data segment.

[0030] In one possible implementation, the mapping information includes the starting position of at least one reference data segment in multiple data segments and the starting position of at least one reference compressed segment compressed by the at least one reference data segment; a determination module is used to determine a first reference data segment in the at least one reference data segment based on the starting position of the first data segment and the starting position of the at least one reference data segment, the first reference data segment being the reference data segment that is before the first data segment and is closest to the first data segment; and the starting position of the first reference compressed segment obtained by compressing the first reference data segment is determined as the starting position of the first compressed segment.

[0031] In one possible implementation, the determination module is further used to determine, in multiple data segments, the second data segment to which the second position belongs based on the starting positions of each data segment, where the second position is the sum of the first starting position and the length of the target data; determine, in the starting positions of each data segment, the starting position of a third data segment adjacent to the second data segment, where the third data segment is after the second data segment; determine, based on the starting position of the third data segment, the starting position of the second compressed segment corresponding to the third data segment; and determine the starting position of the second compressed segment as the ending position of the second compressed data.

[0032] In one possible implementation, the decompression module is used to decompress the second compressed data according to the second starting position and the ending position of the second compressed data to obtain intermediate data; determine the ending position of the target data according to the first starting position and the length of the target data; and intercept the target data from the intermediate data according to the first starting position and the ending position of the target data.

[0033] In a possible implementation, the data group is stored in the HOT, the data group includes multiple data in a storage block of the HOT, and the target data includes at least one data among the multiple data.

[0034] In a possible implementation, the data group is stored in a baseline database of the LSM-Tree, the data group includes data in multiple storage blocks in the baseline database, and the target data includes data in at least one storage block among the multiple storage blocks.

[0035] In a fourth aspect, a data processing device is provided, which includes: an acquisition module for acquiring a data group, the data group including multiple bytes; a division module for dividing the multiple bytes into multiple data segments, the multiple data segments including at least one of a matching data segment or a non-matching data segment, the bytes in the non-matching data segment are different from the bytes in each data segment before the non-matching data segment, the bytes in the matching data segment are the same as the bytes in a matched data segment before the matching data segment, and the distance between the matching data segment and the matched data segment is less than or equal to the reference distance; a compression module for compressing the multiple data segments based on the data segments referenced by the multiple data segments to obtain first compressed data including multiple compressed segments, each compressed segment is compressed by the corresponding data segment, the data segment referenced by any matching data segment is the matched data segment with the same bytes in any matching data segment, and the data segment referenced by any non-matching data segment is any non-matching data segment.

[0036] In one possible implementation, a partitioning module is used to partition multiple bytes into multiple alternative data segments, the multiple alternative data segments include at least one of alternative matching data segments or non-matching data segments, and the bytes in the alternative matching data segment are the same as the bytes in at least one alternative matched data segment before the alternative matching data segment; obtain a position array of the alternative matching data segment, the position array including the position of the alternative matching data segment and the position of at least one alternative matched data segment corresponding to the alternative matching data segment; determine the alternative matching data segment as a matching data segment when there is an alternative distance less than or equal to a reference distance in at least one alternative distance, and the at least one alternative distance is determined based on the position of the alternative matching data segment and the position of at least one alternative matched data segment; and determine the alternative matching data segment as a non-matching data segment when at least one alternative distance is greater than the reference distance.

[0037] In a possible implementation, the device further includes a generation module; the generation module is used to generate mapping information, where the mapping information is used to indicate a mapping relationship between the starting positions of the multiple data segments and the starting positions of the multiple compressed segments.

[0038] In a possible implementation, a data group is stored in a HOT, and the data group includes a plurality of data in a storage block of the HOT.

[0039] In a possible implementation, the data group is stored in a baseline database of the LSM-Tree, and the data group includes data in multiple storage blocks in the baseline database.

[0040] In a fifth aspect, the present application provides a computing device cluster, which includes at least one computing device, each computing device including a processor and a memory; the processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster performs the data processing method provided by the aforementioned first aspect, second aspect, and any possible implementation of the first aspect or the second aspect.

[0041] In a sixth aspect, embodiments of the present application provide a computer program product comprising instructions. When the instructions are executed by a computing device cluster, the computing device cluster performs the data processing method provided in the first aspect, the second aspect, and any possible implementation of the first aspect or the second aspect. The computer program product may be a software installation package. When the functions of the computing device cluster described above are required, the computer program product may be downloaded and executed on the computing device cluster.

[0042] In a seventh aspect, embodiments of the present application provide a computer-readable storage medium comprising computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster performs the data processing method provided in the first aspect, the second aspect, and any possible implementation of the first aspect or the second aspect. The storage medium includes, but is not limited to, volatile memory, such as random access memory, and non-volatile memory, such as flash memory, a hard disk drive (HDD), or a solid state drive (SSD).

[0043] It should be understood that the beneficial effects achieved by the technical solutions of the third to seventh aspects of this application and the corresponding possible implementation methods can be referred to the technical effects of the data processing methods provided for the first aspect, the second aspect, and any possible implementation method of the first aspect or the second aspect, and will not be repeated here. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] FIG1 is a schematic diagram of a general data compression process based on the LZ77 algorithm provided by the related art;

[0045] FIG2 is a schematic diagram of a general data compression process based on the LZ77 algorithm provided by the related art;

[0046] FIG3 is a schematic diagram of an implementation scenario of a data processing method provided in an embodiment of the present application;

[0047] FIG4 is a flow chart of a data compression method provided in an embodiment of the present application;

[0048] FIG5 is a schematic diagram of a HOT-based data processing environment provided in an embodiment of the present application;

[0049] FIG6 is a schematic diagram of an LSM-Tree-based data processing environment provided in an embodiment of the present application;

[0050] FIG7 is a schematic diagram of mapping information provided in an embodiment of the present application;

[0051] FIG8 is a schematic diagram of a data compression process provided by an embodiment of the present application;

[0052] FIG9 is a schematic diagram of a data decompression process provided by an embodiment of the present application;

[0053] FIG10 is a schematic structural diagram of a data processing device provided in an embodiment of the present application;

[0054] FIG11 is a schematic structural diagram of another data processing device provided in an embodiment of the present application;

[0055] FIG12 is a schematic diagram of the hardware structure of a computing device provided in an embodiment of the present application;

[0056] FIG13 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application;

[0057] FIG14 is a schematic diagram of a connection method for a computing device cluster provided in an embodiment of the present application. DETAILED DESCRIPTION

[0058] The terms used in the implementation section of this application are only used to explain the specific embodiments of this application and are not intended to limit this application.

[0059] With the development of cloud computing technology, cloud services based on this technology are becoming increasingly diverse. The complex calculations required to provide these services require a large amount of data to be accessed and, consequently, a large amount of data to be stored. For example, data management systems such as databases are used to store large amounts of data corresponding to various cloud services. As cloud services are updated and developed, the amount of data these management systems need to store increases. However, to store this ever-increasing amount of data, data management systems require continuous system expansion, consuming significant storage costs. Furthermore, this expansion can lead to instability in the data management system, resulting in problems such as errors or loss of stored data.

[0060] Therefore, when storing data, data compression techniques are often used to compress the data. This reduces the size of the compressed data, and storing compressed data can reduce data storage costs and reduce the frequency and potential risks of system expansion. For example, a group of data can be compressed as a whole using data compression techniques. For example, a group of user records in a database can be compressed as a whole. Examples of data compression techniques include LZ77 (a data compression technique) and Huffman compression. LZ77 is a lossless data compression algorithm. Huffman compression achieves data compression by encoding the data.

[0061] During the computational process, if compressed data needs to be calculated, the required compressed data must first be queried, then decompressed to obtain the decompressed data, and the computation is then completed based on the decompressed data. Since the computational process may involve partial queries on compressed data—that is, the data to be queried is part of a compressed set of data—there is no need to decompress all of the compressed data. Instead, partial decompression of the compressed data can be performed to obtain the required data.

[0062] Because different data compression technologies are used to compress data, the resulting compressed data is different, and therefore the methods for partially decompressing different compressed data are different. For example, related art 1 uses data compression technologies such as dictionary encoding or run length encoding (RLE) to compress data based on the characteristics of the compressed data, so that the compressed data can support partial decompression.

[0063] Taking dictionary encoding as an example of related technology 1, related technology 1 can divide the data to be compressed into multiple independent fields according to business needs and business characteristics, find repeated fields that appear frequently in multiple fields, store the values ​​of the repeated fields in an independent data dictionary, and replace the values ​​of the repeated fields in the data to be compressed with the positions of the repeated fields in the data dictionary to obtain the compressed data, thereby completing the data compression. When decompressing the compressed data, the data can be restored according to the positions of the repeated fields in the compressed data in the data dictionary and the values ​​in the data dictionary to complete the data decompression. When part of the data needs to be decompressed, the position in the data dictionary included in the part of the data needs to be restored to the value of the corresponding repeated field to obtain the restored data, thereby achieving the decompression of part of the data.

[0064] However, Related Art 1 can only identify repeated fields in data, but cannot identify redundant, repeated information within fields. Therefore, it cannot compress unidentified repeated information, making it difficult to guarantee the compression rate for repeated information in the data, and thus, to ensure the stability of the compression rate. Furthermore, Related Art 1 requires data to be segmented according to business needs and characteristics, but in some cases, this is difficult to do, making Related Art 1 difficult to apply to compression and decompression scenarios in such situations. Consequently, this technology has poor adaptability to specific scenarios.

[0065] Another related technology in the field of data compression is to divide the data to be compressed into multiple independent parts, compress each part separately, and obtain multiple compressed segments. When it is necessary to decompress part of the data, the compressed segment containing the part of the data to be decompressed can be decompressed.

[0066] However, since repeated information in the data may be located in different parts, data compression performed through the second related technology makes it impossible to identify and eliminate repeated information located in different parts, thereby reducing the overall compression rate of the data. For example, when the data to be compressed is small in size, for example, the data to be compressed is the data in a page of a database, and the length of the data in a page is generally 8 kilobytes (KB) to 16KB. The repetition rate of the smaller data itself may be low, and dividing the data into multiple independent parts may result in an even lower repetition rate of information in each independent part, or even no repetition of information in each independent part, resulting in a lower compression rate.

[0067] In addition, Related Technology 2 compresses multiple parts of data through independent compression processes, requiring multiple calls to the compression algorithm. That is, each time a separate part is compressed, the compression algorithm needs to be called. Calling the compression algorithm may require introducing an index structure for compression, and maintaining independent index structures for different parts to identify duplicate information in different parts. Multiple calls and activations of the compression algorithm trigger the initialization of the index structure, making the startup cost of calling the compression algorithm high and resulting in a waste of resources.

[0068] This application provides a data processing method that can compress and decompress data, and supports partial data decompression. This method has a stable compression rate, strong adaptability to different scenarios, and a low startup cost for calling the compression algorithm, which can avoid resource waste. This data processing method can be implemented based on a common compression algorithm, such as the LZ77 algorithm.

[0069] In order to make the description of the data processing method provided in the embodiment of the present application easier to understand, before describing the data processing method provided in the embodiment of the present application, the general data compression and decompression process based on the LZ77 algorithm is first schematically described.

[0070] For example, referring to FIG1 , a schematic diagram of a general data compression process based on the LZ77 algorithm provided by the related art is shown. First, the data to be compressed is obtained, treated as a byte stream comprising multiple bytes, and the bytes are processed sequentially. Before processing each byte, a determination is made as to whether all bytes in the byte stream have been processed. If there are still bytes that have not been processed, processing proceeds to the next byte.

[0071] Taking byte x as an example, the first step is to find the matching data segment for byte x. For example, the bytes preceding byte x can be searched for the byte identical to byte x. If the bytes preceding byte x contain the same byte as byte x, byte x is temporarily recorded as a matching data segment, and the bytes preceding byte x that are identical to byte x are temporarily recorded as matching data segments. If the bytes preceding byte x do not contain the same byte as byte x, byte x is temporarily recorded as a non-matching data segment. The matching data segment can also be called a matching string, the matching data segment can also be called a matching string, and the non-matching data segment can also be called a non-matching string.

[0072] For example, see Figure 2, which shows a schematic diagram of a general data compression process based on the LZ77 algorithm provided by the related art. The data to be compressed is obtained as ABCDEFBCDGDEFEAFBC, where each letter is a byte. Therefore, the data to be compressed can be considered as a byte stream containing 18 bytes. Since the data to be compressed is in text form, the data to be compressed can also be referred to as the text to be compressed or the original text.

[0073] After compression begins, each byte in the byte stream is processed sequentially. When processing the first byte A, the matching data segment for the first byte A is first searched. Since the first byte A is the first byte in the byte stream, there is no byte identical to the first byte A before it. Therefore, the matching data segment corresponding to the first byte A cannot be found in the byte stream. Therefore, the first byte A can be temporarily recorded as a non-matching data segment. The second byte B, the third byte C, the fourth byte D, the fifth byte E, and the sixth byte F cannot find corresponding matching data segments in the byte stream. Therefore, the second byte B, the third byte C, the fourth byte D, the fifth byte E, and the sixth byte F can all be temporarily recorded as non-matching data segments. When processing the seventh byte B, it is found that the second byte B is identical to the seventh byte B and that the second byte B precedes the seventh byte B. Therefore, the seventh byte B can be temporarily recorded as a matching data segment, and the second byte B can be temporarily recorded as the matching data segment for the seventh byte B. When processing the eighth byte C and the ninth byte D, it is found that the third byte C is the same as the eighth byte C, and the fourth byte D is the same as the ninth byte D. Therefore, the eighth byte C and the ninth byte D can be temporarily recorded as matching data segments, the third byte C can be temporarily recorded as the matched data segment of the eighth byte C, and the fourth byte D can be temporarily recorded as the matched data segment of the ninth byte D.

[0074] In the process of sequentially processing each byte in the byte stream, if several consecutive bytes are identical to several previous consecutive bytes, the consecutive bytes can be merged into a matching data segment, and the previous consecutive bytes can be determined as the matched data segment of the matching data segment. One or more consecutive bytes between two adjacent matching data segments that are temporarily recorded as non-matching data segments can be merged into a non-matching data segment.

[0075] For example, referring to Figure 2, when processing the tenth byte G, no byte identical to G can be found in the byte stream. Therefore, the tenth byte G can be temporarily recorded as a non-matching data segment. The consecutive seventh byte B, eighth byte C, and ninth byte D are identical to the consecutive second byte B, third byte C, and fourth byte D, while the tenth byte G is different from the fifth byte F. Therefore, the consecutive seventh byte B, eighth byte C, and ninth byte D can be combined into a matching data segment BCD, and the consecutive second byte B, third byte C, and fourth byte D can be determined as the matched data segment BCD of this matching data segment.

[0076] Furthermore, since the first byte A through the sixth byte F are temporarily recorded as a non-matching data segment, and the seventh byte B following the sixth byte F has been determined to be a matching data segment, the first byte A through the sixth byte F can be combined into a non-matching data segment ABCDEF. The processing of the other bytes in the byte stream in Figure 2 can refer to the above description and will not be repeated here.

[0077] In some cases, after recording and merging the matching and non-matching data segments, the final matching and non-matching data segments can be determined based on the lengths of the merged matching data segments. If the length of the merged matching data segment is less than a length threshold, the matching data segment can be re-determined as a non-matching data segment.

[0078] For example, the fourteenth byte E in Figure 2 is identical to the twelfth byte E and the fifth byte E. The fourteenth byte E is recorded as a matching data segment E. However, the length of this matching data segment E is 1 bit (B), which is less than the length threshold of 3B. Therefore, this matching data segment E can be re-determined as a non-matching data segment E. The fifteenth byte A is identical to the first byte A. The fifteenth byte A is recorded as a matching data segment A. However, the length of this matching data segment A is 1B, which is less than the length threshold of 3B. Therefore, this matching data segment A can be re-determined as a non-matching data segment A. Because the byte F preceding the fourteenth byte E belongs to the matching data segment DEF, whose length is equal to the length threshold, and the byte F following the fifteenth byte A also belongs to the matching data segment FBC, whose length is equal to the length threshold, the fourteenth byte E and the fifteenth byte A can be merged into a non-matching data segment EA.

[0079] After determining the final matching and non-matching data segments, data compression can be performed based on the characteristics of each matching and non-matching data segment, and the compressed data can be output. For non-matching data segments, the length of the non-matching data segment and the non-matching data segment are output; for matching data segments, the length of the matching data segment and the matching offset of the matching data segment are output. The matching offset of the matching data segment refers to the distance between the matching data segment and the matched data segment of the matching data segment, for example, the byte length between the first byte of the matching data segment and the last byte of the matched data segment.

[0080] Continuing with the byte stream in Figure 2 as an example, the first byte A to the sixth byte F is ultimately determined to be a non-matching data segment ABCDEF with a length of 6B, and the length 6B and the non-matching data segment ABCDEF can be output; the seventh byte B to the ninth byte D is ultimately determined to be a matching data segment BCD with a length of 3B, and the matched data segment BCD of the matching data segment BCD is the second byte B to the fourth byte D, and the distance between the matching data segment BCD and the matched data segment BCD is 2B, that is, the matching displacement of the matching data segment BCD is 2B, and the length 3B and the matching displacement 2B can be output (the unit B of the length and displacement can be ignored); the tenth byte G is ultimately determined to be a non-matching data segment G with a length of 1B, and the length 1B and the non-matching data segment G can be output; the eleventh byte D to the thirteenth byte F are ultimately determined to be a matching data segment 3B in length. The matched data segment DEF of the matching data segment DEF is the fourth byte D to the sixth byte F, and the distance between the matching data segment DEF and the matched data segment DEF is 4B, that is, the matching displacement of the matching data segment DEF is 4B, so the length 3B and the matching displacement 4B can be output; the fourteenth byte E to the fifteenth byte A are finally determined to be a non-matching data segment EA with a length of 2B, so the length 2B and the non-matching data segment EA can be output; the sixteenth byte F to the eighteenth byte C are finally determined to be a matching data segment FBC with a length of 3B, and the matched data segment FBC of the matching data segment FBC is the sixth byte F to the eighth byte C, and the distance between the matching data segment FBC and the matched data segment FBC is 7B, that is, the matching displacement of the matching data segment is 7B, so the length 3B and the matching displacement 7B can be output.

[0081] Since the compression of matching data segments depends on the position of the matched data segments and the distance between them, in some cases, it can be considered that the matching data segments are compressed by referencing the matched data segments. Similarly, the compression of non-matching data segments depends on the value and length of the non-matching data segments, so it can also be considered that the non-matching data segments are compressed by referencing the non-matching data segments themselves.

[0082] After determining the output for each matching data segment and non-matching data segment, the output results can be sorted to obtain compressed data. The compressed data is also in text form, so the compressed data can also be called compressed text. Referring to Figure 2 , the compressed data obtained by the above compression method is: (6) ABCDEF (2, 3, 1) G (4, 3, 2) EA (7, 3). Among them, 6 is the length 6B of the non-matching data segment ABCDEF; 2, 3, and 1 respectively represent the matching displacement 2B of the matching data segment BCD, the length 3B of the matching data segment BCD, and the length 1B of the next non-matching data segment G adjacent to the matching data segment BCD; 4, 3, and 2 respectively represent the matching displacement 4B of the matching data segment DEF, the length 3B of the matching data segment DEF, and the length 2B of the next non-matching data segment EA adjacent to the matching data segment DEF; 7 and 3 respectively represent the matching displacement 7B of the matching data segment FBC and the length 3B of the matching data segment FBC. Since there is no non-matching data segment after the matching data segment FBC, the length of the non-matching data segment after the matching data segment FBC can be omitted.

[0083] After data compression is completed using the above compression method, the data system will store the compressed data instead of the pre-compression data, reducing the space required for data storage and saving storage costs. Later, if the stored data needs to be retrieved, it can be decompressed to obtain the decompressed data for data calculations. If the data to be retrieved is part of the compressed data, it is necessary to partially decompress the part to restore the pre-compression data.

[0084] For example, referring to FIG2 , the data system stores compressed data (6)ABCDEF(2, 3, 1)G(4, 3, 2)EA(7, 3). If the partial data to be called is DEFEAFBC in the original data, the caller sends the starting position of the partial data in the complete pre-compressed data and the length of the partial data to the data management system. Optionally, the starting position of the partial data in the complete pre-compressed data can be the position of the first byte in the partial data in the complete pre-compressed data. The position of each byte in the complete pre-compressed data can be the number of bytes between each byte and the first byte in the complete pre-compressed data (e.g., the number marked above the pre-compressed data in FIG2 ). For example, the first byte D in the partial data DEFEAFBC is the eleventh byte in the complete pre-compressed data, so the starting position of the partial data DEFEAFBC in the complete pre-compressed data is 10. The length of the partial data can be the length of the bytes of the partial data. For example, if the number of bytes of the partial data DEFEAFBC is 8 and the length of each byte is 1B, the length of the partial data DEFEAFBC is 8B.

[0085] After receiving the starting position of the partial data in the complete uncompressed data and the length of the partial data, the data management system determines the partial compressed data corresponding to the partial data in the compressed data and decompresses the partial compressed data to obtain the partial data before compression. The data management system can determine the partial compressed data corresponding to the partial data based on the length and sequence of each data segment contained in each compressed segment in the compressed data and the starting position and length of the partial data.

[0086] For example, based on the above example, the first compressed segment (6) ABCDEF in the compressed data includes the length 6B of the first data segment ABCDEF, the second compressed segment (2, 3, 1) includes the length 3B of the second data segment BCD and the length 1B of the third data segment G, the fourth compressed segment (4, 3, 2) includes the length 3B of the fourth data segment DEF and the length 2B of the fifth data segment EA, and the sixth compressed segment (7, 3) includes the length 3B of the sixth data segment FBC.

[0087] If the partial data DEFEAFBC to be called starts at position 10 in the complete uncompressed data, then it is necessary to find the compressed segment (4, 3, 2) corresponding to the data segment DEF to which the eleventh byte D belongs. The length of the first data segment ABCDEF is 6 bytes, so the eleventh byte D does not belong to the first data segment ABCDEF; the cumulative length of the first data segment ABCDEF and the second data segment BCD is 9 bytes, so the eleventh byte D does not belong to the second data segment BCD; the cumulative length from the first data segment ABCDEF to the third data segment G is 10 bytes, so the eleventh byte D does not belong to the third data segment G. However, it can be determined that the eleventh byte D belongs to the data segment after the third data segment G, that is, the fourth data segment DEF. Therefore, it can be determined that decompression needs to start from the fourth compressed segment (4, 3, 2) corresponding to the fourth data segment DEF.

[0088] After determining the compressed segment to which the starting position of the portion of compressed data to be decompressed belongs, the compressed segment to which the ending position of the portion of compressed data to be decompressed belongs can be determined based on the starting position and length of the portion of data. For example, if the starting position of the portion of data DEFEAFBC to be called is 10 in the entire original data and the length of the portion of data to be called is 8 bytes, then the ending position of the portion of data DEFEAFBC to be called in the complete pre-compressed data can be determined to be 17. The cumulative length of the first data segment ABCDEF to the sixth data segment FBC in the complete pre-compressed data is 18 bytes, so it can be determined that the compressed segment to which the ending position of the portion of data DEFEAFBC to be called belongs is the sixth compressed segment (7, 3) corresponding to the sixth data segment FBC, and the compressed segments to be decompressed are the fourth compressed segments (4, 3, 2) to the sixth compressed segment (7, 3).

[0089] After determining the compressed segments that need to be decompressed, the compressed segments that need to be decompressed can be decompressed in sequence. A compressed segment compressed from a non-matching data segment can be decompressed to obtain the value of the data segment before compression based on the length of the compressed segment and the value of the compressed segment. A compressed segment compressed from a matching data segment can be decompressed to obtain the value of the data segment before compression based on the length of the matching data segment in the compressed segment, the distance between the matching data segment and the referenced matched data segment, and the value of the matched data segment.

[0090] Continuing with the above example, the compressed segments determined to need decompression are the fourth compressed segment (4, 3, 2) to the sixth compressed segment (7, 3). The fourth compressed segment (4, 3, 2) can be decompressed first. According to the format of the value of the fourth compressed segment (4, 3, 2), the fourth data segment obtained by compressing the fourth compressed segment is a matching data segment. Therefore, based on the value in the fourth compressed segment (4, 3, 2), it can be determined that the distance between the matched data segment referenced when compressing the fourth compressed segment (4, 3, 2) and the fourth data segment DEF is 4B. Furthermore, the compressed segment corresponding to the data segment 4B away from the fourth data segment DEF can be determined in the compressed data preceding the fourth compressed segment (4, 3, 2).

[0091] By reading the value of the second compressed segment (2, 3, 1), we know that the length of the third data segment G corresponding to the third compressed segment G is 1B, the length of the second data segment BCD corresponding to the second compressed segment (2, 3, 1) is 3B, and the sum of the lengths of the second data segment BCD and the third data segment G is 4B. This means that the matched data segment referenced by the fourth data segment DEF is the data segment before the second data segment BCD, that is, the first data segment ABCDEF, and the fourth data segment BCD references the last byte F of the first data segment ABCDEF, a total of three bytes forward. Since the first data segment ABCDEF obtained by compression is a non-matching data segment, the last byte F, the second to last byte E, and the third to last byte D in the first compressed segment (6)ABCDEF can be directly read to output the data segment DEF, completing the decompression and restoration of the fourth compressed segment DEF.

[0092] Since the data segment obtained by compression of the fifth compressed segment EA is a non-matching data segment, the value EA of the fifth compressed segment can be directly read and the data segment EA can be output to complete the decompression and restoration of the fifth data segment EA. According to the format of the value of the sixth compressed segment (7, 3), the sixth data segment FBC obtained by compression of the sixth compressed segment (7, 3) is a matching data segment. Therefore, based on the value in the sixth compressed segment (7, 3), it can be determined that the distance between the matched data segment referenced when compressing the sixth compressed segment (7, 3) and the sixth data segment FBC is 7B. Therefore, a data segment with a distance of 7B from the sixth data segment FBC can be determined in the compressed data before the sixth compressed segment (7, 3).

[0093] By reading the values ​​of the second compressed segment (2, 3, 1) and the fourth compressed segment (4, 3, 2), we know that the length of the fourth data segment DEF corresponding to the fourth compressed segment (4, 3, 2) is 3 bytes, the length of the fifth data segment EA corresponding to the fifth compressed segment EA is 2 bytes, and the length of the third data segment G corresponding to the third compressed segment G is 1 byte. The sum of the lengths of the fifth data segment EA through the third data segment G is 6 bytes. This indicates that the matched data segment referenced by the sixth data segment FBC is the data segment before the third data segment G, that is, the second data segment BCD. The sixth data segment FBC references the data segment starting from the second-to-last byte C of the second data segment BCD, referencing three bytes forward. Since the second data segment BCD of the second compressed segment (2, 3, 1) obtained by compression is the matching data segment, it is necessary to first decompress the second compressed segment (2, 3, 1) and restore it to the second data segment BCD. Then, the sixth compressed segment (7, 3) needs to be decompressed based on the second data segment BCD.

[0094] That is, the matched data segment directly referenced during the compression of the sixth data segment FBC includes bytes in the second data segment BCD, and the first data segment ABCDEF is directly referenced during the compression of the second data segment BCD, resulting in the sixth data segment FBC indirectly referencing the first data segment ABCDEF. Therefore, two layers of decompression are required during the decompression of the sixth compressed segment (7, 3).

[0095] From the above examples, it can be seen that when partially decompressing compressed data obtained by a general data compression method based on LZ77, there may be situations where two or even multiple layers of decompression are required, and the need for double or multiple layers of decompression cannot be predicted before decompression. That is, before decompression is completed, the number, size, and location of the compressed segments or data segments actually involved in the decompression cannot be determined, and the decompression efficiency is positively correlated with the number and size of the compressed segments that need to be decompressed. Therefore, it is difficult to guarantee the efficiency of decompression when the number and size of the compressed segments or data segments actually involved in the decompression are uncertain. For scenarios where frequent partial decompression is required, there is no stable decompression boundary and decompression efficiency, which leads to unstable decompression performance. Among them, decompression performance can be determined by combining decompression efficiency and decompression accuracy. If there are differences in decompression performance for different data, it may affect the operation of the business that needs to query data, making it difficult to guarantee the user experience.

[0096] The data processing method provided in the embodiment of the present application can ensure that the decompression boundary is determined before the data is decompressed, and the decompression performance is stabilized. For example, referring to Figure 3, an implementation scenario of the method is shown, and the implementation scenario includes a device 11 for providing cloud computing services. Optionally, the device 11 can be deployed in the cloud, for example, the device 11 can be a data management device in the cloud or other devices in the cloud that need to process data. Different services 12 and data systems 13 can be deployed in the device 11, and the device 11 can store and process data through the data system 13. The data system 13 can also be called a data management system. The data system 13 can be a system for managing data, such as a database. The data system 13 can be installed with a HOT-based data processing environment and an LSM-Tree-based data processing environment.

[0097] The data processing method provided in the embodiment of the present application includes a data compression method and a data decompression method. The data compression method can provide a basis for stabilizing the performance of data decompression. Below, the data compression method provided in the embodiment of the present application is first described. For example, referring to the flow diagram of the data compression method shown in Figure 4, the data compression method includes but is not limited to the following S401 to S403.

[0098] S401, obtaining a data group, where the data group includes multiple bytes.

[0099] A data group is a unit of data compression. A data group may include one or more groups of data, a group of data may include one or more data, and a data may include one or more bytes, thus a data group may include one or more bytes. Since if a data group includes one byte, there is no need to compress the data group, and the embodiments of the present application are intended to compress or decompress the data group, the data processing method when the data group includes only one byte will not be discussed in the embodiments of the present application. Instead, the data processing method and process will be described using the case where the data group includes multiple bytes as an example.

[0100] The present embodiment does not limit the size of the data group. The sizes of data groups in different data processing environments can be the same or different. Below, the present embodiment uses a HOT-based data processing environment and an LSM-Tree-based data processing environment as examples to illustrate the composition and determination process of the data group.

[0101] For example, referring to FIG5 , a schematic diagram of a data processing environment based on HOT is shown. In a data processing environment based on HOT, data is stored in an independent heap located in the HOT, and the heap is a logical collection of multiple data. Each data stored in the heap corresponds to an address for obtaining the data and a key value (Key, K) for querying the data. The address for obtaining the data can be represented by a row identifier (RID). Based on the HOT, a binary (binary, B) +-tree index can be established based on the Key of each data. The bottom node of the index contains multiple tuples. A tuple includes the Key and RID of a data. Each tuple can be represented as<Key,RID> or<K,R> In a B+-Tree, each tuple is arranged in the order of the keys it contains. For example, the bottom node of the B+-Tree in Figure 5 contains 6 tuples, which are<K1,R1> 、<K2,R2> 、<K3,R3> 、<K4,R4> 、<K5,R5> and<K6,R6> , the 6 tuples are arranged in order of K from small to large.

[0102] The physical structure of HOT includes multiple storage blocks (Blocks), each of which stores multiple data. The size of each storage block can be, for example, 8KB-16KB. Since the values ​​(value, V) of different data are different, different data can also be represented by different values, so that different data can be distinguished by different values. The arrangement order of multiple data in each storage block has nothing to do with the K corresponding to the multiple data. For example, the storage block indicated by ① in Figure 5 includes data 2 represented by V2, data 3 represented by V3, and data 6 represented by V6. Data 2 corresponds to K2, data 3 corresponds to K3, and data 6 corresponds to K6. The arrangement order of the three data is data 3-data 6-data 2, and is not arranged in the order of the K of each data.

[0103] In a HOT-based data processing environment, multiple data within a storage block that meet compression conditions can be compressed. The multiple data that meet the compression conditions can be all of the data stored within the storage block, or a portion of the data stored within the storage block. The compression conditions can be set based on experience or specified by the user. For example, the compression condition can be a data length threshold. If the length of the data is greater than or equal to the length threshold, the data can be determined to meet the compression condition. Conversely, if the length of the data is less than the length threshold, the data can be determined to not meet the compression condition.

[0104] Multiple data used for compression in a storage block can be called a data group. That is, in the embodiment of the present application, when the data group is stored in the HOT, the data group can include multiple data in the storage block of the HOT. For example, in the compression process indicated by ② in Figure 5, it is determined that data 2 represented by V2 and data 3 represented by V3 are data that meet the compression conditions, then data 2 and data 3 can be regarded as a data group, that is, the data group includes data 2 and data 3 in the storage block.

[0105] For example, referring to FIG6 , a schematic diagram of a data processing environment based on LSM-Tree is shown. In the LSM-Tree-based storage engine, the stored data includes baseline data and incremental data. The baseline data can be the original data uploaded by the user, and the baseline data is stored in a sorted string table (SSTable), as shown in ① in FIG6 .

[0106] Physically, the smallest unit corresponding to an SSTable is a chunk, also known as a baseline database. The size of a baseline database can range from 64KB to 2MB. A baseline database consists of multiple storage chunks, each of which persistently stores one or more baseline data. Each chunk can be of different size, and can be any size within a reference range. This range can be empirically determined or user-specified, for example, 8KB to 16KB. Logically, each baseline data is arranged in the order of its corresponding K. K is the value used to locate each baseline data, similar to the K in Figure 5. The K corresponding to each baseline data is maintained in independent metadata, which records the distribution of K values ​​for each chunk. Furthermore, each baseline data can be represented by a different value (V). For example, in the SSTable shown in Figure 6, a chunk stores baseline data 1, represented by V1, and baseline data 2, represented by V2. Baseline data 1 corresponds to K1, and baseline data 2 corresponds to K2. Baseline data 1 and baseline data 2 are arranged in ascending order of their corresponding K1 and K2 values.

[0107] Another type of data stored in the LSM-Tree based storage engine, namely incremental data, is used to record updates or modifications to the baseline data, wherein modifications to the baseline data include but are not limited to insertions or deletions. As shown in ④ in Figure 6, the incremental data is stored in a memory table (MemTable). Optionally, the memory table can also be called a metadata table. An incremental data record in the memory table records the update or modification of a baseline data. For example, the first incremental data in the memory table in Figure 6 records the insertion of baseline data 3 represented by V3, and baseline data 3 corresponds to K3; the second incremental data record updates the value of baseline data 5 corresponding to V5 to V5'.

[0108] Since updates or modifications to baseline data do not directly affect the baseline data, when querying the baseline data based on business needs, it is necessary to merge the incremental data with the business-required data in the baseline data, obtain the merged baseline data, and then return it. For example, when querying baseline data 5 based on business needs, it is determined that the incremental data corresponding to baseline data 5 is to update the value of baseline data 5 from V5 to V5'. Therefore, the baseline data 5 and the incremental data corresponding to baseline data 5 can be merged, that is, the value of baseline data 5 is updated from V5 to V5', to obtain the merged baseline data 5, and return the merged baseline data 5.

[0109] Since the size of MemTable is limited, when MemTable grows to a certain size, it will trigger the automatic merging of incremental data and baseline data (as indicated by ⑤ in Figure 6) to generate new baseline data (as indicated by ⑥ in Figure 6). The new baseline data will not directly overwrite the old baseline data, but will be stored in a new (new) SSTable as shown in ⑦ in Figure 6. The new SSTable is located in the newly allocated storage space, and the new baseline data is persistently stored in the new SSTable. The baseline data before the merge is the old baseline data, and the SSTable that stores the old baseline data is the old SSTable.

[0110] The old baseline data will expire after meeting the expiration conditions, and the storage space occupied by the old SSTable will be released. The expiration conditions may include a time threshold after the new baseline data is generated. If the time threshold is reached after the new baseline data is generated, the old baseline data can be deleted, freeing up the storage space occupied by the old SSTable.

[0111] In an embodiment of the present application, data can be compressed during the process of generating new baseline data based on the merging of old baseline data and incremental data. In some cases, the baseline database is the minimum unit of data writing and space allocation during the input / output (I / O) process. Therefore, during the compression process, all baseline data in a baseline database can be compressed as a data group, or multiple consecutive modified storage blocks in a baseline database can be compressed as a data group. That is, when the data group is stored in the baseline database of the LSM-Tree, the data group includes data in multiple storage blocks in the baseline database.

[0112] Regardless of the data processing environment, and regardless of the type of data included in the data group, the data group may include multiple bytes, and the data group including multiple bytes may be compressed. The process of compressing the data group including multiple bytes can refer to S402 and S403 below, and will not be repeated here.

[0113] S402: Divide the multiple bytes into multiple data segments, where the multiple data segments include at least one of a matching data segment and a non-matching data segment, the bytes in the non-matching data segment are different from the bytes in each data segment before the non-matching data segment, the bytes in the matching data segment are the same as the bytes in a matched data segment before the matching data segment, and the distance between the matching data segment and the matched data segment is less than or equal to the reference distance.

[0114] The description of the matching data segment, the non-matching data segment, and the matched data segment can refer to the description of the general compression algorithm based on LZ77 above, and will not be repeated here. In the process of dividing multiple data segments in the embodiment of the present application, it is necessary to constrain the distance between the matching data segment and the matched data segment so that the distance between the final matching data segment obtained by division and the matched data segment is less than or equal to the reference distance. The size of the reference distance can be set based on experience or pre-defined according to user needs.

[0115] The embodiments of the present application do not limit the method for constraining the distance between the matching data segment and the matched data segment. For example, taking the method of constraining the distance between the matching data segment and the matched data segment by using a position array as an example, the process of dividing multiple bytes into multiple data segments may include: dividing multiple bytes into multiple alternative data segments; obtaining the position array of the alternative matching data segments; determining the alternative matching data segment as a matching data segment when there is an alternative distance less than or equal to the reference distance in at least one alternative distance; determining the alternative matching data segment as a non-matching data segment when at least one alternative distance is greater than the reference distance. The multiple alternative data segments include at least one of an alternative matching data segment or a non-matching data segment, and the bytes in the alternative matching data segment are the same as the bytes in at least one alternative matched data segment before the alternative matching data segment. The process of dividing multiple bytes into multiple alternative data segments is the same as the process of dividing multiple bytes into multiple data segments using the general compression algorithm based on LZ77, and will not be repeated here.

[0116] The position array used to constrain the distance between the matching data segment and the matched data segment can be generated in the process of dividing into multiple alternative data segments. The position array may include the position of the alternative matching data segment and the position of at least one alternative matched data segment corresponding to the alternative matching data segment. The position of the alternative matching data segment includes the position of each byte in the alternative matching data segment, and the position of at least one alternative matched data segment includes the position of each byte in at least one alternative matched data segment. In an embodiment of the present application, the position of each byte can be represented by the offset position of each byte in the byte stream, or by the order of each byte in the byte stream.

[0117] The embodiment of the present application provides a symbol system applied to a data processing method, which includes a representation method for each byte, a representation method for each data segment, and a representation method for a position array. Below, the composition and function of the position array are explained through the description of the symbol system. In this symbol system, the data group to be compressed is regarded as a byte stream of length n, and the byte stream is set to S. S[i] represents the i-th byte in S, and S[i..j] represents the data segment in S starting from the i-th byte and ending at the j-1-th byte, that is, the substring from the i-th byte to the j-1-th byte. S[i..j] can also be expressed in the form of S[i, j] or S[ij]. Correspondingly, in this symbol system, the first compressed data after compression is regarded as a byte stream of length m, and the compressed byte stream is set to D. The representation method of each compressed segment in D is the same as the representation method of each data segment in S, and will not be repeated here.

[0118] The length of the position array is n. Each byte in the byte stream corresponds to at least one ancestor position. The ancestor position of each byte is the position of the byte referenced by that byte. The position array records the correspondence between each byte position and its ancestor position. Therefore, position data can also be called an ancestor array.

[0119] The prototype position of a byte in a non-matching data segment is the position of the byte itself. For example, the prototype position of the first byte in a non-matching data segment is the position of the first byte itself. In the symbology provided in the embodiment of the present application, for a non-matching data segment S[i..i+l], Ancestor[i+k] is set to i+k in the position array, where k is greater than or equal to 0 and less than l, and l is the length of the candidate matching data segment and the candidate matched data segment (length, l).

[0120] The prototype positions of the bytes in the candidate matching data segment are the prototype positions of the corresponding bytes in the candidate matched data segment corresponding to the candidate matching data segment. If the candidate matched data segment is a non-matching data segment, that is, the candidate matched data segment does not reference other data segments, then the prototype positions of the bytes in the candidate matching data segment are the positions of the corresponding bytes in the candidate matched data segment corresponding to the candidate matching data segment. If the candidate matched data segment is a matching data segment, that is, the candidate matched data segment references other candidate matched data segments, then the prototype positions of the bytes in the candidate matching data segment may include at least one of the positions of the corresponding bytes in the candidate matched data segment or the positions of the corresponding bytes in other candidate matched data segments.

[0121] For example, if the alternative matched data segment referenced by candidate matching data segment 1 is candidate matched data segment 2, and candidate matched data segment 2 belongs to a non-matching data segment, and the byte referenced by the first byte in candidate matching data segment 1 is the first byte in candidate matched data segment 2, then the prototype position of the first byte in candidate matching data segment 1 is the position in the byte stream of the first byte in candidate matched data segment 2. For another example, if candidate matched data segment 2 belongs to candidate matching data segment 4, and candidate matching data segment 4 references candidate matched data segment 5, and candidate matched data segment 5 belongs to non-matching data segment 6, then candidate matching data segment 1 directly references candidate matched data segment 2, and candidate matching data segment 1 indirectly references candidate matched data segment 5. The first byte in candidate matching data segment 1 directly references the first byte in candidate matched data segment 2, and the first byte in candidate matching data segment 1 indirectly references the first byte in candidate matched data segment 5. Therefore, the prototype position of the first byte in candidate matching data segment 1 may include at least one of the position of the first byte in candidate matched data segment 2 and the position of the first byte in candidate matched data segment 5. In one possible implementation, if the candidate matching data segment is identical to multiple candidate matched data segments, the candidate matched data segment directly referenced by the candidate matching data segment may be the candidate matched data segment that is closest to the candidate matching data segment.

[0122] In the symbology provided in the embodiment of the present application, for an alternative matching data segment S[i..i+l], the device selects the alternative matched data segment directly referenced by the matching data segment S[i..i+l] as S[j..j+l], where j is less than i, that is, j is the byte before i. In the position array, Ancestor[i+k]=Ancestor[j+k] and Ancestor[i+k]=j+k, where k is greater than or equal to 0 and less than 1. That is, the prototype position of S[i+k] in the byte stream includes at least one of the prototype position of S[j+k] or the position of S[j+k].

[0123] In some cases, the position of the candidate matched data segment corresponding to the candidate matching data segment can be called the left boundary of the decompressed candidate matching data segment, so the position array can also be called the left boundary array. In this case, Ancestor[i] represents the left boundary of the decompressed S[i], that is, starting from the Ancestor[i]th character in S, the decompression can be done to get S[i].

[0124] After generating the position array, the distance between each byte position and the prototype position corresponding to each byte can be determined based on the position of each byte in the position array and the prototype position corresponding to each byte, thereby determining at least one candidate distance corresponding to each candidate matching data segment. For any candidate matching data segment, the at least one candidate distance is determined based on the position of the candidate matching data segment and the position of at least one candidate matched data segment.

[0125] In one possible implementation, at least one alternative distance between each candidate matching data segment and at least one candidate matched data segment referenced by each candidate matching data segment can be determined by comparing the position of the first byte in each candidate matching data segment with the prototype position corresponding to the first byte. For example, the candidate matched data segment directly referenced by candidate matching data segment 1 is candidate matched data segment 7, the first byte in candidate matching data segment 1 is at position 46 in the data group, and the first byte in candidate matched data segment 7 is at position 39 in the data group. Then, the alternative distance between candidate matching data segment 1 and candidate matched data segment 7 is 6, which is the number of bytes between the first byte in candidate matching data segment 1 and the first byte in candidate matched data segment 7. If the alternative matched data segment 7 belongs to the alternative matched data segment 8, and the alternative matched data segment 8 directly references the alternative matched data segment 9, then the alternative matched data segment 1 indirectly references the alternative matched data segment 9, and the position of the byte in the alternative matched data segment 9 corresponding to the first byte in the alternative matched data segment 1 is 25, then the alternative distance between the alternative matched data segment 1 and the alternative matched data segment 9 is 20.

[0126] After determining at least one alternative distance corresponding to each candidate matching data segment, whether each candidate matching data segment can be determined as a matching data segment can be determined based on the relative size of the alternative distance and the reference distance. For any candidate matching data segment, if at least one alternative distance corresponding to the candidate matching data segment is greater than the reference distance, it can be considered that if the candidate matching data segment is determined to be a matching data segment, then when decompressing the compressed segment obtained by compressing the matching data segment, it is necessary to decompress the matching data segment based on the compressed segment corresponding to the non-matching data segment that is farther away from the compressed segment, making it difficult to ensure the stability of the decompression. Therefore, the candidate matching data segment can be determined as a non-matching data segment. In the subsequent decompression process of the non-matching data segment, decompression can be achieved through the non-matching data segment itself, thereby improving the stability and efficiency of the decompression.

[0127] If there is an alternative distance less than or equal to the reference distance among at least one alternative distance corresponding to the alternative matching data segment, the alternative matching data segment can be determined as the matching data segment, and the alternative data segment to be matched corresponding to the alternative distance less than or equal to the reference distance is determined as the data segment to be matched corresponding to the matching data segment. If the number of alternative distances less than or equal to the reference distance is multiple, the data segment to be matched with the smallest alternative distance can be determined as the data segment to be matched corresponding to the matching data segment.

[0128] In the symbol system provided in the embodiments of the present application, the reference distance can be identified by W. The role of the reference distance is similar to a sliding window, and is used to control the stable left boundary of each matching data segment. Based on the above description, it can be seen that in the embodiments of the present application, the logic for finding the data segment to be matched is modified as follows: the prerequisite for the matching data segment S[i..i+l] to complete the matching is that for any k (0 <= k < l), (i + k) - Ancestor[i + k] <= W. If this is not satisfied, the alternative matching data segment S[i..i+l] is re-determined as a non-matching data segment. In the embodiments of the present application, by introducing a position array and a reference distance during the compression process, it is controlled that each matching data segment can be decompressed by a data segment to be matched whose distance from the matching data segment is within the reference distance. In some cases, after controlling the distance between the matching data segment and the data segment to be matched according to the position array, the position array can be deleted to reduce space occupancy.

[0129] In addition, when the data processing method provided in the embodiments of the present application compresses data, it does not need to consider the service characteristics of the data, can adapt to various scenarios where data needs to be compressed, and has strong scenario adaptability.

[0130] S403. Compress multiple data segments according to the data segments referred to by the multiple data segments, to obtain first compressed data including multiple compressed segments. Each compressed segment is obtained by compressing the corresponding data segment. The data segment referred to by any matching data segment is the data segment to be matched with the same bytes in any matching data segment, and the data segment referred to by any non-matching data segment is any non-matching data segment.

[0131] The process of compressing multiple data segments and outputting the first compressed data including multiple compressed segments can refer to the aforementioned process of compressing multiple data segments in the general compression algorithm based on LZ77, which will not be repeated here. The embodiment of the present application does not limit the storage location of the first compressed data formed after compression. For example, after the compression is completed, the data group before compression can be deleted, and the compressed first compressed data can be stored in the location originally used to store the data group, or the first compressed data can be stored in a location different from the data group. For example, in the data processing scenario shown in Figure 5, the first compressed data can be stored in the storage block to which the original data group belongs, or it can be stored independently in other locations in the HOT. Regardless of whether the first compressed data is stored in the storage block or in other locations in the HOT (such as the storage location indicated by ③ in Figure 5), it is necessary to ensure that any B+-Tree index established on the basis of the HOT can reference the storage location of the first compressed data. If the storage location of the first compressed data is different from the storage location of the data group, the value of the RID also needs to be changed so that the value of the changed RID indicates the storage location of the first compressed data. For example, as shown in the process indicated by ④ in FIG5 , the value in R2 corresponding to data 2 in the B+-Tree changes from the value pointing to the storage location of data 2 in the data group (indicated by the dotted arrow) to the value pointing to the storage location of data 2 in the first compressed data (indicated by the solid arrow). Correspondingly, the value in R3 corresponding to data 3 in the B+-Tree changes from the value pointing to the storage location of data 3 in the data group (indicated by the dotted arrow) to the value pointing to the storage location of data 3 in the first compressed data (indicated by the solid arrow).

[0132] In order to improve the efficiency of the subsequent decompression process, the positions of multiple matching data segments and multiple non-matching data segments can be recorded during the compression process, so that during the decompression process, the positions of the multiple matching data segments and the positions of the multiple non-matching data segments can be quickly decompressed according to the positions of the multiple matching data segments and the positions of the multiple non-matching data segments. Therefore, in a possible implementation, after compressing multiple data segments based on the data segments referenced by the multiple data segments, mapping information can also be generated, and the mapping information is used to indicate the mapping relationship between the starting positions of the multiple data segments and the starting positions of the multiple compressed segments. The mapping information can also be called an anchor index (Anchor Index), and the mapping information can include the starting position of each data segment in the multiple data segments, and the starting position of each compressed segment corresponding to each data segment. Alternatively, in order to save the space occupied by the mapping information, the mapping information may not include the positions of each data segment and each compressed segment, but may include the starting position of at least one reference data segment and the starting position of at least one reference compressed segment, and the at least one reference data segment corresponds one-to-one to the at least one reference compressed segment.

[0133] Optionally, at least one reference data segment can be first determined in the data group, and then at least one reference compressed segment corresponding to the at least one reference data segment can be determined. The at least one reference data segment can be determined based on the size of the data group. For example, a first step length for selecting a reference data segment can be determined based on the size of the data group, and at least one reference data segment can be determined in the data group based on the first step length. The first step length is the distance between two adjacent reference data segments. The size of the first step length can be set based on experience or specified by the user, and the first step lengths corresponding to different data groups can be the same or different. For example, if the size of the data group is 8KB, 1KB can be used as the first step length. The first data segment whose distance from its pre-compression position to the starting position of the data group is no less than 1KB, 2KB, ..., 7KB can be determined, respectively. The seven data segments determined are the reference data segments, and the starting position of each reference data segment can then be recorded. Subsequently, the reference compressed segments obtained by compressing each reference data segment need to be determined, and the starting position of each reference compressed segment in the first compressed data needs to be recorded.

[0134] Referring to Figure 7 , a schematic diagram of mapping information is shown. If the length of the data group before compression is 18 bytes, and the step length for determining the reference data segment is 6 bytes, then the first data segment BCD whose distance between its pre-compression position and the start position of the data group is no less than 6 bytes can be determined as a reference data segment, and the starting position of data segment BCD is recorded as 6 in the mapping information. Subsequently, the first data segment EA whose distance between its pre-compression position and the start position of the data group is no less than 12 bytes can be determined as another reference data segment, and the starting position of data segment EA can be recorded as 13 in the mapping information.

[0135] Taking the example of a compressed segment with a length of 1B obtained by compressing the matching data segment, after determining the starting position of the reference data segment BCD, it can be determined that the reference compressed segment corresponding to the reference data segment BCD is (2, 3, 1). In this case, the starting position of the reference compressed segment (2, 3, 1) can be recorded in the mapping information as 7. If the reference compressed segment corresponding to the reference data segment EA is EA, the distance from the starting position of the reference compressed segment EA in the mapping information can be 10.

[0136] Optionally, at least one reference compressed segment can be first determined in the first compressed data, and then at least one reference data segment corresponding to the at least one reference compressed segment can be determined. The at least one reference compressed segment can be determined based on the size of the compressed first compressed data. For example, a second step size for selecting a reference compressed segment can be determined based on the size of the first compressed data, and at least one reference compressed segment can be determined in the first compressed data based on the second step size. The second step size is the distance between two adjacent reference compressed segments. The size of the second step size can be set empirically or specified by the user, and the second step sizes corresponding to first compressed data of different sizes can be the same or different. For example, if the size of the data set is 4KB, a first step size of 1KB can be used. The first compressed segments whose distances from the compressed position to the starting position of the first compressed data are no less than 1KB, 2KB, and 3KB can be determined, respectively. The three determined compressed segments are the reference compressed segments, and the starting positions of each reference compressed segment can be recorded. Subsequently, the reference data segments for each reference compressed segment obtained by compression need to be determined, and the starting positions of each reference data segment in the data set need to be recorded.

[0137] In some cases, whether the mapping information records the starting positions of each data segment and each compressed segment, or the starting positions of at least one reference data segment and at least one reference compressed segment, both can be represented using the symbology provided in embodiments of the present application. In this symbology, the starting position of each data segment or the starting position of at least one reference data segment can be represented by AS, and the starting position of each compressed segment or the starting position of at least one reference compressed segment can be represented by AD.

[0138] In this case, the distribution characteristics of the starting position in the mapping information can be represented by an ordered binary array, that is, the mapping relationship between the starting position of the data segment and the starting position of the corresponding compressed segment data segment can be recorded by an ordered binary array. Among them, the ordered binary array can be represented as<AS,AD> or<AD,AS> .

[0139] Since the mapping information can indicate the mapping relationship between the starting positions of multiple data segments and the starting positions of multiple compressed segments, during the decompression process, the starting position of the compressed segment corresponding to the starting position of any data segment can be efficiently determined according to the mapping information, or the starting position of the data segment corresponding to the starting position of any compressed segment can be efficiently determined.

[0140] Based on the above description of the data compression method in the data processing method provided in the embodiment of the present application, it can be known that in the embodiment of the present application, after compression, not only the first compressed data obtained by compression can be output, but also the mapping information can be output. Therefore, when storing the content output after the compression is completed, it is necessary to determine the storage format according to the format of the output content, and perform persistent storage of the output content according to the storage format. For example, if the output content includes two parts, the first compressed data and the mapping information, the storage format may be a format for storing the first compressed data and the mapping information separately or a format for storing the first compressed data and the mapping information together. Different storage formats may be determined for different data processing scenarios. For example, in the data processing scenario based on HOT as shown in FIG5 , the storage format may be a format for storing the first compressed data and the mapping information together. In the data processing scenario based on LSM-Tree as shown in FIG6 , the storage format may be a format for storing the first compressed data and the mapping information separately. The first compressed data may be stored in SSTable, and the mapping information may be stored in MemTable.

[0141] For example, referring to FIG8 , a schematic diagram of a data compression process is shown. First, a data group is acquired and a determination is made as to whether all bytes have been processed. If not, the next byte is processed and a determination is made as to whether a matching data segment can be found for that byte. If a matching data segment can be found, the byte is temporarily recorded as a matching data segment. If no matching data segment can be found, the byte is temporarily recorded as a non-matching data segment. If all bytes have been processed, the final matching and non-matching data segments are determined based on the position array and reference distance, and the first compressed data and mapping information are output.

[0142] In the present application, by controlling the distance between the matching data segment and the matched data segment in the data group to be less than or equal to the reference distance, during decompression, the decompression of each compressed segment that needs to be decompressed can be achieved based on the compressed segment corresponding to the data segment whose distance between each data segment does not exceed the reference distance. Furthermore, before decompression, the boundary of the compressed data that needs to be decompressed can be determined based on the reference distance, and the compressed data can be decompressed within the boundary, thereby ensuring stable decompression performance.

[0143] The above describes the data compression process in the data processing method provided by the embodiment of the present application. Below, the data decompression process corresponding to the data compression process is described. For example, referring to FIG9 , a schematic diagram of a data decompression process is shown. The data decompression process includes but is not limited to the following S901 to S903.

[0144] S901, obtain a data processing instruction, the data processing instruction instructs to decompress target data from first compressed data, the first compressed data is obtained by compressing the data group to which the target data belongs, the data processing instruction includes the first starting position of the target data in the data group and the length of the target data.

[0145] The target data may be part of the data in the data group or all of the data in the data group. In the embodiment of the present application, the decompression process is described with the target data being part of the data in the data group.

[0146] Based on the above description, it can be known that a data group can be stored in a HOT, and the data group includes multiple data in a storage block of the HOT. In this case, the target data can include at least one data among the multiple data.

[0147] In an LSM-Tree-based data processing environment, data groups can also be stored in an LSM-Tree baseline database. To control read access discreteness and reduce the maintenance cost of incremental data changes, LSM-Tree-based data processing environments use block groups as the minimum unit of I / O writes and storage blocks as the minimum unit of I / O reads. Therefore, a data group includes data within multiple storage blocks in a block group in the baseline database, and the target data can include data within at least one of the multiple storage blocks.

[0148] Furthermore, in an LSM-Tree-based data processing environment, frequently accessed storage blocks can be stored in a block cache. For example, referring to FIG6 , the block cache shown in FIG6 ② is used to cache frequently accessed storage blocks. In some cases, data can be read in units of storage blocks during the data read process indicated in FIG6 ③.

[0149] The embodiments of the present application do not limit the method for obtaining data processing instructions. For example, when target data needs to be obtained during business operations, the device can autonomously generate a data processing instruction to instruct decompressing the first compressed data to obtain the target data, thereby obtaining the data processing instruction. Alternatively, the device can also receive data processing instructions sent by other devices or input by the user to obtain the data processing instruction.

[0150] If the data processing instruction is generated autonomously by the device, the device can first obtain the first starting position of the target data in the data group and the length of the target data before generating the data processing instruction. Taking a HOT-based data processing environment as an example, the device can use the B+-Tree index to search based on the K corresponding to the target data to determine the RID corresponding to the target data. The first starting position of the target data in the data group is determined based on the RID, and the length of the target data is determined based on the length of at least one data item included in the target data.

[0151] In the symbol system provided in the embodiment of the present application, the first starting position of the target data can be expressed as OFF, and the length of the target data can be expressed as (size, SZ).

[0152] S902, based on the first starting position and the reference distance, determine the second starting position of the second compressed data in the first compressed data, the second compressed data is used to decompress to obtain the target data, and the reference distance indicates the maximum value of the distance between any data segment in the target data and the data segment referenced by any data segment.

[0153] The embodiments of the present application do not limit the method for determining the second starting position based on the first starting position and the reference distance. In one possible implementation, the data group includes multiple data segments, the first compressed data includes multiple compressed segments, and each compressed segment is compressed by the corresponding data segment. In this implementation, the process of determining the second starting position may include: obtaining the starting position of each data segment in the data group; determining the first data segment to which the first position belongs among the multiple data segments based on the starting position of each data segment, where the first position is determined based on the difference between the first starting position and the reference distance; determining the starting position of a first compressed segment obtained by compressing the first data segment based on the starting position of the first data segment; and determining the starting position of the first compressed segment as the second starting position.

[0154] The difference between the first starting position and the reference distance can indicate the first position that is before the first starting position and whose distance from the first starting position is the reference distance. Since the distance between any data segment in the target data and the data segment referenced by any data segment is less than or equal to the reference distance, that is, the data segment referenced by any data segment in the target data segment is located at the first position or after the first position, decompression is started from the starting position of the first compressed segment corresponding to the first data segment belonging to the first position, which can ensure that the target data is obtained by decompression.

[0155] In one possible implementation, obtaining the starting position of each data segment in a data group may include: obtaining the length of each data segment; determining the starting position of each data segment in the data group based on the length of each data segment and the order of each compressed segment in the first compressed data, the order of any data segment in the data group is the same as the order of the compressed segments obtained by compressing any data segment in the first compressed data.

[0156] Since the order of any data segment in the data group is the same as the order of the compressed segments obtained by compressing any data segment in the first compressed data, the order of each data segment in the data group can be determined according to the order of each compressed segment in the first compressed data, and then the initial position of each data segment in the data group can be accurately determined according to the length of each data segment and the order of each data segment in the data group.

[0157] Exemplarily, since the lengths of multiple data segments are stored in multiple compressed segments, the lengths of each data segment and the order of each compressed segment in the first compressed data can be determined by traversing each compressed segment, and the cumulative lengths of all data segments before each data segment can be determined, thereby determining the starting position of each data segment in the data group.

[0158] For example, by reading the lengths of each data segment stored in multiple compressed segments, it is determined that the length of the first data segment corresponding to the first compressed segment is 3B. Then, it can be determined that the starting position of the first data segment in the data group is 0, and the starting position of the second data segment corresponding to the second compressed segment in the data group is 3.

[0159] In one possible implementation, the first position for determining the first data segment among the starting positions of the plurality of data segments can be determined based on the difference between the first starting position and a reference distance. Exemplarily, the difference between the first starting position and the reference distance is calculated. If the difference is greater than or equal to 0, the difference can be determined as the first position. If the difference is less than 0, 0 can be determined as the first position. That is, if the difference is greater than or equal to 0, it indicates that the position indicated by the difference is a position in the data group. If the difference is less than 0, it indicates that the difference is not a position in the data group. This also indicates that the distance between the first starting position and the starting position of the data group is less than or equal to the reference distance. Therefore, the starting position of the data group can be determined as the first position, and 0 can be determined as the first position.

[0160] In the symbology provided in the embodiments of the present application, the first position can be represented as TARGET_OFF, and the relationship between the first position, the first starting position, and the reference distance can be represented by TARGET_OFF = MAX(OFF - W, 0), where MAX is a function that finds the maximum value.

[0161] After determining the starting position of each data segment and the first position, the first data segment to which the first position belongs can be determined from the plurality of data segments based on the starting position of each data segment, thereby further determining the starting position of the first data segment. For example, among the starting positions of each data segment, a starting position A that is closest to the first position and smaller than or equal to the first position can be determined, and a starting position B that is closest to the first position and larger than or equal to the first position can be determined.

[0162] If both starting position A and starting position B are the same as the first position, it means that the first position is the starting position of the first data segment, and the data segment with the first position as the starting position can be determined as the first data segment. If starting position A is different from starting position B, and starting position A is smaller than the first position, and starting position B is larger than the first position, it means that the first position is located within the data segment with starting position A as the starting position, and the data segment with starting position A as the starting position can be determined as the first data segment. If starting position A is different from starting position B, and one of starting position A or starting position B is the same as the first position, the data segment corresponding to the same starting position as the first position can be determined as the first data segment.

[0163] In one possible implementation, the starting position of the first compressed segment obtained by compressing the first data segment is determined based on the starting position of the first data segment, including: obtaining mapping information, the mapping information is used to indicate the mapping relationship between the starting positions of multiple data segments and the starting positions of multiple compressed segments; and determining the starting position of the first compressed segment based on the mapping information and the starting position of the first data segment. The mapping information is the mapping information generated during the compression process. Since the mapping information may include the starting position of each data segment and the starting position of each compressed segment, and may also include the starting position of at least one reference data segment and the starting position of at least one reference compressed segment, the determination of the starting position of the first compressed segment based on the mapping information and the starting position of the first data segment can be divided into different situations according to the different contents included in the mapping information, for example, it can be divided into the following situations A1 and A2.

[0164] In case A1, the mapping information includes the starting position of each data segment and the starting position of each compressed segment. In case A1, in the process of determining the starting position of the first compressed segment based on the mapping information and the starting position of the first data segment, the starting position of the first compressed segment that has a mapping relationship with the first data segment can be directly found through the starting position of the first data segment. For example, in some cases, the mapping information stores a pair corresponding to the starting position of each data segment and the starting position of each compressed segment, and then the pair to which the starting position AS1 of the first data segment belongs can be found in each pair.<AS1,AD1> or a binary<AD1,AS1> , AD1 in the tuple is the starting position of the first compressed segment, and the starting position of the first compressed segment is the second starting position of the second compressed data.

[0165] In case A2, the mapping information includes the starting position of at least one reference data segment in each data segment and the starting position of at least one reference compressed segment compressed from the at least one reference data segment. In case A2, in determining the starting position of the first compressed segment based on the mapping information and the starting position of the first data segment, a first reference data segment is determined in the at least one reference data segment based on the starting position of the first data segment and the starting position of the at least one reference data segment. The first reference data segment is the reference data segment that precedes the first data segment and is closest to the first data segment. The starting position of the first reference compressed segment corresponding to the first reference data segment is determined as the starting position of the first compressed segment.

[0166] Taking the binary group corresponding to the starting position of at least one reference data segment and the starting position of at least one reference compressed segment stored in the mapping information as an example, the first reference data segment closest to the starting position AS1 of the first data segment and following the first data segment can be found according to the starting position of at least one reference data segment in each binary group. The starting position of the first reference data segment is AS2, and the binary group to which AS2 belongs is used to find the first reference data segment that is closest to the starting position AS1 of the first data segment and follows the first data segment.<AS2,AD2> or a binary<AD2,AS2> The starting position AD2 of the first reference compression segment is determined. The starting position AD2 of the first reference compression segment is the starting position of the first compression segment, and the starting position of the first compression segment is the second starting position of the second compressed data.

[0167] S903 , decompressing the second compressed data according to the second starting position and the ending position of the second compressed data to obtain target data, where the ending position of the second compressed data is determined based on the first starting position and the length of the target data.

[0168] In a possible implementation, before decompressing the second compressed data according to the second starting position and the ending position of the second compressed data to obtain the target data, it is also necessary to determine the ending position of the second compressed data. The embodiment of the present application does not limit the method for determining the ending position of the second compressed data. For example, the second data segment to which the second position belongs can be determined in multiple data segments according to the starting position of each data segment, and the second position is the sum of the first starting position and the length of the target data; the starting position of the third data segment adjacent to the second data segment is determined in the starting position of each data segment, and the third data segment is after the second data segment; the starting position of the second compressed segment corresponding to the third data segment is determined according to the starting position of the third data segment; and the starting position of the second compressed segment is determined as the ending position of the second compressed data.

[0169] The sum of the first starting position and the length of the target data can indicate the end position of the target data in the data group, that is, the second position can indicate the end position of the target data in the data group, and the second data segment to which the second position belongs is the data segment to which the end position of the target data belongs. Therefore, the starting position of the third data segment adjacent to the second data segment and after the second data segment is located after the end position of the target data, and decompression is stopped from the starting position of the second compressed segment corresponding to the third data segment, which can ensure that the target data is completely decompressed.

[0170] In one possible implementation, the method for determining the second data segment to which the second position belongs among multiple data segments based on the starting positions of each data segment can refer to the method for determining the first data segment to which the first position belongs among multiple data segments based on the starting positions of each data segment in S902, and will not be repeated here.

[0171] After determining the second data segment, a third data segment following and adjacent to the second data segment can be determined, and the starting position of the third data segment can be determined from the starting positions of the various data segments. The starting position of the second compressed segment corresponding to the third data segment can then be determined based on the starting position of the third data segment. The second compressed segment corresponding to the third data segment can be a compressed segment obtained by compressing the third data segment, or it can be a reference compressed segment corresponding to a reference data segment that is closest to and following the third data segment. The method for determining the starting position of the second compressed segment based on the starting position of the third data segment can refer to the method for determining the starting position of the first compressed segment based on the starting position of the first data segment in S902, and will not be further described here. The determined starting position of the second compressed segment is the ending position of the second compressed data.

[0172] After determining the second starting position and the second compressed data ending position, the second compressed data can be decompressed according to the second starting position and the second compressed data ending position to obtain the target data. In some cases, the data range of the second compressed data is larger than the data range of the target data, that is, the data obtained by decompressing the second compressed data includes the target data and data other than the target data. In this case, decompressing the second compressed data according to the second starting position and the second compressed data ending position to obtain the target data includes: decompressing the second compressed data according to the second starting position and the second compressed data ending position to obtain intermediate data; determining the target data ending position based on the first starting position and the length of the target data; and intercepting the target data from the intermediate data based on the first starting position and the target data ending position.

[0173] Optionally, the process of decompressing the second compressed data according to the second starting position and the ending position of the second compressed data to obtain the intermediate data can refer to the description of the general data decompression process based on LZ77 described above, and will not be repeated here. In addition, in one possible implementation, decompression of the second compressed data can also be started at the second starting position, and after completing the decompression of each compressed segment, the ending position of the data segment obtained by decompressing each compressed segment is recorded. When the recorded ending position of a data segment is greater than or equal to the second position, decompression is stopped to obtain the intermediate data.

[0174] The decompressed intermediate data includes the target data, so the target data can be intercepted from the intermediate data. The starting position of the target data is the first starting position, so before intercepting the target data, it is also necessary to determine the end position of the target data. Exemplarily, the end position of the target data can be determined as the sum of the first starting position and the length of the target data. For example, if the first starting position is 9 and the length of the target data is 56B, it can be determined that the end position of the target data is 64. Then, the data between the first starting position 9 and the end position 64 can be intercepted from the first data to obtain the target data.

[0175] In one possible implementation, the intermediate data includes a compressed segment that is not referenced by any data segment in the target data, and the compressed segment referenced by the compressed segment is located outside the intermediate data. That is, the compressed segment referenced by the compressed segment is located before the second starting position, resulting in the compressed segment being unable to be decompressed and restored. However, since no data segment in the target data references a compressed segment that cannot be decompressed and restored, the presence of such a compressed segment in the intermediate data does not affect the acquisition of the target data.

[0176] In addition, when the present application needs to query or call part of the data (such as a single-point query on a data segment or field), there is no need to decompress all the compressed data, which improves the decompression efficiency, reduces resource waste, and avoids affecting the operating efficiency of the business that needs to call or query the data.

[0177] In summary, since the reference distance indicates the maximum value of the distance between any data segment in the target data and the data segment referenced by any data segment, the data segment referenced by each data segment in the target data may be located within the target data, or within a portion of the data whose distance from the first starting position is less than or equal to the reference distance. Then, based on the first starting position and the reference distance, the second starting position of the second compressed data used for decompressing the target data can be determined, and it is ensured that the target data can be obtained by decompressing the second compressed data starting from the second starting position. In addition, based on the first starting position and the length of the target data, the end position of the second compressed data can be accurately determined, and decompression is stopped at the end position of the second compressed data. This not only ensures that the decompressed data includes the target data, but also avoids decompression of compressed data that does not need to be decompressed, thereby improving decompression efficiency and reducing resource waste.

[0178] The present application uses the determined second starting position and the ending position of the second compressed data as the decompression boundary, so that the decompression performance according to the clear decompression boundary is stable.

[0179] The above describes the data processing method of the embodiment of the present application. Corresponding to the above method, the embodiment of the present application also provides a data processing device. Figure 10 or Figure 11 is a schematic diagram of the structure of the data processing device provided in the embodiment of the present application. The device can implement the aforementioned data processing method through software, hardware, or a combination of both. The device can be applied to a device that provides cloud computing services. It should be understood that the device may include more additional modules than the modules shown or omit some of the modules shown therein, and the embodiment of the present application is not limited to this.

[0180] The data processing device shown in FIG10 includes:

[0181] An acquisition module 1001 is used to acquire a data processing instruction, where the data processing instruction indicates decompressing target data from first compressed data. The first compressed data is obtained by compressing the data group to which the target data belongs. The data processing instruction includes a first starting position of the target data in the data group and the length of the target data. A determination module 1002 is used to determine a second starting position of the second compressed data in the first compressed data based on the first starting position and a reference distance. The second compressed data is used to decompress the target data. The reference distance indicates the maximum value of the distance between any data segment in the target data and a data segment referenced by any data segment. A decompression module 1003 is used to decompress the second compressed data according to the second starting position and the end position of the second compressed data to obtain the target data. The end position of the second compressed data is determined based on the first starting position and the length of the target data.

[0182] In one possible implementation, the first compressed data includes multiple compressed segments, the data group includes multiple data segments, and each compressed segment is obtained by compressing the corresponding data segment; the determination module 1002 is used to obtain the starting position of each data segment in the data group; based on the starting position of each data segment, determine the first data segment to which the first position belongs in the multiple data segments, and the first position is determined based on the difference between the first starting position and the reference distance; based on the starting position of the first data segment, determine the starting position of the first compressed segment corresponding to the first data segment; and determine the starting position of the first compressed segment as the second starting position.

[0183] In one possible implementation, the determination module 1002 is used to obtain the length of each data segment; based on the length of each data segment and the order of each compressed segment in the first compressed data, the starting position of each data segment in the data group is determined, and the order of any data segment in the data group is the same as the order of the compressed segments obtained by compressing any data segment in the first compressed data.

[0184] In one possible implementation, the determination module 1002 is used to obtain mapping information, where the mapping information is used to indicate a mapping relationship between the starting positions of multiple data segments and the starting positions of multiple compressed segments; and determine the starting position of the first compressed segment based on the mapping information and the starting position of the first data segment.

[0185] In one possible implementation, the mapping information includes the starting position of at least one reference data segment in multiple data segments and the starting position of at least one reference compressed segment compressed by the at least one reference data segment; the determination module 1002 is used to determine the first reference data segment in the at least one reference data segment based on the starting position of the first data segment and the starting position of the at least one reference data segment, the first reference data segment being the reference data segment that is before the first data segment and is closest to the first data segment; and the starting position of the first reference compressed segment obtained by compressing the first reference data segment is determined as the starting position of the first compressed segment.

[0186] In one possible implementation, the determination module 1002 is further used to determine, in multiple data segments, the second data segment to which the second position belongs based on the starting positions of each data segment, where the second position is the sum of the first starting position and the length of the target data; determine, in the starting positions of each data segment, the starting position of a third data segment adjacent to the second data segment, where the third data segment is after the second data segment; determine, based on the starting position of the third data segment, the starting position of the second compressed segment corresponding to the third data segment; and determine the starting position of the second compressed segment as the ending position of the second compressed data.

[0187] In one possible implementation, the decompression module 1003 is used to decompress the second compressed data according to the second starting position and the ending position of the second compressed data to obtain intermediate data; determine the ending position of the target data according to the first starting position and the length of the target data; and intercept the target data in the intermediate data according to the first starting position and the ending position of the target data.

[0188] In a possible implementation, the data group is stored in the HOT, the data group includes multiple data in a storage block of the HOT, and the target data includes at least one data among the multiple data.

[0189] In a possible implementation, the data group is stored in a baseline database of the LSM-Tree, the data group includes data in multiple storage blocks in the baseline database, and the target data includes data in at least one storage block among the multiple storage blocks.

[0190] The data processing device shown in FIG11 includes:

[0191] An acquisition module 1101 is used to acquire a data group, which includes multiple bytes; a division module 1102 is used to divide the multiple bytes into multiple data segments, where the multiple data segments include at least one of a matching data segment or a non-matching data segment, the bytes in the non-matching data segment are different from the bytes in each data segment before the non-matching data segment, the bytes in the matching data segment are the same as the bytes in a matched data segment before the matching data segment, and the distance between the matching data segment and the matched data segment is less than or equal to the reference distance; a compression module 1103 is used to compress the multiple data segments based on the data segments referenced by the multiple data segments to obtain first compressed data including multiple compressed segments, each compressed segment is compressed by the corresponding data segment, the data segment referenced by any matching data segment is the matched data segment with the same bytes in any matching data segment, and the data segment referenced by any non-matching data segment is any non-matching data segment.

[0192] In one possible implementation, the partitioning module 1102 is used to partition multiple bytes into multiple alternative data segments, where the multiple alternative data segments include at least one of an alternative matching data segment or a non-matching data segment, and the bytes in the alternative matching data segment are the same as the bytes in at least one alternative matched data segment before the alternative matching data segment; obtain a position array of the alternative matching data segment, where the position array includes the position of the alternative matching data segment and the position of at least one alternative matched data segment corresponding to the alternative matching data segment; determine the alternative matching data segment as a matching data segment when there is an alternative distance less than or equal to a reference distance in at least one alternative distance, and the at least one alternative distance is determined based on the position of the alternative matching data segment and the position of at least one alternative matched data segment; and determine the alternative matching data segment as a non-matching data segment when at least one alternative distance is greater than the reference distance.

[0193] In a possible implementation, the device further includes a generation module; the generation module is used to generate mapping information, where the mapping information is used to indicate a mapping relationship between the starting positions of the multiple data segments and the starting positions of the multiple compressed segments.

[0194] In a possible implementation, a data group is stored in a HOT, and the data group includes a plurality of data in a storage block of the HOT.

[0195] In a possible implementation, the data group is stored in a baseline database of the LSM-Tree, and the data group includes data in multiple storage blocks in the baseline database.

[0196] It should be understood that the devices provided in FIG. 10 or FIG. 11 above are merely examples of the division of the functional modules described above when implementing their functions. In actual applications, the functions described above can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the devices and method embodiments provided in the above embodiments are based on the same concept. The specific implementation process is detailed in the method embodiment and will not be repeated here.

[0197] The present application also provides a computing device that can be configured as a server in the above-mentioned implementation environment. Referring to Figure 12 , Figure 12 is a schematic diagram of the hardware structure of a computing device provided in an embodiment of the present application. As shown in Figure 12 , computing device 600 includes: a bus 602, a processor 604, a memory 606, and a communication interface 608. The processor 604, the memory 606, and the communication interface 608 communicate with each other via bus 602. It should be understood that the present application does not limit the number of processors and memories in computing device 600.

[0198] Bus 602 may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, among others. Buses may be classified as address buses, data buses, control buses, and the like. For ease of illustration, FIG12 shows only one bus line, but this does not imply a single bus or type of bus. Bus 602 may include a path for transmitting information between various components of computing device 600 (e.g., memory 606, processor 604, and communication interface 608).

[0199] The processor 604 may include any one or more processors such as a central processing unit (CPU), a graphics processing unit (GPU), a microprocessor (MP), or a digital signal processor (DSP).

[0200] The memory 606 may include volatile memory, such as random access memory (RAM). The processor 604 may also include non-volatile memory, such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid state drive (SSD).

[0201] The memory 606 stores executable program codes, and the processor 604 executes the executable program codes to respectively implement the functions of the aforementioned modules, thereby implementing the data processing method. In other words, the memory 606 stores instructions for executing the data processing method.

[0202] The communication interface 608 uses a transceiver module such as, but not limited to, a network interface card or a transceiver to implement communication between the computing device 600 and other devices or a communication network.

[0203] The present application also provides a computing device cluster. The computing device cluster includes at least one computing device. The computing device can be configured as a server in the above implementation environment, such as a central server, an edge server, or a local server in a local data center.

[0204] Figure 13 is a schematic diagram of the structure of a computing device cluster provided in an embodiment of the present application. As shown in Figure 13, the computing device cluster includes at least one computing device 600. The memory 606 of one or more computing devices 600 in the computing device cluster may store the same instructions for executing the data processing method.

[0205] In some possible implementations, the memory 606 of one or more computing devices 600 in the computing device cluster may also store partial instructions for executing the data processing method. In other words, the combination of one or more computing devices 600 can jointly execute the instructions for executing the data processing method.

[0206] It should be noted that the memory 606 in different computing devices 600 in the computing device cluster can store different instructions, each for executing part of the functions of the data processing apparatus. In other words, the instructions stored in the memory 606 in different computing devices 600 can implement the functions of one or more modules in each module.

[0207] In some embodiments, one or more computing devices in a computing device cluster can be connected via a network. The network can be a wide area network (WAN) or a local area network (LAN), among others. FIG. 14 is a schematic diagram of a connection method for a computing device cluster provided in an embodiment of the present application. As shown in FIG. 14 , two computing devices 600 are connected via a network. Specifically, each computing device is connected to the network via a communication interface within the computing device.

[0208] It should be understood that the functions of the computing device 600 shown in FIG. 14 may also be performed by multiple computing devices 600 .

[0209] The present application also provides a computer program product comprising instructions. The computer program product may be software or a program product comprising instructions that can be run on a computing device or stored in any available medium. When the computer program product is run on at least one computing device, the computer program product causes the at least one computing device to execute a data processing method.

[0210] The present application also provides a computer-readable storage medium. The computer-readable storage medium can be any available medium that can be stored by a computing device or a data storage device such as a data center that contains one or more available media. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that instruct the computing device to execute the data processing method.

[0211] In the above embodiments, all or part of the embodiments may be implemented by software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described herein are generated. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrated therein. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive).

[0212] In this application, the terms "first," "second," and the like are used to distinguish between identical or similar items having substantially the same function or effect. It should be understood that "first," "second," and "nth" do not have a logical or temporal dependency, nor do they limit the quantity or order of execution. It should also be understood that although the following description uses the terms "first," "second," and the like to describe various elements, these elements should not be limited by these terms. These terms are simply used to distinguish one element from another.

[0213] It should also be understood that in the various embodiments of the present application, the size of the serial number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0214] In this application, the term "at least one" means one or more, and the term "plurality" means two or more. For example, "plurality of second devices" means two or more second devices. The terms "system" and "network" are often used interchangeably herein.

[0215] It should be understood that the terminology used in the description of the various examples herein is for the purpose of describing particular examples only and is not intended to be limiting. As used in the description of the various examples and the appended claims, the singular forms "a," "an," and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0216] It should also be understood that the term "and / or" as used herein refers to and encompasses any and all possible combinations of one or more of the listed items. The term "and / or" describes an association between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this application generally indicates that the associated objects are in an "or" relationship.

[0217] It should also be understood that the terms “if” and “if” may be interpreted to mean “when” or “upon” or “in response to determining” or “in response to detecting.” Similarly, the phrases “if it is determined that ” or “if [stated condition or event] is detected” may be interpreted to mean “upon determining ” or “in response to determining ” or “upon detecting [stated condition or event]” or “in response to detecting [stated condition or event],” depending on the context.

[0218] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and such modifications or substitutions should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.

[0219] In the above embodiments, all or part of the embodiments may be implemented using software, hardware, firmware, or any combination thereof. When implemented using software, all or part of the embodiments may be implemented in the form of program structure information. The program structure information includes one or more program instructions. When the program instructions are loaded and executed on a computing device, all or part of the processes or functions described in the embodiments of the present application are generated.

[0220] Those skilled in the art will understand that all or part of the steps to implement the above embodiments may be accomplished by hardware, or may be accomplished by a program to instruct the relevant hardware, and the program may be stored in a computer-readable storage medium, and the above-mentioned storage medium may be a read-only memory, a disk or an optical disk, etc.

[0221] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.), and signals involved in this application are all authorized by the user or fully authorized by all parties, and the collection, use, and processing of relevant data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the data involved in this application were obtained with full authorization.

[0222] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the protection scope of the technical solutions of the embodiments of the present application.

Claims

1. A data processing method, characterized in that: The method comprises: Obtaining a data processing instruction, the data processing instruction instructing to decompress target data from first compressed data, the first compressed data being obtained by compressing a data group to which the target data belongs, the data processing instruction including a first starting position of the target data in the data group and a length of the target data; Determining a second starting position of second compressed data in the first compressed data based on the first starting position and a reference distance, the second compressed data being used to decompress the target data, the reference distance indicating a maximum value of a distance between any data segment in the target data and a data segment referenced by any data segment; The second compressed data is decompressed according to the second starting position and the ending position of the second compressed data to obtain the target data, wherein the ending position of the second compressed data is determined based on the first starting position and the length of the target data.

2. The method according to claim 1, characterized in that The first compressed data includes a plurality of compressed segments, the data group includes a plurality of data segments, and each compressed segment is obtained by compressing a corresponding data segment; The determining, based on the first starting position and the reference distance, a second starting position of the second compressed data in the first compressed data comprises: Obtaining the starting position of each data segment in the data group; Determine, according to the starting positions of the respective data segments, a first data segment to which a first position belongs among the plurality of data segments, wherein the first position is determined based on a difference between the first starting position and the reference distance; Determine, according to the starting position of the first data segment, the starting position of the first compressed segment corresponding to the first data segment; The starting position of the first compression section is determined as the second starting position.

3. The method according to claim 2, characterized in that The obtaining the starting position of each data segment in the data group includes: Obtaining the length of each data segment; The starting position of each data segment in the data group is determined according to the length of each data segment and the order of each compressed segment in the first compressed data, and the order of any data segment in the data group is the same as the order of compressed segments obtained by compressing any data segment in the first compressed data.

4. The method according to claim 2 or 3, characterized in that: The determining, according to the starting position of the first data segment, the starting position of the first compressed segment corresponding to the first data segment includes: Acquire mapping information, where the mapping information is used to indicate a mapping relationship between starting positions of the plurality of data segments and starting positions of the plurality of compressed segments; The starting position of the first compressed segment is determined according to the mapping information and the starting position of the first data segment.

5. The method according to claim 4, characterized in that The mapping information includes a starting position of at least one reference data segment among the multiple data segments and a starting position of at least one reference compressed segment obtained by compressing the at least one reference data segment; determining the starting position of the first compressed segment according to the mapping information and the starting position of the first data segment includes: Determine a first reference data segment in the at least one reference data segment according to the starting position of the first data segment and the starting position of the at least one reference data segment, wherein the first reference data segment is a reference data segment that is before the first data segment and is closest to the first data segment; A starting position of a first reference compressed segment obtained by compressing the first reference data segment is determined as a starting position of the first compressed segment.

6. The method according to any one of claims 2 to 5, characterized in that: Before decompressing the second compressed data according to the second starting position and the end position of the second compressed data to obtain the target data, the method further includes: Determine, according to the starting positions of the respective data segments, a second data segment to which a second position belongs among the plurality of data segments, wherein the second position is the sum of the first starting position and the length of the target data; Determine a starting position of a third data segment adjacent to the second data segment among the starting positions of the data segments, wherein the third data segment is after the second data segment; Determine, according to the starting position of the third data segment, the starting position of the second compressed segment corresponding to the third data segment; The starting position of the second compression segment is determined as the ending position of the second compressed data.

7. The method according to any one of claims 1 to 6, characterized in that: The decompressing the second compressed data according to the second starting position and the ending position of the second compressed data to obtain the target data includes: decompressing the second compressed data according to the second starting position and the end position of the second compressed data to obtain intermediate data; Determining an end position of the target data according to the first starting position and the length of the target data; The target data is intercepted from the intermediate data according to the first starting position and the end position of the target data.

8. The method according to any one of claims 1 to 7, characterized in that: The data group is stored in a heap organization table HOT, the data group includes a plurality of data in a storage block of the HOT, and the target data includes at least one data among the plurality of data.

9. The method according to any one of claims 1 to 7, characterized in that: The data group is stored in a baseline database of a log structure merge tree, the data group includes data in a plurality of storage blocks in the baseline database, and the target data includes data in at least one storage block of the plurality of storage blocks.

10. A data processing device, characterized in that: The device comprises: an acquisition module, configured to acquire a data processing instruction, wherein the data processing instruction instructs to decompress target data from first compressed data, wherein the first compressed data is obtained by compressing a data group to which the target data belongs, and the data processing instruction includes a first starting position of the target data in the data group and a length of the target data; a determination module, configured to determine a second starting position of second compressed data in the first compressed data based on the first starting position and a reference distance, wherein the second compressed data is used to decompress the target data, and the reference distance indicates a maximum value of a distance between any data segment in the target data and a data segment referenced by any data segment; A decompression module is used to decompress the second compressed data according to the second starting position and the end position of the second compressed data to obtain the target data, and the end position of the second compressed data is determined based on the first starting position and the length of the target data.

11. The device according to claim 10, characterized in that The first compressed data includes a plurality of compressed segments, the data group includes a plurality of data segments, and each compressed segment is obtained by compressing a corresponding data segment; The determination module is used to obtain the starting position of each data segment in the data group; determine the first data segment to which the first position belongs among the multiple data segments according to the starting positions of the each data segment, the first position being determined based on the difference between the first starting position and the reference distance; determine the starting position of a first compressed segment corresponding to the first data segment according to the starting position of the first data segment; and determine the starting position of the first compressed segment as the second starting position.

12. The device according to claim 11, characterized in that The determination module is used to obtain the length of each data segment; determine the starting position of each data segment in the data group according to the length of each data segment and the order of each compressed segment in the first compressed data, and the order of any data segment in the data group is the same as the order of compressed segments obtained by compressing any data segment in the first compressed data.

13. The device according to claim 11 or 12, characterized in that The determination module is used to obtain mapping information, where the mapping information is used to indicate a mapping relationship between the starting positions of the multiple data segments and the starting positions of the multiple compressed segments; and determine the starting position of the first compressed segment based on the mapping information and the starting position of the first data segment.

14. The device according to claim 13, characterized in that The mapping information includes the starting position of at least one reference data segment among the multiple data segments and the starting position of at least one reference compressed segment compressed by the at least one reference data segment; the determination module is used to determine a first reference data segment in the at least one reference data segment according to the starting position of the first data segment and the starting position of the at least one reference data segment, the first reference data segment being a reference data segment that is before the first data segment and is closest to the first data segment; and determine the starting position of the first reference compressed segment obtained by compressing the first reference data segment as the starting position of the first compressed segment.

15. The device according to any one of claims 11 to 14, characterized in that: The determination module is further used to determine, according to the starting positions of the respective data segments, a second data segment to which a second position belongs among the multiple data segments, the second position being the sum of the first starting position and the length of the target data; to determine, among the starting positions of the respective data segments, a starting position of a third data segment adjacent to the second data segment, the third data segment being after the second data segment; to determine, according to the starting position of the third data segment, a starting position of a second compressed segment corresponding to the third data segment; and to determine the starting position of the second compressed segment as the ending position of the second compressed data.

16. The device according to any one of claims 10 to 15, characterized in that: The decompression module is used to decompress the second compressed data according to the second starting position and the end position of the second compressed data to obtain intermediate data; determine the end position of the target data according to the first starting position and the length of the target data; and intercept the target data in the intermediate data according to the first starting position and the end position of the target data.

17. The device according to any one of claims 10 to 16, characterized in that: The data group is stored in a heap organization table HOT, the data group includes a plurality of data in a storage block of the HOT, and the target data includes at least one data among the plurality of data.

18. The device according to any one of claims 10 to 16, characterized in that: The data group is stored in a baseline database of a log structure merge tree, the data group includes data in a plurality of storage blocks in the baseline database, and the target data includes data in at least one storage block of the plurality of storage blocks.

19. A computing device cluster, characterized in that: The computing device cluster includes at least one computing device, each computing device includes a processor, and the processor is coupled to a memory; The processor of the at least one computing device is used to execute instructions stored in the memory of the at least one computing device, so that the computing device cluster executes the data processing method according to any one of claims 1 to 9.

20. A computer program product comprising instructions, characterized in that When the instructions are executed by a computing device cluster, the computing device cluster executes the data processing method according to any one of claims 1 to 9.

21. A computer-readable storage medium, characterized in that: The method comprises computer program instructions. When the computer program instructions are executed by a computing device cluster, the computing device cluster executes the data processing method according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Data processing method, device, equipment and computer readable storage medium

    CN119210460B

  • Data processing method and device and electronic equipment

    CN113535709A

  • Compressed file processing method and device, computer equipment and storage medium

    CN115033381A

  • Data compression method and device, data decompression method and device, electronic equipment and chip

    CN116192154A

  • Data compression method and device, equipment and storage medium

    CN116418348A