A data compression method, a data decompression method and a device

By not compressing the root-sub-node and compressing the leaf-sub-node in the data storage of the tree structure, the problem of decompressing the entire data when querying specific data is solved, and query efficiency and storage space utilization are improved.

CN119210461BActive Publication Date: 2025-06-24HUAWEI CLOUD COMPUTING TECHNOLOGIES CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410386764.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2023-11-09
Filing Date
2024-03-29
Publication Date
2025-06-24
Estimated Expiration
2044-03-29

AI Technical Summary

Technical Problem

When using existing data compression algorithms, the entire compressed data needs to be decompressed when querying specific data, resulting in low query efficiency.

Method used

In the data storage of the tree structure, the data indicated by the root-sub node is not compressed and the data indicated by the leaf-sub node is compressed, and the data indicated by the leaf-sub node is only decompressed during query.

Benefits of technology

It improves the data query efficiency, reduces the storage space occupied by the data, and improves the utilization rate of the storage space.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119210461B_ABST
    Figure CN119210461B_ABST
Patent Text Reader

Abstract

A data compression method, a data decompression method and a device are disclosed, relating to the technical field of data storage. The data compression method includes: obtaining second data corresponding to first data, and then compressing the second data to obtain compressed data corresponding to the first data. The storage structure of the second data is a tree structure, and the bottom layer nodes of the tree structure include root-child nodes and leaf-child nodes having an index relationship with the root-child nodes; the compressed data includes data indicated by the root-child nodes and data indicated by the compressed leaf-child nodes. Since when querying data in the tree structure, it is to query data indicated by child nodes in sequence from the data indicated by the root node of the tree structure, the present application does not compress the data indicated by the root-child nodes. Therefore, when the computing device queries the data level indicated by the root-child nodes in the compressed data, it does not need to decompress the compressed data, which can improve the data query efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims the priority of a Chinese patent application with the application number 202311489545.9 and the application title "Data Compression Method and Computing Device", which was filed with the National Intellectual Property Administration on November 09, 2023. The entire content of this application is incorporated herein by reference. Technical Field

[0002] This application relates to the field of data storage technologies, and in particular, to a data compression method, a data decompression method, and a device. Background Art

[0003] Data compression technology has become a key technology in fields such as storage today. To improve the query efficiency of the storage system for compressed data, the storage device can construct data into a B-Tree or B+-Tree storage structure for storage, and then compress the data stored in the B-Tree or B+-Tree storage structure. For example, a prefix coding algorithm that compresses the common prefix of data, or a general compression algorithm that compresses the entire data. However, in the solutions adopting the above compression algorithms, when specific data needs to be queried, the specific position of the specific data in the compressed data cannot be known, resulting in the need to decompress all the compressed data to query the specific data, and the query efficiency is relatively low. Summary of the Invention

[0004] This application provides a data compression method, a data decompression method, and a device to solve the problem that when specific data needs to be queried, the specific position of the specific data in the compressed data cannot be known, resulting in the need to decompress all the compressed data to query the specific data, and the query efficiency is relatively low.

[0005] In a first aspect, this application provides a data compression method. This data compression method can be applied to a computer system or to a computing device that supports the computer system to implement the data compression method. For example, the computing device is a storage device. Here, taking the computing device to execute the data compression method provided in this embodiment as an example for description, the data compression method includes: obtaining second data corresponding to first data, and then compressing the second data to obtain compressed data corresponding to the first data. Among them, the storage structure of the second data is a tree structure, the bottom layer nodes of the tree structure include root-child nodes and leaf-child nodes having an index relationship with the root-child nodes, the first data includes the data indicated by the root-child nodes and the leaf-child nodes, and the compressed data includes the data indicated by the root-child nodes and the compressed data indicated by the leaf-child nodes.

[0006] When querying data in a tree structure, the data indicated by the root node of the tree structure is queried sequentially to the data indicated by the child nodes. In this application, the data indicated by the root-child nodes is not compressed. Therefore, when querying the data level of the root-child nodes in the compressed data, there is no need to decompress the compressed data, which can improve the data query efficiency. In addition, compressing the data indicated by the leaf-child nodes reduces the storage space occupied by the data and improves the utilization rate of the storage space.

[0007] In a possible scenario, the above tree structure can be a B+ tree (B+-Tree).

[0008] In a possible scenario, the data format of the first data is specifically a key-value pair (Key-Value). The data indicated by the leaf-child nodes is stored sequentially in the storage area of the leaf-child nodes according to the Key.

[0009] In this application, storing the data sequentially according to the Key is beneficial to determining that the data to be decompressed belongs to the Nth group according to the data indicated by the root-child nodes, and then decompressing the Nth group to obtain the data to be decompressed, which can improve the data query efficiency.

[0010] In a possible scenario, the storage structure of the first data can be one or more of the following: tree structure, hash table, skip list.

[0011] In a possible implementation, the storage area of the root-child nodes is used to store the first sub-data that meets the conditions, and the storage area of the leaf-child nodes is used to store the second sub-data that does not meet the conditions. Among them, the first data includes the first sub-data and the second sub-data.

[0012] Exemplarily, determine the first sub-data that meets the conditions from the first data and migrate the first sub-data to the storage area of the root-child nodes. In addition, determine the second sub-data that meets the conditions from the first data and migrate the second sub-data to the storage area of the leaf-child nodes.

[0013] In this application, according to the conditions, data with different attributes are respectively allocated to the storage areas corresponding to the root-child nodes or leaf-child nodes. Then, when compressing the second data, only the data indicated by the root-child nodes will not be compressed, and the data indicated by the leaf-child nodes will be compressed. Therefore, when querying the data level of the root-child nodes in the compressed data, there is no need to decompress the compressed data, which can improve the data query efficiency. And compressing the data indicated by the leaf-child nodes reduces the storage space occupied by the data and improves the utilization rate of the storage space.

[0014] In a possible scenario, the above conditions are used to indicate that the access metric of the data is greater than or equal to the threshold.

[0015] Exemplarily, the above access metrics include: access frequency or access rate. Furthermore, the first sub-data can be used to indicate hot data, and the second sub-data can be used to indicate cold data.

[0016] In this application, based on the access metrics, the hot data and cold data in the data indicated by the bottommost node are determined, and the hot data is migrated to the storage area of the root-sub node without compression. Furthermore, when storing Key-Value in the storage area of the root-sub node, since the data level indicated by the root-sub node in the compressed data is queried, there is no need to decompress the compressed data. Therefore, the Key-Value can be directly obtained, improving the efficiency of querying hot data.

[0017] Exemplarily, the access frequency or access rate of each data in the first data is determined from the metadata corresponding to the first data, and then the first sub-data whose access frequency or access rate meets the threshold and the second sub-data whose access frequency or access rate does not meet the threshold are determined.

[0018] In another possible scenario, the identifier numbers of two adjacent data indicated by the root-sub node have an interval of K + 1, where K is an integer greater than or equal to 1.

[0019] Exemplarily, K can indicate the number of second sub-data included in a group of second sub-data, that is, the data volume. The above identifier number can be a Key.

[0020] In this application, by defining the identifier numbers of two adjacent data in the storage area of the root-sub node, the data in the root-sub node can regularly indicate the data in the leaf-sub node, so that the speed of indexing the data in the leaf-sub node according to the data in the root-sub node can be faster, thereby improving the query efficiency of the data.

[0021] Exemplarily, K satisfies the following formula: K(K + X) ≥ N. Where X is a constant and N is the data volume included in a bottommost node.

[0022] In a possible implementation, the storage area of the bottommost node in the compressed data is also used to store: the right boundary or the left and right boundaries in the storage area of the leaf-sub node for each group of data indicated by the leaf-sub node. The left boundary of the first group of data is used to indicate the minimum byte length from the start address of the storage area corresponding to the leaf-sub node to the position where the first group of data is stored, and the right boundary of the first group of data is used to indicate the maximum byte length from the start address of the storage area corresponding to the leaf-sub node to the position where the first group of data is stored. The first group of data is any group of data indicated by the leaf-sub node.

[0023] In a possible scenario, the minimum byte length and the maximum byte length of the above first group of data match the Key in the first group of data.

[0024] In a possible scenario, the data indicated by the root - child node indicates the right boundary or both the left and right boundaries of each group of data.

[0025] Exemplarily, since there is an index relationship between the root - child node and the leaf - child node, and the minimum byte length and the maximum byte length of the first group of data match the Key in the first group of data, thus, based on the data in the root - child node, the right boundary or both the left and right boundaries of each group of data in the leaf - child node can be determined.

[0026] In the present application, when it is determined that the data belongs to the Nth group of data from the index relationship between the root - child node and the leaf - child node, the right boundary or both the left and right boundaries of the Nth group of data can be determined according to the stored right boundary or both the left and right boundaries of each group of data. Further, the right boundary or both the left and right boundaries of the Nth group of data are passed to the decompression engine for decompression, avoiding decompressing the entire bottom - most node, reducing the decompression area passed to the decompression engine, thereby significantly improving the decompression efficiency and enhancing the query speed of the data.

[0027] In a second aspect, the present application provides a data decompression method. This data decompression method can be applied to a computer system or to a computing device that supports the computer system to implement the data decompression method. For example, the computing device is a storage device. Here, taking the computing device executing the data decompression method provided in this embodiment as an example for description, the data decompression method includes: obtaining a data decompression request including the identifier of the first data, and querying the compressed data according to the identifier of the first data. Further, determining the target storage area in the compressed data that matches the identifier of the first data, and decompressing the data in the target storage area to obtain the first data. Among them, the storage structure of the compressed data is a tree structure, and the bottom - most nodes of the tree structure include root - child nodes and leaf - child nodes having an index relationship with the root - child nodes. The compressed data includes the data indicated by the root - child nodes and the compressed data indicated by the leaf - child nodes, and the target storage area is the storage area of the leaf - child nodes.

[0028] Since when querying the data of the tree structure, it is queried from the data indicated by the root node of the tree structure to the data indicated by the child nodes in sequence, and in the present application, the data indicated by the root - child nodes is not compressed. Therefore, when querying to the level of the data indicated by the root - child nodes in the compressed data, there is no need to decompress the compressed data, which can improve the data query efficiency. Also, when querying the data, according to the root - child nodes and in accordance with the index relationship, the smaller area (target storage area) where the data to be queried is located in the storage area of the leaf - child nodes can be accurately located. Further, only the compressed data in this smaller area needs to be decompressed, without decompressing all the data, reducing the amount of data for decompressing the compressed data, thereby further improving the data query efficiency.

[0029] In a possible scenario, the data format of the first data is Key-Value pairs, and the data indicated by the leaf nodes is stored in the storage area of the leaf nodes in sequence according to the Key.

[0030] In a possible scenario, the tree structure includes a B-tree.

[0031] In a possible implementation, the storage area of the bottommost node of the compressed data is also used to store: for each group of data indicated by the leaf nodes, the right boundary or both the left and right boundaries in the storage area of the leaf nodes. The left boundary of the first group of data is used to indicate the minimum byte length from the start address of the storage area corresponding to the leaf node to the position where the first group of data is stored, and the right boundary of the first group of data is used to indicate the maximum byte length from the start address of the storage area corresponding to the leaf node to the position where the first group of data is stored. The first group of data is any group of data indicated by the leaf nodes.

[0032] In a possible scenario, the data indicated by the root node indicates the right boundary or both the left and right boundaries of each group of data.

[0033] In a possible implementation, determining the target storage area in the compressed data that matches the identifier of the first data includes: according to the identifier of the first data, sequentially querying from the root node of the compressed data to the bottommost node, and then comparing the identifier of the first data with the data in the root node to determine the right boundary or both the left and right boundaries of the second group of data. Among them, the storage area of the leaf nodes in the bottommost node stores the first data or the address information of the first data. The second group of data includes the first data, and the right boundary or both the left and right boundaries of the second group of data are used to indicate the target storage area.

[0034] In this application, since there is an index relationship between the root node and the leaf nodes, therefore, the smaller area (i.e., the target storage area) where the first data is located in the storage area of the leaf nodes can be accurately located. Furthermore, only the compressed data in this smaller area needs to be decompressed, without decompressing all the data, reducing the amount of data for decompressing the compressed data, thereby improving the query efficiency of the data.

[0035] In a possible implementation, the identifier of the first data is compared with the data in the root-child node to determine the right boundary or the left and right boundaries of the second set of data, including: determining the first identifier and / or the second identifier in the identifiers of the data indicated by the root-child node that satisfy the first policy with the identifier of the first data. Then, from the index relationship, the second set of data in the leaf-child node corresponding to the first identifier and / or the second identifier is determined, so as to obtain the right boundary or the left and right boundaries of the storage area of the second set of data in the leaf-child node. Wherein, the first policy is used to indicate that the identifier of the first data is less than the first identifier, or the identifier of the first data is greater than the second identifier, or between the identifier of the first data is greater than or equal to the first identifier and less than or equal to the second identifier, and the first identifier is less than the second identifier.

[0036] In the present application, the identifier indicated by the root-child node is queried according to the identifier of the first data, and then the first identifier and / or the second identifier is determined. Since there is an index relationship between the root-child node and the leaf-child node, therefore, the smaller area (i.e., the target storage area) where the first data is located in the storage area of the leaf-child node is accurately located, and then only the compressed data in this smaller area needs to be decompressed, without decompressing all the data, reducing the amount of data for decompressing the compressed data, thereby improving the query efficiency of the data.

[0037] In a possible implementation, the storage area of the root-child node is used to store the first sub-data that meets the conditions, and the storage area of the leaf-child node is used to store the second sub-data that does not meet the conditions. Wherein, the first data includes the first sub-data and the second sub-data.

[0038] In a possible case, the condition is used to indicate that the access metric of the data is greater than or equal to the threshold.

[0039] Exemplarily, the access metric includes: the access frequency or the access rate; the first sub-data is used to indicate the hot data, and the second sub-data is used to indicate the cold data.

[0040] In another possible case, the interval between the identifier numbers of two adjacent data indicated by the root-child node is K + 1, and K is an integer greater than or equal to 1.

[0041] Exemplarily, K satisfies the following formula, K(K + X) ≥ N. Wherein, X is a constant, and N is the amount of data included in one bottommost node.

[0042] For more detailed content of the data decompression method provided in this aspect, reference can be made to the description of the first aspect above, and details are not elaborated here.

[0043] In a third aspect, the present application provides a data compression device. The data compression device is applied to a computer system or a computing device that supports the computer system to implement the data compression method, and the data compression device includes various modules for executing the data compression method in the first aspect or any optional implementation manner of the first aspect. The data compression device includes:

[0044] An acquisition module, configured to acquire second data corresponding to first data. The storage structure of the second data is a tree structure, and the bottom-level nodes of the tree structure include root-child nodes and leaf-child nodes having an index relationship with the root-child nodes; the first data includes the data indicated by the root-child nodes and the leaf-child nodes.

[0045] A compression module, configured to compress the second data to obtain compressed data corresponding to the first data, where the compressed data includes the data indicated by the root-child nodes and the compressed data indicated by the leaf-child nodes.

[0046] The data compression device provided in this aspect may include modules for executing the methods in any implementation manner in the first aspect.

[0047] In a fourth aspect, the present application provides a data decompression device. The data decompression device is applied to a computer system or a computing device that supports the computer system to implement the data decompression method, and the data decompression device includes various modules for executing the data decompression method in the first aspect or any optional implementation manner of the first aspect. The data decompression device includes:

[0048] An acquisition module, configured to acquire a data decompression request, where the data decompression request includes an identifier of the first data.

[0049] A query module, configured to query compressed data according to the identifier of the first data. Among them, the storage structure of the compressed data is a tree structure, and the bottom-level nodes of the tree structure include root-child nodes and leaf-child nodes having an index relationship with the root-child nodes, and the compressed data includes the data indicated by the root-child nodes and the compressed data indicated by the leaf-child nodes.

[0050] A determination module, configured to determine a target storage area in the compressed data that matches the identifier of the first data, where the target storage area is the storage area of the leaf-child nodes.

[0051] A decompression module, configured to decompress the data in the target storage area to obtain the first data.

[0052] The data decompression device provided in this aspect may include modules for executing the methods in any implementation manner in the second aspect.

[0053] Fifth aspect, the present application provides a computing device. The computing device includes a memory and a processor. The memory is used to store instructions. The processor executes the instructions to implement the method in the above first aspect or any possible implementation manner in the first aspect. For more details of the computing device in this aspect, reference can be made to the description of the following storage device or processing device, which will not be elaborated here.

[0054] Sixth aspect, the present application provides a computer-readable storage medium. A computer program or instructions are stored in the storage medium. When the computer program or instructions are executed by a computing device, the method in the above first aspect or any optional implementation manner in the first aspect is implemented.

[0055] Seventh aspect, the present application provides a computer program product. The computer program product includes a computer program or instructions. When the computer program or instructions are executed by a computing device, the method in the above first aspect or any optional implementation manner in the first aspect is implemented.

[0056] For the beneficial effects of the above second aspect to fifth aspect, reference can be made to the description of the first aspect or any implementation manner in the first aspect, which will not be elaborated here. Based on the implementation manners provided in the above aspects of the present application, further combinations can be made to provide more implementation manners. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] Figure 1 It is a schematic structural diagram of a data storage system provided by the present application;

[0058] Figure 2 It is a schematic flowchart of a data compression method provided by the present application;

[0059] Figure 3 It is a schematic structural diagram of the storage structure of the bottom-layer nodes of the first data and the second data provided by the present application;

[0060] Figure 4 It is a schematic structure of a B+-tree constructed based on HOT provided by the present application Figure 1 ;

[0061] Figure 5 It is a schematic structure of a B+-tree constructed based on HOT provided by the present application Figure 2 ;

[0062] Figure 6 It is a schematic structural diagram of an IOT provided by the present application;

[0063] Figure 7 It is a schematic flowchart of a data decompression method provided by the present application;

[0064] Figure 8Schematic structural diagram of a data compression device provided by this application;

[0065] Figure 9 Schematic structural diagram of a data decompression device provided by this application. Detailed implementation manners

[0066] In the field of data storage, to improve the query efficiency of data, data indexing methods are often used to achieve fast query of target data. For example, in a database, a tree structure index is often adopted. Since the data corresponding to the entire tree structure often occupies a large storage space, therefore, the data corresponding to the tree structure can be compressed through data compression technology to reduce the occupation of storage space. However, before querying or using the compressed data, the compressed data needs to be decompressed, resulting in device performance loss.

[0067] To reduce the device performance loss caused by decompressing the compressed data before querying or using it as described above. The following provides two possible processing solutions.

[0068] Processing solution 1: Use lightweight data encoding to compress data. The lightweight data encoding can include prefix encoding (which can also be called prefix encoding algorithm), run length encoding (RLE), suffix encoding, etc.

[0069] The following takes compressing data using prefix encoding in lightweight data encoding as an example for illustration: According to the orderliness of the data in the tree structure, the common prefix of adjacent data is extracted and only stored once, and the remaining suffix parts are stored separately, so as to achieve the purpose of data compression.

[0070] However, when querying specific data in the compressed data obtained by the above prefix encoding, since the compressed data is divided into a common prefix and a suffix part and the data is incomplete, it is impossible to determine the specific position of the specific data in the compressed data. Therefore, the compressed data needs to be fully decompressed before querying to obtain the specific data, resulting in low query efficiency for the specific data. The aforementioned full decompression of the compressed data means splicing the above common prefix and suffix parts to obtain complete data.

[0071] While the above solution has low query efficiency for the compressed data, since the proportion of the common prefix and suffix parts of the data changes according to the similarity between the data, the compression ratio of the data is unstable. In other words, the compression ratio of the data is directly related to the data distribution.

[0072] Processing solution 2: Use a general compression algorithm to compress data. The general compression algorithm can include LZ77 (Lempel-Ziv 77).

[0073] LZ77 usually compresses all the data within a single node in the tree - structured index. Compared with the above - mentioned processing solution 1, the general compression algorithm can provide a more stable compression ratio, thus significantly reducing the storage space occupied by the compressed data, and thereby reducing the storage cost.

[0074] However, when querying specific data in the compressed data obtained by the above - mentioned LZ77, it is still impossible to know the specific location of the specific data in the compressed data. Furthermore, it is necessary to decompress the data in the entire index node to query the aforementioned specific data, resulting in a low query efficiency for the specific data.

[0075] Based on the processing solution 2, the general compression algorithm is used to selectively compress the data in the tree - structured index. For example, compress the data in relatively inactive index nodes and do not compress the data in relatively active index nodes.

[0076] However, when the specific data to be queried is located in a relatively inactive index node, the data in the entire relatively inactive index node still needs to be decompressed to query the aforementioned specific data, resulting in a low query efficiency for the specific data.

[0077] Based on this, the present application provides a data processing method. In this method, the data is compressed data, the storage structure of the compressed data is a tree structure, and the bottom - most nodes of the tree structure include root - child nodes and leaf - child nodes having an index relationship with the root - child nodes. The compressed data includes the data indicated by the root - child nodes and the compressed data indicated by the leaf - child nodes. The method includes: compressing the data to obtain the above - mentioned compressed data. Or, the method includes: decompressing the above - mentioned compressed data to obtain the data indicated by the data decompression request.

[0078] For example, the data processing method is a data compression method. The data compression method includes: obtaining the second data corresponding to the first data, and then compressing the second data to obtain the compressed data corresponding to the first data. The storage structure of the second data is the same as the storage structure of the above - mentioned compressed data. The first data includes the data indicated by the root - child nodes and the leaf - child nodes.

[0079] Again, for example, the data processing method is a data decompression method. The data decompression method includes: obtaining a data decompression request including the identifier of the first data, querying the decompressed data according to the identifier of the first data. Then determining the target storage area in the compressed data that matches the identifier of the first data, and decompressing the data in the target storage area to obtain the first data. The target storage area is the storage area of the leaf - child nodes.

[0080] When querying data in a tree structure, the data indicated by the root node of the tree structure is queried sequentially to the data indicated by the child nodes. In this application, the data indicated by the root-child nodes is not compressed. Therefore, when querying the data level of the root-child nodes in the compressed data, there is no need to decompress the compressed data, which can improve the data query efficiency. Also, when querying data, according to the root-child nodes and the index relationship, the smaller area (target storage area) where the data to be queried is located in the storage area of the leaf-child nodes can be accurately located. Then, only the compressed data in this smaller area needs to be decompressed, without decompressing all the data, reducing the amount of data for decompressing the compressed data, thereby further improving the data query efficiency.

[0081] The terms used in the implementation part of this application are only used to explain the specific embodiments of this application, rather than aiming to limit this application. First, some concepts that this application may involve will be briefly introduced below.

[0082] B-Tree is a multi-way balanced tree. The storage area corresponding to the B-Tree stores an ordered set of a group of Keys. The B-Tree is composed of multiple index nodes, and these multiple index nodes can be divided into at least two layers. Each index node stores multiple Keys. When querying data in the B-Tree, based on the given Key, starting from the topmost index node of the B-Tree, binary search is performed layer by layer until the bottommost index node is queried.

[0083] B+-Tree. The difference between B+-Tree and B-Tree is that in B+-Tree, all Keys are only stored in the bottommost index nodes, and the intermediate index nodes only store the Keys for indexing.

[0084] Index node is the physical storage unit of data. A B+-Tree and a B-Tree are composed of multiple index nodes. Index nodes are also called index blocks or index pages, and their sizes are fixed, usually 8KB or 16KB.

[0085] Next, in conjunction with the accompanying drawings, the application scenarios of the above data compression method and data decompression method will be described. As Figure 1 shown, Figure 1 is a schematic structural diagram of a data storage system provided by this application. The data storage system includes a host 110 and a storage device 120. In Figure 1 the application scenario shown, users access data through application programs. The computer running these application programs can be called a "host". The host 110 can be a physical machine or a virtual machine. Physical machines include but are not limited to desktop computers, servers, laptop computers, and mobile devices.

[0086] In a possible example, the host 110 accesses the storage device 120 through a network to access data. For example, the network may include a switch 130.

[0087] In another possible example, the host 110 can also communicate with the storage device 120 through a wired connection. For example, a universal serial bus (USB) or a peripheral component interconnect express (PCIe) bus, etc.

[0088] Figure 1 The storage device 120 shown can be a centralized storage system. The characteristic of a centralized storage system is that there is a unified entry, and all data from external devices has to pass through this entry, which is the engine 121 of the centralized storage system. The engine 121 is the most core component in the centralized storage system, and many advanced functions of the storage system are implemented therein.

[0089] Such as Figure 1 As shown, there may be one or more controllers in the engine 121. Figure 1 Taking the case where the engine 121 includes one controller as an example. In a possible example, if the engine 121 has multiple controllers, there may be a mirror channel between any two controllers to implement the function of mutual backup between any two controllers, thereby avoiding the unavailability of the entire storage device 120 caused by hardware failures. It should be understood that if the engine 121 includes multiple controllers, the engine 121 can also be referred to as the array controller of the storage device 120.

[0090] The engine 121 also includes a front-end interface 1211 and a back-end interface 1214. The front-end interface 1211 is used to communicate with the host 110 to provide data access services for the host 110. The back-end interface 1214 is used to communicate with hard disks to expand the capacity of the storage device 120. Through the back-end interface 1214, the engine 121 can connect more hard disks to form a very large storage resource pool.

[0091] In terms of hardware, such as Figure 1As shown, the controller includes at least a processor 1212 and a memory 1213. The processor 1212 is a central processing unit (CPU), which is used to process data access requests from outside the storage device 120 (server or other storage systems), and is also used to process requests generated inside the storage device 120. Exemplarily, when the processor 1212 receives a data compression request or a data decompression request sent by the host 110 through the front-end interface 1211, it will temporarily store the data in these data compression requests in the memory 1213. When the total amount of data in the memory 1213 reaches a certain threshold, the processor 1212 sends the data stored in the memory 1213 to at least one of the mechanical hard disks 1221, mechanical hard disks 1222, solid state drive (SSD) 1223, or other hard disks 1224 through the back-end port for persistent storage.

[0092] The memory 1213 refers to the internal memory that directly exchanges data with the processor. It can read and write data at any time, and the speed is very fast. It serves as the temporary data memory for the operating system or other running programs. The memory includes at least two types of memories. For example, the memory can be either a random access memory or a read only memory (ROM). For example, the random access memory is DRAM, or SCM. DRAM is a semiconductor memory, which, like most random access memories (RAM), belongs to a volatile memory device. However, DRAM and SCM are only exemplary descriptions in this embodiment, and the memory can also include other random access memories, such as static random access memory (SRAM), etc. For the read only memory, for example, it can be a programmable read only memory (PROM), an erasable programmable read only memory (EPROM), etc.

[0093] In addition, the memory 1213 can also be a dual in-line memory module or a dual in-line memory module (DIMM), that is, a module composed of dynamic random access memory (DRAM), and can also be an SSD. In practical applications, multiple memories 1213 and different types of memories 1213 can be configured in the controller. The number and type of the memory 1213 are not limited in this embodiment. In addition, the memory 1213 can be configured to have a power retention function. The power retention function means that when the system loses power and then powers on again, the data stored in the memory 1213 will not be lost. A memory with a power retention function is called a non-volatile memory.

[0094] The memory 1213 stores software programs. The processor 1212 can implement writing and reading of compressed data by running the software programs in the memory 1213. For example, according to the first address space indicated by the read request and a preset address mapping table, the second compressed data is read from the second address space, and the second compressed data is decompressed according to the compression information indicated by the read request to obtain the first compressed data.

[0095] As Figure 1 shown, in this system, the engine 121 may not have a hard disk slot. The hard disk needs to be placed in the hard disk enclosure 122, and the backend interface 1214 communicates with the hard disk enclosure 122. The backend interface 1214 exists in the form of an adapter card in the engine 121. Two or more backend interfaces 1214 can be used simultaneously on one engine 121 to connect multiple hard disk enclosures. Alternatively, the adapter card can also be integrated on the motherboard. In this case, the adapter card can communicate with the processor 1212 through the PCIe bus.

[0096] It should be noted that Figure 1 only one engine 121 is shown in

[0097] the hard disk enclosure 122 includes a control unit 1225 and several hard disks. The control unit 1225 can have various forms. In one case, the hard disk enclosure 122 belongs to an intelligent disk enclosure, such as Figure 1As shown, the control unit 1225 includes a CPU and a memory. The CPU is used to perform operations such as address translation and data reading and writing. The memory is used to temporarily store data to be written to the hard disk or data read from the hard disk and to be sent to the controller. In another case, the control unit 1225 is a programmable electronic component, such as a data processing unit (DPU). The DPU has the versatility and programmability of a CPU, but is more specialized and can operate efficiently on network data packets, storage requests, or analysis requests. The DPU is distinguished from the CPU by a high degree of parallelism (requiring the processing of a large number of requests). Optionally, the DPU here can also be replaced by a processing chip such as a graphics processing unit (GPU), a neural-network processing unit (NPU), etc. Usually, the number of control units 1225 can be one, or two or more. The functions of the control unit 1225 can be offloaded to the network card 1226. In other words, in this embodiment, the hard disk enclosure 122 does not have a control unit 1225 inside, but the network card 1226 is used to complete data reading and writing, compression, decompression, and other computing functions. At this time, the network card 1226 is a smart network card. It can include a CPU and a memory. The CPU is used to perform operations such as compression, decompression, and data reading and writing. The memory is used to temporarily store data to be written to the hard disk or data read from the hard disk and to be sent to the controller. It can also be a programmable electronic component, such as a DPU. There is no ownership relationship between the network card 1226 and the hard disks in the hard disk enclosure 122, and the network card 1226 can access any hard disk in the hard disk enclosure 122 (such as Figure 1 the shown mechanical hard disk 1221, mechanical hard disk 1222, solid-state hard disk 1223, and other hard disks 1224), so it is more convenient to expand the hard disk when the storage space is insufficient.

[0098] According to the type of communication protocol between the engine 121 and the hard disk enclosure 122, the hard disk enclosure 122 may be a hard disk enclosure of a serial attached small computer system interface (SAS), or an NVMe (Non-Volatile Memory express) hard disk enclosure, as well as other types of hard disk enclosures. The SAS hard disk enclosure uses the SAS 3.0 protocol, and each enclosure supports 25 SAS hard disks. The engine 121 is connected to the hard disk enclosure 122 through an on-board SAS interface or a SAS interface module. The NVMe hard disk enclosure is more like a complete computer system, and the NVMe hard disk is inserted into the NVMe hard disk enclosure. The NVMe hard disk enclosure is then connected to the engine 121 through an RDMA port.

[0099] In an alternative implementation, the storage device 120 is a centralized storage system integrated with a disk controller. The storage device 120 does not have the above-mentioned hard disk enclosure 122, and the engine 121 is used to manage multiple hard disks connected through hard disk slots. The functions of the hard disk slots can be implemented by the backend interface 1214.

[0100] In another alternative implementation, Figure 1 the storage device 120 shown is a distributed storage system. The distributed storage system includes a computing device cluster and a storage device cluster. The computing device cluster includes one or more computing devices, and the computing devices can communicate with each other. The computing device can be a physical computing device, such as a server, a desktop computer, or a controller of a storage array, etc. Hardware-wise, the computing device can include a processor, a memory, and a network card, etc. Among them, the processor is a CPU, which is used to process data access requests from outside the computing device or requests generated inside the computing device. Exemplarily, when the processor receives a write data request sent by a user, it will temporarily store the data in these write data requests in the memory. When the total amount of data in the memory reaches a certain threshold, the processor sends the data stored in the memory to the storage device for persistent storage. In addition, the processor is also used to calculate or process data, such as metadata management, deduplication, data compression, decompression, virtualizing storage space, and address translation, etc. In one example, any computing device can access any storage device in the storage device cluster through the network. The storage device cluster includes multiple storage devices. A storage device includes one or more controllers, a network card, and multiple hard disks, and the network card is used to communicate with the computing device.

[0101] Exemplarily, the storage structure of the data stored in the storage device 120 can be a tree structure, such as a B-Tree or a B+-Tree, etc.

[0102] Exemplarily, in Figure 1In the data storage system, a user creates a compressed file in an industry-standard format (which can be referred to as the first compressed data), such as a JPEG (Joint Photographic Experts Group) file, on the host 110. The storage device 120 compresses the JPEG file using a preset compression algorithm to obtain a file in a private format (which can be referred to as the second compressed data), and the storage device 120 stores the second compressed data in the storage device. Herein, the file in the private format refers to the format of the file obtained after the storage device 120 compresses the JPEG file according to the preset compression algorithm. When the user initiates a data decompression request on the host 110, the storage device 120 obtains the corresponding second compressed data according to the data decompression request, decompresses the second compressed data to obtain the above-mentioned first compressed data, and finally the storage device 120 returns the first compressed data to the host 110.

[0103] In this application, when the storage device 120 compresses and decompresses data, it can be processed by a compression engine in the storage device, and the compression engine can be a compression engine that supports a high compression ratio. The function of the compression engine can be implemented by the processor 1212 in the storage device 120. For example, the compression engine can support compressing data using the above-mentioned general compression algorithm (such as LZ77) or a lightweight compression algorithm.

[0104] The following separately describes how the processing device compresses or decompresses data.

[0105] The first possible scenario: Based on the data storage system shown in Figure 1 A possible implementation manner for a processing device to compress data is provided, as shown in Figure 2 Shown in Figure 2 is a schematic flowchart of the data compression method provided in this application. The content shown in this embodiment can be executed by the processing device 200, and the processing device 200 can be the storage device 120 in Figure 1 . Please refer to Figure 2 . The data compression method provided in this embodiment includes steps S210 to S230.

[0106] S210. The processing device 200 obtains the first data.

[0107] The storage structure of the first data can be a tree structure, such as a B+-Tree or a B-Tree. The data included in each index node in the tree structure is arranged in sequence. For example, it is arranged in sequence according to the value of Key in the data. The data format of the first data can be in the form of key-value pairs (Key-Value).

[0108] In other examples of this application, the storage structure of the first data can also be a hash table. In a hash table, a hash function is used to map a Key to a bucket (hash bucket), and key-value pairs are stored within the bucket. Through the calculation of the hash function, specific Key-Value can be quickly searched for and accessed.

[0109] The storage structure of the above-mentioned first data is only an example provided in this application and should not be construed as a limitation to this application. In other examples of this application, any storage structure that can store data in Key-Value format can be the storage structure of the first data. For example, a skip list. Alternatively, the storage structure of the first data may not be limited, as long as the format of the first data is Key-Value.

[0110] In a possible implementation, the processing device 200 obtains the first data, including: the processing device 200 obtains the first data from the mechanical hard disk 1221, the mechanical hard disk 1222, the solid-state hard disk 1223, or other hard disks 1224. In other words, the processing device 200 obtains the first data locally.

[0111] For example, the processing device 200 obtains a data compression request a sent by the host 110, and this data compression request carries the specific location of the first data in the above-mentioned hard disk, such as a logical block address (LBA). The processing device 200 obtains the first data from the above-mentioned hard disk according to this LBA.

[0112] In a possible implementation, the processing device 200 obtains the first data, including: the processing device 200 obtains the first data sent by the host 110.

[0113] For example, the processing device 200 obtains a data compression request forwarded by the host 110 through the switch 130, and this data compression request carries the first data.

[0114] S220. The processing device 200 obtains the second data corresponding to the first data.

[0115] Among them, the storage structure of the second data is a tree structure, and the bottom-level nodes of this tree structure include root-child nodes and leaf-child nodes that have an index relationship with the root-child nodes. The first data includes the data indicated by the root-child nodes and the leaf-child nodes.

[0116] In a possible scenario, the above-mentioned tree structure can be a B+-Tree.

[0117] Such as Figure 2As shown, the storage structure of the second data is a B+-Tree. The B+-Tree includes multiple index nodes, which can be divided into at least two layers. All the data (Value) corresponding to the Keys is only stored in the bottommost index nodes (the bottommost nodes), and only some of the Keys are stored in the intermediate index nodes.

[0118] Exemplarily, the storage area corresponding to the bottommost nodes includes the storage data corresponding to the root-child nodes and the leaf-child nodes.

[0119] In a possible implementation, the processing device 200 obtaining the second data corresponding to the first data includes: the processing device 200 organizing the storage structure of the first data to obtain the second data.

[0120] Regarding the content that the processing device 200 organizes the storage structure of the first data to obtain the second data, the following three possible situations are provided.

[0121] In the first possible situation, the storage structure of the first data is a B+-Tree.

[0122] Since the overall storage structures of the first data and the second data are both B+-Trees, therefore, the processing device 200 only needs to organize the data indicated by the bottommost nodes in the first data to obtain the second data.

[0123] Taking one of the bottommost nodes (such as node a) among the multiple bottommost nodes included in the first data as an example for illustration. The processing device 200 determines the first sub-data that meets the conditions from the data indicated by node a, and migrates the first sub-data to the storage area of the root-child nodes. And, the processing device 200 determines the second sub-data that does not meet the conditions from the data indicated by node a, and migrates the second sub-data to the storage area of the leaf-child nodes.

[0124] Among them, the first sub-data is arranged in sequence according to the Key of the first sub-data in the storage area of the root-child nodes. For example, K3 / K6. The second sub-data is arranged in sequence according to the Key of the second sub-data in the storage area of the leaf-child nodes, for example, K1 / K2 / K4 / K5 / K7 / K8.

[0125] In this application, the processing device 200 distributes data with different attributes to the storage areas corresponding to the root-child nodes or leaf-child nodes according to conditions. Then, when compressing the second data, only the data indicated by the root-child nodes will not be compressed, and the data indicated by the leaf-child nodes will be compressed. Therefore, when the processing device 200 queries the data level indicated by the root-child nodes in the compressed data, it does not need to decompress the compressed data, which can improve the data query efficiency. And compressing the data indicated by the leaf-child nodes reduces the storage space occupied by the data and improves the utilization rate of the storage space.

[0126] In a possible example, the above conditions are used to indicate that the access metric of the data is greater than or equal to a threshold. The access metric includes: access frequency or access rate.

[0127] The processing device 200 determines the access frequency or access rate of each data (Key-Value) from the metadata of node a, and then determines the first sub-data whose access frequency or access rate meets the threshold, and the second sub-data whose access frequency or access rate does not meet the threshold.

[0128] Among them, the above first sub-data can be hot data, and the second sub-data can be cold data, and the thresholds corresponding to the access frequency or access rate are different. For example, the threshold corresponding to the access frequency is a (such as 50 times), and the threshold corresponding to the access rate is b (such as 0.2).

[0129] Exemplarily, the data indicated by node a includes K (Key) 1-V (Value) 1, K2-V2, K3-V3, K4-V4, K5-V5. The corresponding access frequencies of K1-V1, K2-V2, K3-V3, K4-V4, K5-V5 are 40, 20, 60, 20, 40 respectively. The processing device 200 compares the access frequencies corresponding to K1-V1, K2-V2, K3-V3, K4-V4, K5-V5 with the threshold a to determine that K1-V1, K2-V2, K4-V4, K5-V5 are the second sub-data, and K3-V3 is the first sub-data. Then, the processing device 300 migrates K2-V2 to the storage area corresponding to the root-child node, and migrates K1-V1, K2-V2, K4-V4, K5-V5 to the storage area corresponding to the leaf-child node.

[0130] It should be noted that the second sub-data on the left and right sides of the Key of the first sub-data can be used as two groups of second sub-data respectively. For example, K1-V1 and K2-V2 can be used as one group of second sub-data, and K4-V4 and K5-V5 can be used as another group of second sub-data.

[0131] In the present application, the processing device 200 determines the hot data and cold data in the data indicated by the bottommost node according to the access metrics, and migrates the hot data to the storage area of the root - child node without compression. Furthermore, when storing Key - Value in the storage area of the root - child node, since the processing device 200 does not need to decompress the compressed data when querying the data level indicated by the root - child node in the compressed data, the Key - Value can be directly obtained, thereby improving the efficiency of the processing device 200 when querying hot data.

[0132] In a second possible example, the above condition is used to indicate that the identifier number interval between two adjacent data indicated by the root - child node is K + 1, where K is an integer greater than or equal to 1.

[0133] K can indicate the number of second sub - data included in a group of second sub - data, that is, the data volume. Since the data in the bottommost node of the first data is arranged in sequence according to the Key, after the processing device 200 determines each group of second sub - data, it obtains the largest Key in the group of second sub - data, and uses the Key - Value corresponding to the largest Key + 1 as the first sub - data. The above identifier number is also the Key.

[0134] For the above K, a possible calculation method is provided below. K satisfies the following formula: K(K + X) ≥ N. Where X is a constant, and N is the data volume included in a bottommost node.

[0135] The above constant can be configured by the user, X is an integer greater than or equal to 1. N is the number of Key - Value included in a bottommost node.

[0136] Exemplarily, K is the smallest positive integer that satisfies the above formula.

[0137] Exemplarily, as Figure 3 shown, Figure 3 is a schematic diagram of the storage structure of the bottommost nodes of the first data and the second data provided by the present application. Figure 3 shows one of the bottommost nodes (node b) included in the second data, Figure 3 where a shows the arrangement of the data in the bottommost node of the first data, that is, the data in node b of the first data is arranged in ascending order according to the Key. The processing device 200 receives X configured by the user as 2, and determines that the data volume in node b is 8, and then determines that K is 2 according to the above formula.

[0138] Furthermore, as Figure 3 shown in b, Figure 3Figure b shows the logical structure of the root-child nodes and leaf-child nodes in the second data. The processing device 200 takes the data corresponding to K1 / K2 as a set of first sub-data, the data corresponding to K3 as second sub-data, the data corresponding to K4 / K5 as another set of first sub-data, the data corresponding to K6 as second sub-data, and the data corresponding to K7 / K8 as yet another set of first sub-data.

[0139] Thus, as Figure 3 shown in Figure c, Figure 3 Figure c shows the storage structure of the compressed second data. The processing device 300 migrates the data corresponding to K3 and K6 to the storage area corresponding to the root-child nodes, and migrates the data corresponding to K1, K2, K3, K4, K7, and K8 to the storage area corresponding to the leaf-child nodes. In this storage area, the data is stored in sequence according to the Key, and the data corresponding to K1, K2, K3, K4, K7, and K8 is compressed.

[0140] In this application, storing the data in sequence according to the Key is beneficial for the processing device 200 to determine that the data to be decompressed belongs to the Nth group according to the data indicated by the root-child nodes, and then decompress this Nth group to obtain the data to be decompressed, which can improve the query efficiency of the data.

[0141] It should be noted that for the content shown above, Figure 3 only K3 and K6 in K3-V3 and K6-V6 are migrated to the storage area corresponding to the root-child nodes for storage, and the complete data of K3-V3 and K6-V6 is still stored in the storage area corresponding to the leaf-child nodes, which is not shown in the figure. Furthermore, the processing device 200 can take the data corresponding to K1, K2, and K3 as a set of first sub-data, and the data corresponding to K4, K5, and K6 as another set of first sub-data.

[0142] In the second possible scenario, the storage structure of the first data is a B-Tree.

[0143] The processing device 200 converts the storage structure of the first data from a B-Tree to a B+-Tree, and then organizes the data indicated by the bottommost nodes in the B+-Tree to obtain the second data.

[0144] Since each layer of nodes in the B-Tree stores Key-Value, the processing device 200 needs to migrate the Key-Value stored in each layer of nodes to the bottommost nodes for storage, and only retain the Key for indexing in the foregoing layers of nodes.

[0145] Regarding the content that the processing device 200 organizes the data indicated by the bottommost nodes in the B+-Tree to obtain the second data, reference can be made to the description in the first possible scenario above, and details are not repeated here.

[0146] In the third possible scenario, the storage structure of the first data is a skip list.

[0147] The processing device 200 converts the storage structure of the first data from a skip list to a B+-Tree, and then organizes the data indicated by the bottommost nodes in the B+-Tree to obtain the second data.

[0148] Exemplarily, the processing device 200 reads out all the data in the skip list and then stores the read data into the constructed B+-Tree.

[0149] Regarding the content that the processing device 200 organizes the data indicated by the bottommost nodes in the B+-Tree to obtain the second data, reference can be made to the description in the first possible scenario above, and details are not elaborated here.

[0150] The above content is only the possible scenarios provided by this application and should not be construed as a limitation to this application. In other embodiments of this application, the processing device 200 additionally constructs a B+-Tree according to the first data for indexing the first data.

[0151] S230. The processing device 200 compresses the second data to obtain the compressed data corresponding to the first data.

[0152] Among them, the compressed data includes the data indicated by the root-child nodes and the compressed data indicated by the leaf-child nodes.

[0153] In a possible implementation manner, the processing device 200 compresses the second data to obtain the compressed data corresponding to the first data, including: the processing device 200 only compresses the data indicated by the leaf-child nodes in the second data to obtain the compressed data corresponding to the first data.

[0154] The processing device 200 may use a compression algorithm to compress the data indicated by the leaf-child nodes to obtain the compressed data corresponding to the first data.

[0155] For example, the compressed data obtained by using the above compression algorithm supports partial decompression. The foregoing compression algorithm may be the LZ77 algorithm, and the LZ77 algorithm can achieve partial decompression based on the right boundary of the data.

[0156] In a possible scenario, the above compression algorithm can achieve partial decompression based on the left and right boundaries of the data.

[0157] For the description of the left boundary or the right boundary, reference can be made to the content of the storage area of the bottommost nodes below, and details are not elaborated here.

[0158] In a possible embodiment, the storage area of the bottommost node in the above compressed data is also used to store: each group of data indicated by the leaf-child node, at the right boundary or the left and right boundaries in the storage area of the leaf-child node.

[0159] Among them, the left boundary of the first group of data is used to indicate the minimum byte length from the start address of the storage area corresponding to the leaf-child node to the position where the first group of data is stored. The right boundary of the first group of data is used to indicate the maximum byte length from the start address of the storage area corresponding to the leaf-child node to the position where the first group of data is stored, and the first group of data is any group of data indicated by the leaf-child node.

[0160] The minimum byte length and the maximum byte length of the above first group of data match the Key in the first group of data.

[0161] In a possible situation, the data indicated by the root-child node indicates the right boundary or the left and right boundaries of each group of data.

[0162] In a possible example, the storage area of the bottommost node stores the space (length) occupied by each data (Key-Value). The right boundary of the Nth group of data among the multiple groups of data in the leaf-child node can be determined according to the sum of the spaces occupied by the Key-Value in the first N groups of data, and the left boundary of the Nth group of data among the multiple groups of data in the leaf-child node can be determined according to the sum of the spaces occupied by the Key-Value in the first N - 1 groups of data.

[0163] Since there is an index relationship between the root-child node and the leaf-child node, and the minimum byte length and the maximum byte length of the first group of data match the Key in the first group of data, the processing device 200 can determine the right boundary or the left and right boundaries of each group of data in the leaf-child node according to the data in the root-child node.

[0164] Exemplarily, the data indicated by the root-child node are the data corresponding to K3 and K6, and the data indicated by the corresponding leaf-child node only include three groups of data: K1 / K2, K4 / K5, and K7 / K8. The processing device 200 can determine the right boundary or the left and right boundaries of the first group of data K1 / K2 in the leaf-child node according to K3 in the root-child node; can determine the right boundary or the left and right boundaries of the second group of data K4 / K5 in the leaf-child node according to K3 and K6 in the root-child node; can determine the right boundary or the left and right boundaries of the third group of data K7 / K8 in the leaf-child node according to K6 in the root-child node.

[0165] It should be noted that the right boundary or the left and right boundaries of the above first group of data are used to indicate the right boundary or the left and right boundaries before the first group of data is compressed.

[0166] In this application, when the processing device 200 determines that the data belongs to the Nth group of data from the index relationship between the root-subnode and the leaf-subnode, it can determine the right boundary or the left and right boundaries of the Nth group of data according to the right boundary or the left and right boundaries of each group of data stored above, and then transfer the right boundary or the left and right boundaries of the Nth group of data to the compression engine for decompression, avoiding decompressing the entire bottommost node, reducing the decompression area transferred to the compression engine, thus significantly improving the decompression efficiency and enhancing the query speed of the data.

[0167] In a possible embodiment, as Figure 4 shown, Figure 4 is a schematic structure diagram of a B+-tree constructed based on HOT provided by this application Figure 1 . The type of the first data is a heap-organized table (HOT). HOT stores the data in an independent heap. The heap logic is a set of data. Each data stored in the heap has a unique address, that is, the data address (row identifier, RID). A B+-tree can be established based on the Key on the HOT. The bottommost nodes of the B+-tree include multiple Key-RIDs and are arranged in sequence according to the Key.

[0168] Based on the content in S220 above, the processing device 200 divides the bottommost nodes of the B+-tree into root-subnodes and leaf-subnodes. The data indicated by the root-subnodes includes K3 and K6, and the data indicated by the leaf-subnodes includes K1-R(RID)1, K2-R2, K3-R3, K4-R4, K5-R5, K6-R6, K7-R7, K8-R8. The data indicated by the leaf-subnodes is stored in sequence according to the Key in the storage area corresponding to the leaf-subnodes.

[0169] The processing device 200 compresses the data indicated by the leaf-subnodes as a whole, so that the compressed data obtained includes the data indicated by the root-subnodes (K3 and K6) and the compressed data indicated by the leaf-subnodes (compressed K1-R1, K2-R2, K3-R3, K4-R4, K5-R5, K6-R6, K7-R7, K8-R8).

[0170] In a possible example, the processing device 200 may still store the compressed data in the storage area corresponding to the bottommost nodes. The storage area of the bottommost nodes includes the storage area indicated by the leaf-subnodes.

[0171] In another possible example, the processing device 200 stores the compressed data in other storage areas, or stores the root-subnodes in storage areas outside the storage area of the bottommost nodes.

[0172] For the content of the overall compression of the data indicated by the leaf - child nodes by the processing device 200, reference can be made to the description of S230 above, which will not be elaborated here.

[0173] In a possible scenario, the processing device 200 compresses the data indicated by the leaf - child nodes under all the bottom - most nodes in the above B+-tree.

[0174] In another possible scenario, the processing device 200 determines to compress the data indicated by the leaf - child nodes under some of the bottom - most nodes in the above B+-tree according to the screening conditions.

[0175] The above screening conditions can be that the access frequency or access rate of the data in each bottom - most node is less than or equal to the threshold c. If the access frequency of the data in the bottom - most node a is 100, which is less than the threshold c (120), then the processing device 200 compresses the data indicated by the leaf - child nodes under the bottom - most node a. If the access frequency of the data in the bottom - most node a is 150, which is greater than the threshold c (120), then the processing device 200 does not compress the data in the bottom - most node a.

[0176] If the access rate of the data in the bottom - most node b is 0.2, which is less than the threshold c (0.4), then the processing device 200 compresses the data indicated by the leaf - child nodes under the bottom - most node b. If the access rate of the data in the bottom - most node b is 0.5, which is greater than the threshold c (0.4), then the processing device 200 does not compress the data in the bottom - most node b.

[0177] In a possible scenario, the root - child nodes in the B+-tree can store data in the Key - Value form. As Figure 5 shown, Figure 5 is the structural schematic of the B+-tree constructed based on HOT provided by this application Figure 2 , the processing device 200 can divide the bottom - most nodes of the B+-tree into root - child nodes and leaf - child nodes. The data indicated by the root - child nodes includes K3 - R3 and K6 - R6, and the data indicated by the leaf - child nodes includes K1 - R1, K2 - R2, K4 - R4, K5 - R5, K7 - R7, K8 - R8. The data indicated by the leaf - child nodes is stored sequentially according to the Key in the storage area corresponding to the leaf - child nodes.

[0178] In this application, the processing device 300 stores the data such as RID that occupies less space into the storage area corresponding to the root - child nodes. When the occupation of the storage space is not high, it can improve the efficiency of querying the RID indicated in the root - child nodes, and then read the data indicated by the RID, thereby improving the efficiency of the entire data query process.

[0179] In another possible embodiment, asFigure 6 As shown Figure 6 is a schematic structural diagram of the IOT provided by this application. The type of the first data is an index-organized table (IOT). The IOT itself is a B+-Tree. The data is directly stored in the bottommost nodes of the B+-Tree. Each bottommost node contains multiple Key-Values and is arranged in sequence according to the Key. Only the Keys for indexing are stored in the other nodes of the B+-Tree.

[0180] Based on the content in S220 above, the processing device 200 divides the bottommost nodes of the B+-tree into root-sub nodes and leaf-sub nodes. The data indicated by the root-sub nodes includes K3 and K6, and the data indicated by the leaf-sub nodes includes K1-V1, K2-V2, K3-V3, K4-V4, K5-V5, K6-V6, K7-V7, K8-V8. The data indicated by the leaf-sub nodes is stored in sequence according to the Key in the storage area corresponding to the leaf-sub nodes.

[0181] The processing device 200 performs overall compression on the data indicated by the leaf-sub nodes, so that the obtained compressed data includes the data indicated by the root-sub nodes (K3 and K6) and the compressed data indicated by the leaf-sub nodes (compressed K1-V1, K2-V2, K3-V3, K4-V4, K5-V5, K6-V6, K7-V7, K8-V8). The storage area of the root-sub nodes is only used to store the Keys of the data for indexing the data indicated by the leaf-sub nodes.

[0182] For more details of this embodiment, reference can be made to the description of the embodiment shown above Figure 4 and will not be elaborated here.

[0183] In this application, only the Keys are stored in the root-sub nodes for indexing, avoiding the situation that the amount of Value data may be large, resulting in more storage space occupation, thus improving the utilization rate of the storage space. And, both the Keys and Values are stored in the storage area of the leaf-sub nodes, and then overall compression is performed on the data indicated by the leaf-sub nodes, further reducing the occupation of the storage space and improving the utilization rate of the storage space.

[0184] In a possible scenario, the root-child nodes in the B+-tree can store data in the form of Key-Value. For example, the processing device 200 can divide the bottom-level nodes of the B+-tree into root-child nodes and leaf-child nodes. The data indicated by the root-child nodes includes K3-V3 and K6-V6, and the data indicated by the leaf-child nodes includes K1-V1, K2-V2, K4-V4, K5-V5, K7-V7, and K8-V8. The data indicated by the leaf-child nodes is stored in sequence according to the Key in the storage area corresponding to the leaf-child nodes.

[0185] The second possible scenario: After the processing device 200 compresses the second data to obtain the compressed data, the processing device 200 receives a data decompression request from the host 110 and decompresses part of the data in the compressed data. As Figure 7 shown, Figure 7 is a schematic flowchart of the data decompression method provided by this application. The content shown in this embodiment can be executed by the processing device 200, and the processing device 200 can be Figure 1 the storage device 120 in. Please refer to Figure 7 , the data decompression method provided by this embodiment includes steps S710 to S740.

[0186] S710. The processing device 200 obtains a data decompression request.

[0187] The data decompression request includes the identifier of the first data. For example, the identifier of the first data is the Key of the first data.

[0188] In a possible implementation manner, the processing device 200 obtains a data decompression request, including: the processing device 200 obtains the data decompression request sent by the host 110.

[0189] The data decompression request can be generated by a user's operation on the host 110. For example, the host 110 receives a click operation, a swipe operation, or a voice operation, etc. on the front-end display interface of the host 110 by the user, and generates a data decompression request.

[0190] Exemplarily, the host 110 receives the user's click on the control component corresponding to the first data on the front-end display interface of the host 110, then determines the identifier of the first data, and encapsulates the identifier of the first data into the data decompression request, so as to send the data decompression request carrying the identifier of the first data to the processing device 200.

[0191] S720. The processing device 200 queries the compressed data according to the identifier of the first data.

[0192] Among them, the storage structure of the compressed data is a tree structure. The bottom - most nodes of this tree structure include root - child nodes and leaf - child nodes having an index relationship with the root - child nodes. The compressed data includes the data indicated by the root - child nodes and the compressed data indicated by the leaf - child nodes.

[0193] For the data indicated by the root - child nodes and the compressed data indicated by the leaf - child nodes, reference can be made to the content of the data stored in the storage area of the root - child nodes and the storage area of the leaf - child nodes shown above. Details are not described here. Figure 2 Shown in the storage area of the root - child nodes and the storage area of the leaf - child nodes as above, the content of the data stored therein will not be elaborated here.

[0194] Since the storage structure of the compressed data is a tree structure, the processing device 200 sequentially queries the data indicated by the child nodes from the data indicated by the root node in the compressed data. For example, the processing device 200 starts querying from the first - layer index node of the compressed data, then indexes to the second - layer index node according to the content queried in the first layer, and then queries in the second - layer index node, and indexes to the third - layer index node according to the content queried in the second - layer index node, repeating the foregoing steps until the Nth group of data where the first data is located is queried in the compressed data. N is an integer greater than or equal to 1.

[0195] When the processing device 200 queries the data indicated by the root - child nodes of the bottom - most nodes in the compressed data, it first compares the identifier of the first data with the data in the root - child nodes to determine the second group of data to which the first data belongs.

[0196] In a possible implementation, the processing device 200 determines the first identifier and / or the second identifier in the identifier of the data indicated by the root - child nodes that satisfy the first policy with the identifier of the first data, and then determines the second group of data in the leaf - child nodes corresponding to the first identifier and / or the second identifier from the index relationship.

[0197] For example, the first policy is used to indicate that the identifier of the first data is less than the first identifier.

[0198] In this example, the first identifier is the smallest identifier indicated by the root - child nodes. Exemplarily, the storage area of the root - child nodes includes K3, K6, and the first identifier is K3.

[0199] Exemplarily, as Figure 3 shown in b Figure 3 shown in b in the figure is one of the bottom - most nodes in the tree structure. That is, the processing device 200 determines through the above - mentioned query steps that the first data is located at Figure 3The bottommost node shown in b. If the identifier of the first data is K1, the processing device 200 performs a binary search within the root-subnode and determines that K1 is less than K3 (the first identifier) in the root-subnode. Therefore, based on the index relationship between the root-subnode and the leaf-subnode, it is determined that the data corresponding to K1 is located in the first set of data indicated by the leaf-subnode. This first set of data can also be referred to as the data indicated by the first leaf-subnode among multiple leaf-subnodes.

[0200] For another example, this first strategy can also be used to indicate that the identifier of the first data is greater than the second identifier.

[0201] In this example, the second identifier is the maximum identifier indicated by the root-subnode. Exemplarily, the storage area of the root-subnode includes K3 and K6, and the second identifier is K6.

[0202] Exemplarily, as Figure 3 shown in b, if the identifier of the first data is K7, the processing device 200 performs a binary search within the root-subnode and determines that K7 is greater than K6 (the second identifier) in the root-subnode. Therefore, based on the index relationship between the root-subnode and the leaf-subnode, it is determined that the data corresponding to K6 is located in the third set of data indicated by the leaf-subnode. This third set of data can also be referred to as the data indicated by the third leaf-subnode among multiple leaf-subnodes.

[0203] Again, this first strategy is used to indicate that the identifier of the first data is greater than or equal to the first identifier and less than or equal to the second identifier, and the first identifier is less than the second identifier.

[0204] In this example, the first identifier and the second identifier are the identifiers adjacent to the identifier of the first data among the identifiers indicated by the root-subnode, or the first identifier or the second identifier is the identifier that is the same as the identifier of the first data among the identifiers indicated by the root-subnode. Exemplarily, the storage area of the root-subnode includes K3, K6, and K9. If the identifier of the first data is K5, then the first identifier is K3 and the second identifier is K6.

[0205] Exemplarily, as Figure 3 shown in b, if the identifier of the first data is K5, the processing device 200 performs a binary search within the root-subnode and determines that K5 is greater than K3 (the first identifier) in the root-subnode and K5 is less than K6 (the second identifier) in the root-subnode. Therefore, based on the index relationship between the root-subnode and the leaf-subnode, it is determined that the data corresponding to K5 is located in the second set of data indicated by the leaf-subnode.

[0206] If the identifier of the first data is K6, the processing device 200 performs a binary search within the root - child node and determines that K6 is equal to K6 (the first identifier) in the root - child node. Therefore, based on the index relationship between the root - child node and the leaf - child node, it is determined that the data corresponding to K6 is located in the second group of data indicated by the leaf - child node.

[0207] It should be noted that the first group of data, the second group of data, etc. indicated by the leaf - child node are arranged in sequence according to the data decompression direction. In other examples of this application, they can also have other names, which are not limited in this application.

[0208] In this application, the processing device 200 queries the identifier indicated by the root - child node according to the identifier of the first data, and then determines the first identifier and / or the second identifier. Since there is an index relationship between the root - child node and the leaf - child node, the smaller area (i.e., the target storage area) where the first data is located in the storage area of the leaf - child node can be accurately located. Then, only the compressed data in this smaller area needs to be decompressed, without decompressing all the data, reducing the amount of data for decompressing the compressed data, thereby improving the query efficiency of the data.

[0209] S730. The processing device 200 determines the target storage area in the compressed data that matches the identifier of the first data.

[0210] In this embodiment, the target storage area is the storage area of the leaf - child node.

[0211] The first data or the address information of the first data is stored in the storage area of the leaf - child node.

[0212] The following takes the case where the first data is stored in the storage area of the leaf - child node as an example for illustration.

[0213] In a possible implementation manner, for the processing device 200 to determine the target storage area in the compressed data that matches the identifier of the first data, it includes: the processing device 200 determines the right boundary or the left and right boundaries corresponding to the Nth group of data where the first data is located from the metadata corresponding to the bottom - most node, so as to obtain the target storage area that matches the identifier of the first data.

[0214] In a possible scenario, the metadata including the right boundary or the left and right boundaries of each group of data is stored in the storage area of the bottom - most node of the compressed data. The right boundary or the left and right boundaries of each group of data are stored in sequence in the storage area of the bottom - most node according to the arrangement of each group of data. For example, the right boundary or the left and right boundaries of each group of data are arranged in the storage area of the bottom - most node as follows: the right boundary or the left and right boundaries of the first group of data, the right boundary or the left and right boundaries of the second group of data, the right boundary or the left and right boundaries of the third group of data are arranged in sequence.

[0215] Exemplarily, the processing device 200 obtains the Nth right boundary or left and right boundaries from the metadata according to the Nth group of data to which the first data belongs, so as to obtain the target storage area matching the identifier of the first data.

[0216] In another possible scenario, the storage area of the bottommost node of the compressed data stores the space occupied by each group of data, such as length, or the space occupied by each Key-Value.

[0217] Exemplarily, in the case where the storage area of the bottommost node of the compressed data stores the space occupied by each group of data. The processing device 200 determines that the length of the right boundary of the first data and the start address a of the storage area corresponding to the leaf-child node is N multiplied by the space occupied by each group of data, and then adds the foregoing start address a plus N multiplied by the space occupied by each group of data to obtain the right boundary of the first data.

[0218] The processing device 200 determines that the length of the left boundary of the first data and the start address a of the storage area corresponding to the leaf-child node is (N - 1) multiplied by the space occupied by each group of data, and then adds the foregoing start address a plus (N - 1) multiplied by the space occupied by each group of data to obtain the left boundary of the first data.

[0219] Exemplarily, in the case where the storage area of the bottommost node of the compressed data stores the space occupied by each Key-Value. The processing device 200 determines that the length of the right boundary of the first data and the start address a of the storage area corresponding to the leaf-child node is N multiplied by the number of Key-Values included in each group of data and then multiplied by the space occupied by each Key-Value, and then adds the foregoing start address a plus the length of the start address a which is N multiplied by the number of Key-Values included in each group of data and then multiplied by the space occupied by each Key-Value to obtain the right boundary of the first data.

[0220] The processing device 200 determines that the length of the left boundary of the first data and the start address a of the storage area corresponding to the leaf-child node is (N - 1) multiplied by the number of Key-Values included in each group of data and then multiplied by the space occupied by each Key-Value, and then adds the foregoing start address a plus the length of the start address a which is (N - 1) multiplied by the number of Key-Values included in each group of data and then multiplied by the space occupied by each Key-Value to obtain the left boundary of the first data.

[0221] Exemplarily, in the storage area of the bottom - most node of the compressed data, the space occupied by each Key - Value is stored, and the space occupied by each Key - Value is inconsistent. The processing device 200 determines that the length from the right boundary of the first data to the starting address a of the storage area corresponding to the leaf - child node is the sum of the spaces occupied by each Key - Value in the first N groups of data, and then adds the sum of the spaces occupied by each Key - Value in the first N groups of data to the aforementioned starting address a to obtain the right boundary of the first data.

[0222] The processing device 200 determines that the length from the left boundary of the first data to the starting address a of the storage area corresponding to the leaf - child node is the sum of the spaces occupied by each Key - Value in the first N - 1 groups of data, and then adds the sum of the spaces occupied by each Key - Value in the first N - 1 groups of data to the aforementioned starting address a to obtain the left boundary of the first data.

[0223] In a possible example, the metadata is stored at a fixed position in the storage area of the bottom - most node, such as at the starting part of the storage area.

[0224] For the content of the Nth group of data where the first data is located, reference can be made to the description of S920 above, which will not be elaborated here. For the content of the right boundary or the left and right boundaries, reference can be made to the embodiment where the storage area of the bottom - most node in the above - mentioned compressed data is also used to store the right boundary or the left and right boundaries, which will not be elaborated here.

[0225] S740. The processing device 200 decompresses the data in the target storage area to obtain the first data.

[0226] The target storage area, that is, the above - mentioned right boundary or the left and right boundaries, is the boundary before data compression.

[0227] Regarding the content that the processing device 200 decompresses the data in the target storage area to obtain the first data, the following provides two possible implementation manners.

[0228] In the first possible implementation manner, the processing device 200 starts decompressing the data in the leaf - child node from the starting address a of the storage area corresponding to the leaf - child node until the decompressed data reaches the above - mentioned right boundary. The processing device 200 queries the first data corresponding to the identifier of the first data from the decompressed data.

[0229] In the second possible implementation manner, the processing device 200 determines the left boundary of the data a corresponding to the left boundary of the first data from the correspondence between the boundary before data compression and the boundary after data compression, and then starts decompressing the compressed data from the left boundary of the data a until the decompressed data reaches the above - mentioned right boundary. The processing device 200 queries and obtains the first data from the decompressed data according to the identifier of the first data.

[0230] The correspondence between the boundaries of the above data before compression and after compression is the metadata stored in the bottommost node.

[0231] Regarding the above content, a possible example is provided below.

[0232] As Figure 6 shown, when the data decompression request received by the processing device 200 carries the identifier K4, query the compressed IOT, that is, query from the root node of the compressed IOT to the bottommost node, and determine that the data corresponding to K4 is located in the bottommost node a. Furthermore, the processing device 200 queries the root-child nodes (K3 / K6) in the bottommost node a according to K4, determines that K4 is between K3 and K6, and thus determines that the data corresponding to K4 is located in the second group of data (i.e., the second leaf node). The processing device 200 determines the space occupied by each Key-Value in the first two groups of data according to the space occupied by each Key-Value maintained in the bottommost node, so as to obtain the right boundary of the second group of data. The processing device 200 starts decompressing the data in the leaf-child node from the starting address a of the storage area corresponding to the leaf-child node until the decompressed data reaches the above right boundary, and then queries the Value corresponding to K4 in the decompressed data.

[0233] In a possible embodiment, if the address information of the first data is stored in the storage area of the above leaf-child node. That is, the data stored in the leaf-child node is Key-RID.

[0234] The processing device 200 decompresses the data in the target storage area to obtain the first data, including: the processing device 200 starts decompressing the data in the leaf-child node from the starting address a of the storage area corresponding to the leaf-child node until the decompressed data reaches the above right boundary. The processing device 200 queries the RID of the first data corresponding to the identifier of the first data from the decompressed data, and then reads the Value indicated by the RID.

[0235] Alternatively, the processing device 200 determines the left boundary of the data a corresponding to the left boundary of the first data from the correspondence between the boundaries of the data before compression and after compression, and then starts decompressing the compressed data from the left boundary of the data a until the decompressed data reaches the above right boundary. The processing device 200 queries the RID of the first data corresponding to the identifier of the first data from the decompressed data, and then reads the Value indicated by the RID.

[0236] In a possible embodiment, the data is stored in the form of Key-Value in the storage area of the above root-child node. Furthermore, after the processing device 200 queries that the data in the root-child node matches the identifier of the first data, the matching data can be directly returned.

[0237] It is understandable that, in order to implement the functions in the above embodiments, the processing device includes the corresponding hardware structures and / or software modules for executing each function. Those skilled in the art should easily realize that, in combination with the units and method steps of each example described in the embodiments disclosed in the present application, the present application can be implemented in the form of hardware or a combination of hardware and computer software. Whether a certain function is executed in the form of hardware or computer software driving hardware depends on the specific application scenarios and design constraints of the technical solution.

[0238] In the foregoing, in combination with Figures 1 to 6 , the data compression method provided according to this embodiment is described in detail. Next, in combination with Figure 8 , the data compression device provided according to this embodiment will be described.

[0239] Figure 8 FIG. 13 is a schematic structural diagram of a data compression device provided by the present application. This data compression device can be used to implement the functions of the processing device in the above method embodiments, and thus can also achieve the beneficial effects possessed by the above method embodiments. In this embodiment, this data compression device can be a module (such as a chip) of the processing device.

[0240] As Figure 8 shown, the data compression device 800 includes an acquisition module 810 and a compression module 820. The data compression device 800 is used to implement the functions in the method embodiments shown in the above Figures 1 to 6 .

[0241] The acquisition module 810 is used to acquire second data corresponding to the first data; the storage structure of the second data is a tree structure, and the bottom-layer nodes of the tree structure include root-child nodes and leaf-child nodes having an index relationship with the root-child nodes; the first data includes the data indicated by the root-child nodes and the leaf-child nodes.

[0242] The compression module 820 is used to compress the second data to obtain compressed data corresponding to the first data, and the compressed data includes the data indicated by the root-child nodes and the compressed data indicated by the leaf-child nodes.

[0243] It should be noted that, in other embodiments, the acquisition module 810 can be used to execute any step in the management method, and the compression module 820 can be used to execute any step in the management method. The steps to be implemented by the acquisition module 810 and the compression module 820 can be specified as needed, and all functions of the data compression device 800 are implemented by the acquisition module 810 and the compression module 820 respectively implementing different steps in the data compression method.

[0244] It should be noted that the processing device according to the embodiments of the present application may correspond to the data compression device 800 in the embodiments of the application, and may correspond to the execution of the method according to the embodiments of the present application Figures 1 to 6 the corresponding corresponding entity, and the operations and / or functions of each module in the data compression device 800 are respectively for realizing Figures 1 to 6 the corresponding processes of each method in the corresponding embodiments in, for the sake of brevity, will not be elaborated here

[0245] In the above text, in combination with Figure 7 , the data decompression method provided according to this embodiment is described in detail. Next, in combination with Figure 9 , the data decompression device provided according to this embodiment will be described

[0246] Figure 9 FIG. is a schematic structural diagram of a data decompression device provided by the present application. This data decompression device can be used to implement the functions of the processing device in the above method embodiments, and thus can also achieve the beneficial effects possessed by the above method embodiments. In this embodiment, this data decompression device may be a module (such as a chip) of the processing device

[0247] As Figure 9 shown, the data decompression device 900 includes an acquisition module 910, a query module 920, a determination module 930, and a decompression module 940. The data decompression device 900 is used to implement the functions in the method embodiments shown in the above Figure 7

[0248] The acquisition module 910 is used to acquire a data decompression request, and the data decompression request includes an identifier of the first data

[0249] The query module 920 is used to query compressed data according to the identifier of the first data; wherein, the storage structure of the compressed data is a tree structure, the bottom layer nodes of the tree structure include root-child nodes and leaf-child nodes having an index relationship with the root-child nodes, and the compressed data includes the data indicated by the root-child nodes and the compressed data indicated by the leaf-child nodes

[0250] The determination module 930 is used to determine, in the compressed data, a target storage area that matches the identifier of the first data, and the target storage area is the storage area of the leaf-child nodes

[0251] The decompression module 940 is used to decompress the data in the target storage area to obtain the first data

[0252] ​It should be noted that, in other embodiments, the obtaining module 910 may be used to execute any step in the management method, the querying module 920 may be used to execute any step in the management method, the determining module 930 may be used to execute any step in the management method, and the decompressing module 940 may be used to execute any step in the management method. The steps to be implemented by the obtaining module 910, the querying module 920, the determining module 930, and the decompressing module 940 can be specified as needed. By implementing different steps in the data decompression method through the obtaining module 910, the querying module 920, the determining module 930, and the decompressing module 940 respectively, all functions of the data decompression device 900 are realized.

[0253] It should be noted that the processing device according to the embodiment of the present application may correspond to the data decompression device 900 in the embodiment of the application, and may correspond to executing the method according to the embodiment of the present application Figure 7 The corresponding corresponding entity, and the operations and / or functions of each module in the data decompression device 900 are respectively for realizing Figure 7 the corresponding processes of each method in the corresponding embodiment. For the sake of brevity, they will not be elaborated here.

[0254] When the data compression device or the data decompression device implements the data compression method or the data decompression method shown in any of the foregoing figures through software, the data compression device and its respective units or the data decompression device and its respective units may also be software modules. The above data compression method or data decompression method is implemented by the processor calling the software module. The processor may be a CPU, implemented by an ASIC, or a programmable logic device (PLD). The above PLD may be a complex programmable logic device (CPLD), a field programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.

[0255] For a more detailed description of the above data compression device or data decompression device, reference may be made to the relevant descriptions in the embodiments shown in the foregoing figures, which will not be elaborated here. It can be understood that the data compression device or data decompression device shown in the foregoing figures is only an example provided in this embodiment. According to different data types or compression methods, the data compression device or data decompression device may include more or fewer units, which are not limited in this application.

[0256] When the data compression device or the data decompression device is implemented by hardware, the hardware can be implemented by a processor or a chip. The chip includes an interface circuit and a control circuit. The interface circuit is configured to receive data from other devices outside the processor and transmit it to the control circuit, or send the data from the control circuit to other devices outside the processor.

[0257] The control circuit and the interface circuit are configured to implement the method of any possible implementation manner in the above embodiments through logic circuits or execution of code instructions. The beneficial effects can be referred to the description of any aspect in the above embodiments, and will not be elaborated here.

[0258] It can be understood that the processor in the embodiments of the present application can be a CPU, an NPU or a GPU, and can also be other general-purpose processors, digital signal processors (DSP), ASICs, FPGAs or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. The general-purpose processor can be a microprocessor or any conventional processor.

[0259] In addition, Figure 8 the data compression device 800 shown, and Figure 9 the data decompression device 900 shown can also be implemented by a storage device, such as Figure 1 the storage device 120 shown in, or a data storage system including a storage device, etc.

[0260] The method steps in the embodiments of the present application can also be implemented by a processor executing software instructions. The software instructions can be composed of corresponding software modules, and the software modules can be stored in a RAM, flash memory, ROM, PROM, EPROM, electrically erasable programmable read-only memory (EEPROM), register, hard disk, removable hard disk, CD-ROM or any other form of storage medium well-known in the art. An exemplary storage medium is coupled to the processor so that the processor can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor. The processor and the storage medium can be located in an ASIC. In addition, the ASIC can be located in a communication device. Of course, the processor and the storage medium can also exist as discrete components in the communication device.

[0261] The present application also provides a chip system, which includes a processor for implementing the functions of the storage device in the above method. In a possible design, the chip system further includes a memory for storing program instructions and / or data. The chip system can be composed of chips or can include chips and other discrete devices.

[0262] An embodiment of the present application also provides a computer program product containing instructions. The computer program product may be software or a program product containing instructions that can run on a computing device or be stored in any available medium. When the computer program product runs on at least one computing device, at least one computing device is caused to execute the above compression method or decompression method.

[0263] An embodiment of the present application also provides a computer-readable storage medium. The computer-readable storage medium may be any available medium that a computing device can store or a data storage device such as a data center containing one or more available media. The available medium may be a magnetic medium (e.g., a floppy disk, a hard disk, a magnetic tape), an optical medium (e.g., a DVD), or a semiconductor medium (e.g., a solid-state drive). The computer-readable storage medium includes instructions that direct a computing device to execute the compression method or the decompression method.

[0264] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer programs or instructions. When the computer program or instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are executed in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, a network device, a user device, or other programmable devices. The computer program or instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer program or instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center in a wired or wireless manner. The computer-readable storage medium may be any available medium that a computer can access or a data storage device such as a server or data center integrating one or more available media. The available medium may be a magnetic medium, such as a floppy disk, a hard disk, a magnetic tape; it may also be an optical medium, such as a digital video disc (DVD); or it may be a semiconductor medium, such as a solid-state drive (SSD).

[0265] In various embodiments of the present application, if there is no special indication and logical conflict, the terms and / or descriptions between different embodiments are consistent and can be cross-referenced to each other. The technical features in different embodiments can be combined to form new embodiments according to their inherent logical relationships. The various numerical numbers involved in the embodiments of the present application are only for the convenience of description and are not used to limit the scope of the embodiments of the present application. The magnitude of the sequence numbers of the above processes does not mean the order of execution, and the order of execution of each process should be determined by its function and inherent logic.

Claims

1. A data compression method, characterized in that: The method comprises: Acquire second data corresponding to the first data; the storage structure of the second data is a tree structure, the bottom-level nodes of the tree structure include a root-child node and a leaf-child node having an index relationship with the root-child node; the first data includes the root-child node and the data indicated by the leaf-child node; compressing the second data to obtain compressed data corresponding to the first data, wherein the compressed data includes the data indicated by the root-child node and the compressed data indicated by the leaf-child node; Among them, the storage area of ​​the root-child node is used to store the first sub-data that meets the condition, and the storage area of ​​the leaf-child node is used to store the second sub-data that does not meet the condition, and the condition is used to indicate that the access index of the data is greater than or equal to a threshold, or, the identification number interval of two adjacent data indicated by the root-child node is K plus 1, and K is an integer greater than or equal to 1.

2. The method according to claim 1, characterized in that The access index includes: access frequency or access frequency; the first sub-data is used to indicate hot data, and the second sub-data is used to indicate cold data.

3. The method according to claim 1, characterized in that K satisfies the following formula, K(K+X)≥N; Among them, X is a constant, and N is the amount of data included in a bottom-level node.

4. The method according to any one of claims 1 to 3, characterized in that The storage area of ​​the bottom-level node in the compressed data is also used to store: the right boundary or the left and right boundaries of each group of data indicated by the leaf-child node in the storage area of ​​the leaf-child node; The left boundary of the first group of data is used to indicate the minimum byte length from the first address of the storage area corresponding to the leaf-child node to the location where the first group of data is stored, and the right boundary of the first group of data is used to indicate the maximum byte length from the first address of the storage area corresponding to the leaf-child node to the location where the first group of data is stored. The first group of data is any group of data indicated by the leaf-child node.

5. The method according to claim 4, characterized in that The data indicated by the root-child node indicates the right boundary or the left and right boundaries of each group of data.

6. The method according to any one of claims 1 to 3, characterized in that The data format of the first data is a key-value pair Key-Value, and the data indicated by the leaf-child node is stored in the storage area of ​​the leaf-child node in sequence according to the Key.

7. The method according to any one of claims 1 to 3, characterized in that The tree structure is a B+-tree.

8. A data decompression method, characterized in that: The method comprises: Obtaining a data decompression request, where the data decompression request includes an identifier of the first data; According to the identifier of the first data, the compressed data is queried; wherein the storage structure of the compressed data is a tree structure, the bottom-level nodes of the tree structure include a root-child node and a leaf-child node having an index relationship with the root-child node, and the compressed data includes the data indicated by the root-child node and the compressed data indicated by the leaf-child node; Determine a target storage area in the compressed data that matches the identifier of the first data, the target storage area being a storage area of ​​a leaf-child node; Decompressing the data in the target storage area to obtain the first data; Among them, the storage area of ​​the root-child node is used to store the first sub-data that meets the condition, and the storage area of ​​the leaf-child node is used to store the second sub-data that does not meet the condition, the first data includes the first sub-data and the second sub-data, and the condition is used to indicate that the access index of the data is greater than or equal to the threshold, or, the identification number interval of two adjacent data indicated by the root-child node is K plus 1, and K is an integer greater than or equal to 1.

9. The method according to claim 8, characterized in that The storage area of ​​the bottom-most node of the compressed data is also used to store: the right boundary or the left and right boundaries of each group of data indicated by the leaf-child node in the storage area of ​​the leaf-child node; The left boundary of the first group of data is used to indicate the minimum byte length from the first address of the storage area corresponding to the leaf-child node to the location where the first group of data is stored, and the right boundary of the first group of data is used to indicate the maximum byte length from the first address of the storage area corresponding to the leaf-child node to the location where the first group of data is stored. The first group of data is any group of data indicated by the leaf-child node.

10. The method according to claim 9, characterized in that The data indicated by the root-child node indicates the right boundary or the left and right boundaries of each group of data.

11. The method according to claim 10, characterized in that The determining, in the compressed data, a target storage area that matches the identifier of the first data includes: According to the identifier of the first data, query the bottom-level nodes in sequence from the root node of the compressed data; the first data or the address information of the first data is stored in the storage area of ​​the leaf-child node in the bottom-level node; The identifier of the first data is compared with the data in the root-child node to determine the right boundary or the left and right boundaries of the second group of data; the second group of data includes the first data, and the right boundary or the left and right boundaries of the second group of data are used to indicate the target storage area.

12. The method according to claim 11, characterized in that The step of comparing the identifier of the first data with the data in the root-child node to determine the right boundary or the left and right boundaries of the second set of data includes: Determine, among the identifiers of the data indicated by the root-child node, a first identifier and / or a second identifier that satisfies a first policy with the identifier of the first data; the first policy is used to indicate that: the identifier of the first data is less than the first identifier, or the identifier of the first data is greater than the second identifier, or the identifier of the first data is greater than or equal to the first identifier and less than or equal to the second identifier, and the first identifier is less than the second identifier; Determine, from the index relationship, a second group of data in a leaf-child node corresponding to the first identifier and / or the second identifier; The second set of data is obtained at the right boundary or the left boundary of the storage area of ​​the leaf-child node.

13. The method according to claim 8, characterized in that The access index includes: access frequency or access frequency; the first sub-data is used to indicate hot data, and the second sub-data is used to indicate cold data.

14. The method according to claim 8, characterized in that K satisfies the following formula, K(K+X)≥N; Among them, X is a constant, and N is the amount of data included in a bottom-level node.

15. The method according to any one of claims 8 to 14, characterized in that The data format of the first data is a key-value pair Key-Value, and the data indicated by the leaf-child node is stored in the storage area of ​​the leaf-child node in sequence according to the Key.

16. The method according to any one of claims 9 to 14, characterized in that The tree structure includes a B+-tree.

17. A data compression device, characterized in that: The device comprises: an acquisition module, configured to acquire second data corresponding to the first data; the storage structure of the second data is a tree structure, the bottom-level nodes of the tree structure include a root-child node and a leaf-child node having an index relationship with the root-child node; the first data includes the root-child node and the data indicated by the leaf-child node; A compression module, used for compressing the second data to obtain compressed data corresponding to the first data, wherein the compressed data includes the data indicated by the root-child node and the compressed data indicated by the leaf-child node; Among them, the storage area of ​​the root-child node is used to store the first sub-data that meets the condition, and the storage area of ​​the leaf-child node is used to store the second sub-data that does not meet the condition, and the condition is used to indicate that the access index of the data is greater than or equal to a threshold, or, the identification number interval of two adjacent data indicated by the root-child node is K plus 1, and K is an integer greater than or equal to 1.

18. The device according to claim 17, characterized in that The access index includes: access frequency or access frequency; the first sub-data is used to indicate hot data, and the second sub-data is used to indicate cold data.

19. The device according to claim 17, characterized in that K satisfies the following formula, K(K+X)≥N; Among them, X is a constant, and N is the amount of data included in a bottom-level node.

20. The device according to any one of claims 17 to 19, characterized in that The storage area of ​​the bottom-level node in the compressed data is also used to store: the right boundary or the left and right boundaries of each group of data indicated by the leaf-child node in the storage area of ​​the leaf-child node; The left boundary of the first group of data is used to indicate the minimum byte length from the first address of the storage area corresponding to the leaf-child node to the location where the first group of data is stored, and the right boundary of the first group of data is used to indicate the maximum byte length from the first address of the storage area corresponding to the leaf-child node to the location where the first group of data is stored. The first group of data is any group of data indicated by the leaf-child node.

21. The device according to claim 20, characterized in that The data indicated by the root-child node indicates the right boundary or the left and right boundaries of each group of data.

22. The device according to any one of claims 17 to 19, characterized in that The data format of the first data is a key-value pair Key-Value, and the data indicated by the leaf-child node is stored in the storage area of ​​the leaf-child node in sequence according to the Key.

23. The device according to any one of claims 17 to 19, characterized in that The tree structure is a B+-tree.

24. A data decompression device, characterized in that: The device comprises: An acquisition module, configured to acquire a data decompression request, wherein the data decompression request includes an identifier of the first data; A query module, configured to query compressed data according to the identifier of the first data; wherein the storage structure of the compressed data is a tree structure, the bottom-level nodes of the tree structure include a root-child node and a leaf-child node having an index relationship with the root-child node, and the compressed data includes data indicated by the root-child node and compressed data indicated by the leaf-child node; A determination module, configured to determine a target storage area in the compressed data that matches the identifier of the first data, wherein the target storage area is a storage area of ​​a leaf-child node; a decompression module, configured to decompress the data in the target storage area to obtain the first data; Among them, the storage area of ​​the root-child node is used to store the first sub-data that meets the condition, and the storage area of ​​the leaf-child node is used to store the second sub-data that does not meet the condition, the first data includes the first sub-data and the second sub-data, and the condition is used to indicate that the access index of the data is greater than or equal to the threshold, or, the identification number interval of two adjacent data indicated by the root-child node is K plus 1, and K is an integer greater than or equal to 1.

25. The device according to claim 24, characterized in that The storage area of ​​the bottom-most node of the compressed data is also used to store: the right boundary or the left and right boundaries of each group of data indicated by the leaf-child node in the storage area of ​​the leaf-child node; The left boundary of the first group of data is used to indicate the minimum byte length from the first address of the storage area corresponding to the leaf-child node to the location where the first group of data is stored, and the right boundary of the first group of data is used to indicate the maximum byte length from the first address of the storage area corresponding to the leaf-child node to the location where the first group of data is stored. The first group of data is any group of data indicated by the leaf-child node.

26. The device according to claim 25, characterized in that The data indicated by the root-child node indicates the right boundary or the left and right boundaries of each group of data.

27. The device according to claim 26, characterized in that The determining, in the compressed data, a target storage area that matches the identifier of the first data includes: According to the identifier of the first data, query the bottom-level nodes in sequence from the root node of the compressed data; the first data or the address information of the first data is stored in the storage area of ​​the leaf-child node in the bottom-level node; The identifier of the first data is compared with the data in the root-child node to determine the right boundary or the left and right boundaries of the second group of data; the second group of data includes the first data, and the right boundary or the left and right boundaries of the second group of data are used to indicate the target storage area.

28. The device according to claim 27, characterized in that The step of comparing the identifier of the first data with the data in the root-child node to determine the right boundary or the left and right boundaries of the second set of data includes: Determine, among the identifiers of the data indicated by the root-child node, a first identifier and / or a second identifier that satisfies a first policy with the identifier of the first data; the first policy is used to indicate that: the identifier of the first data is less than the first identifier, or the identifier of the first data is greater than the second identifier, or the identifier of the first data is greater than or equal to the first identifier and less than or equal to the second identifier, and the first identifier is less than the second identifier; Determine, from the index relationship, a second group of data in a leaf-child node corresponding to the first identifier and / or the second identifier; The second set of data is obtained at the right boundary or the left boundary of the storage area of ​​the leaf-child node.

29. The device according to claim 24, characterized in that The access index includes: access frequency or access frequency; the first sub-data is used to indicate hot data, and the second sub-data is used to indicate cold data.

30. The device according to claim 24, characterized in that K satisfies the following formula, K(K+X)≥N; Among them, X is a constant, and N is the amount of data included in a bottom-level node.

31. The device according to any one of claims 24 to 30, characterized in that The data format of the first data is a key-value pair Key-Value, and the data indicated by the leaf-child node is stored in the storage area of ​​the leaf-child node in sequence according to the Key.

32. The device according to any one of claims 24 to 30, characterized in that The tree structure includes a B+-tree.

33. A computer device, characterized in that: The method comprises a memory and a processor, wherein the memory is used to store computer instructions; when the processor executes the computer instructions, the method according to any one of claims 1 to 16 is implemented.

34. A computer-readable storage medium, characterized in that: The storage medium stores a computer program or instruction, and when the computer program or instruction is executed by a processing device, the method according to any one of claims 1 to 16 is implemented.

35. A computer program product comprising a computer program or instructions, characterized in that When the computer program or instruction is executed on a processing device, the method according to any one of claims 1 to 16 is implemented.

Citation Information

Patent Citations

  • A trie tree node compression method and device based on a double array

    CN109446198A

  • Computer product, information retrieval method, and information retrieval apparatus

    US20100131476A1