System for realizing distributed index tree
By decoupling computing, memory, and storage resources in a distributed index tree to form a single index tree structure, the problems of dependent resource expansion and unbalanced load in existing technologies are solved, and independent resource expansion and efficient data operations are achieved.
Patent Information
- Application Number
- CN202410290594.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-14
- Publication Date
- 2025-09-16
AI Technical Summary
Existing distributed index tree clusters cannot achieve independent expansion of computing resources, memory resources, and storage resources, and there is a problem of load imbalance, especially when hot data is distributed in a single data partition, which can easily lead to overload of a single server.
Deploy the memory table of the distributed index tree in the distributed memory cluster of the memory layer, deploy the log files and data files in the distributed storage cluster of the storage layer, and deploy multiple computing nodes in the computing layer to decouple computing resources, memory resources, and storage resources, forming a single index tree structure and avoiding data partitioning.
It achieves independent expansion of a single resource, avoids load imbalance and single server overload problems, and improves resource utilization and data operation efficiency.
Smart Images

Figure CN120653642A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of database index technology, and in particular to a system for implementing a distributed index tree. Background Art
[0002] An index tree is a database index structure. In related technologies, some index trees (such as Log-Structured-Merge Tree (LSM Tree)) can be divided into two parts: memory storage and hard disk storage. The memory storage part is used to store the index tree's memory table (MemTable), and the disk storage part is used to store the index tree's log files and data files. When storing data based on an index tree, in order to increase capacity, a distributed index tree cluster is usually used to partition the data set. Each partition is managed by an index tree to manage the data within the partition. The capacity of the entire cluster increases with the scale of the cluster, and there is no upper limit on the capacity. However, in actual applications, the expansion granularity of this distributed index tree cluster is the server node. When expanding based on the server node, the computing resources, memory resources, and storage resources will increase simultaneously, resulting in the inability to achieve independent expansion of a single resource. In addition, if hot data is distributed in a single data partition, the problem of overloading a single server will also arise. Summary of the Invention
[0003] The present application provides a system for implementing a distributed index tree, which is used to solve the problem that the current distributed index tree cluster cannot achieve independent expansion of computing resources, memory resources and storage resources, and has load imbalance.
[0004] To solve the above technical problems, the present application provides a system for implementing a distributed index tree, including a computing layer, a memory layer, and a storage layer, wherein:
[0005] The memory layer includes a distributed memory cluster, and the distributed memory cluster is used to store a memory table of the index tree;
[0006] The storage layer includes a distributed storage cluster, and the distributed storage cluster is used to store log files and data files of the index tree;
[0007] The computing layer includes a distributed computing cluster, and computing nodes in the distributed computing cluster share the memory table, the log file, and the data file.
[0008] In an embodiment of the present application, since the memory table of the distributed index tree can be deployed in a distributed memory cluster of the memory layer, the log files and data files of the distributed index tree can be deployed in a distributed storage cluster of the storage layer, and multiple computing nodes (i.e., distributed computing clusters) can be deployed in the computing layer, the computing resources, memory resources, and storage resources of the distributed index tree can be decoupled, thereby enabling independent expansion of a single resource (computing resource, memory resource, or storage resource). In addition, since multiple computing nodes can share the memory table, log files, and data files of the distributed index tree, the distributed index tree is a single index tree. When data is stored based on the index tree, there is no need to partition the data, thereby avoiding the problem of load imbalance caused by frequent access to a single data partition when reading and writing data. BRIEF DESCRIPTION OF THE DRAWINGS
[0009] In order to more clearly illustrate the technical solutions in this application or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments recorded in this application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.
[0010] Figure 1 This is a schematic diagram of the structure of a distributed LSM Tree cluster in related technologies;
[0011] Figure 2 This is a schematic diagram of the structure of a system for implementing a distributed index tree according to an embodiment of the present application;
[0012] Figure 3 This is a schematic diagram of switching between three states of a memory table in an embodiment of the present application;
[0013] Figure 4 This is a schematic diagram of an embodiment of the present application performing a data write operation based on a distributed index tree;
[0014] Figure 5 This is a schematic diagram of persisting data in a memory table according to an embodiment of the present application;
[0015] Figure 6 This is a schematic diagram of an embodiment of the present application performing a data read operation based on a distributed index tree;
[0016] Figure 7 This is a schematic diagram of synchronizing metadata information of data files according to an embodiment of the present application. DETAILED DESCRIPTION
[0017] At present, for some index trees that include both memory storage and hard disk storage, when using distributed index tree clusters to partition and store data, there is a problem of being unable to independently expand a single resource (computing resources, memory resources or storage resources). In addition, when hot data is distributed in a single data partition, there will also be a problem of single server overload caused by frequent access to a single data partition.
[0018] Take the LSM Tree as an example. LSM Tree is a data index structure based on a hard disk. It has been widely used in databases, file systems, big data and other fields, such as RocksDB and LevelDB. However, RocksDB and LevelDB are both single-node services, and the LSM Tree used is a single LSM Tree. The index scale of a single LSM Tree is limited by the resources of a single server and has a relatively small capacity. In order to increase capacity, in related technologies, a distributed LSM Tree cluster can be used for data storage. When using a distributed LSM Tree cluster for data storage, the data set can be partitioned, and each partition is managed by an LSM Tree index structure. As the cluster scales, there is no capacity limit for the entire cluster, which can effectively increase capacity.
[0019] However, in actual applications, when a distributed LSM Tree cluster is expanded, the granularity is server nodes. Therefore, when a single resource in the cluster is expanded, other resources in the cluster are also increased simultaneously, resulting in a decrease in the utilization of these other resources. For example, a distributed LSM Tree cluster includes computing resources, memory resources, and storage resources. When computing resources are expanded at the granularity of server nodes, memory resources and storage resources are also increased simultaneously. However, memory resources and storage resources do not actually need to be expanded. Therefore, after expanding memory resources and storage resources, the utilization of memory resources and storage resources will decrease. Similarly, when the memory resources (or storage resources) of the cluster are expanded at the granularity of server nodes, the utilization of the computing resources and storage resources (or computing resources and memory resources) of the cluster will also decrease. Therefore, when a distributed LSM Tree cluster is expanded, it is impossible to independently scale a single resource. In addition, when storing data based on a distributed LSM Tree cluster, the data set is divided into multiple data partitions, and multiple data partitions are difficult to adjust dynamically. Therefore, when hot data is distributed in a single data partition, the data partition will be accessed frequently, which can easily lead to single server overload and load imbalance.
[0020] In related technologies, an improvement solution is to separate the storage and computing resources of the distributed LSM Tree cluster, that is, to separate the cluster's computing resources and memory resources from the cluster's storage resources. Figure 1 shown. Figure 1 The distributed LSMTree cluster in the LSMTree consists of 4 LSM Trees, each of which has independent computing resources (i.e. Figure 1 CPU resources shown) and memory resources (i.e. Figure 1 RAM resources shown), but can share storage resources (i.e. Figure 1 Storage resources shown). Figure 1 The distributed LSM Tree cluster shown decouples the storage resources of the cluster from the computing resources and memory resources of the cluster, which can achieve independent expansion of storage resources, but still cannot achieve independent expansion of computing resources or memory resources. In addition, due to Figure 1 The distributed LSM Tree cluster shown also manages data in partitions, so the problem of single server overload still exists.
[0021] An embodiment of the present application provides a system for implementing a distributed index tree. The system can deploy the memory table of the distributed index tree in a distributed memory cluster of the memory layer, deploy the log files and data files of the distributed index tree in a distributed storage cluster of the storage layer, and deploy multiple computing nodes (i.e., distributed computing clusters) in the computing layer. In this way, the computing resources, memory resources, and storage resources of the distributed index tree can be decoupled, thereby enabling independent expansion of a single resource (computing resources, memory resources, or storage resources). In addition, since multiple computing nodes can share the memory table, log files, and data files of the distributed index tree, the distributed index tree is a single index tree. When data is stored based on the index tree, there is no need to partition the data. Therefore, when reading and writing data, the problem of load imbalance caused by frequent access to a single data partition can be avoided.
[0022] In order to help those skilled in the art better understand the technical solutions of this application, the following will clearly and completely describe the technical solutions of this application in conjunction with the drawings of one or more embodiments of this application. Obviously, the described embodiments are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of this application.
[0023] The terms "first," "second," and the like in this application and the claims are used to distinguish similar objects and are not used to describe a particular order or precedence. It should be understood that such terms are interchangeable where appropriate so that this application can be implemented in sequences other than those illustrated or described herein. In addition, the term "and / or" in this application and the claims refers to at least one of the connected objects, and the character " / " generally indicates that the connected objects are in an "or" relationship.
[0024] It should be noted that the index tree in each embodiment of the present application can be an LSM tree or other index trees, such as a B+ tree, etc., which is not specifically limited here.
[0025] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.
[0026] Figure 2 This is a structural diagram of a system for implementing a distributed index tree according to an embodiment of the present application.
[0027] Figure 2 The system shown includes a computing layer 21, a memory layer 22, and a storage layer 23. The memory layer 22 includes a distributed memory cluster (which can also be expressed as a shared memory cluster). The distributed memory cluster is used to store the memory table of the index tree. The number of memory tables can be multiple, which can be determined based on actual needs and is not specifically limited here. The storage layer 23 includes a distributed storage cluster, which is used to store the log files and data files of the index tree. The computing layer 21 includes a distributed computing cluster. The computing nodes in the distributed computing cluster share the memory tables in the memory layer 22 and the log files and data files in the storage layer 23.
[0028] A distributed memory cluster is composed of multiple memory servers (also called memory nodes), and the memories of multiple memory servers constitute a distributed shared memory pool. The distributed memory cluster stores the memory table of the distributed index tree, specifically the distributed shared memory pool stores the memory table of the distributed index tree, that is, the memory table of the distributed index tree is deployed in the distributed shared memory pool. A memory table (MemTable) is a table that uses a memory engine, that is, a table placed in memory. A memory table is a data structure that can be used to save recently updated data. In an embodiment of the present application, when storing data, the memory table can organize the data in order according to the key. In some embodiments, the index form of the memory table can be a skip list or a B+ tree, etc., which is not specifically limited here.
[0029] The distributed storage cluster is a distributed file system that is responsible for storing the log files and data files of the distributed index tree. The log file is used to store the operation log of the data. When performing a write operation on the data, the operation log can be written to the log file first, and then the data can be written to the memory table. This can avoid the problem of being unable to perform fault recovery when the memory table loses data due to an abnormality. The data file is used to store persistent data. Specifically, when the data in the memory table is persistently stored, the data can be stored in the data file. The data file may include multiple layers of files, the data in the top layer file is the latest, and the data in the bottom layer file is the oldest. In some embodiments, when the index tree is an LSMTree, the data file may be a sorted string table (Sort String Table, SST) file. The SST file is a persistent KV (Key-Value) storage structure that can store data in Key sorting to achieve efficient retrieval. The data file is a read-only structure and generally cannot be changed after creation.
[0030] A distributed computing cluster consists of multiple compute nodes, each running an instance of an index tree service, similar to RocksDB and LevelDB. Multiple compute nodes share in-memory tables in the memory layer and log and data files in the storage layer. Each compute node can access in-memory tables in the memory layer through a key-value interface and log and data files in the storage layer through a file interface.
[0031] Based on the system provided by the embodiment of the present application, since the memory table of the distributed index tree can be deployed in the distributed memory cluster of the memory layer, the log files and data files of the distributed index tree can be deployed in the distributed storage cluster of the storage layer, and multiple computing nodes (i.e., distributed computing clusters) can be deployed in the computing layer, the computing resources, memory resources, and storage resources of the distributed index tree can be decoupled, thereby achieving independent expansion of a single resource (computing resource, memory resource, or storage resource). In addition, since multiple computing nodes can share the memory table, log files, and data files of the distributed index tree, the distributed index tree is a single index tree. When data is stored based on the index tree, there is no need to partition the data, so that when reading and writing data, the problem of load imbalance caused by frequent access to a single data partition can be avoided.
[0032] The memory table of the distributed index tree can include multiple states, and memory tables in different states can perform different functions. In the related art, when the memory table switches state, it involves the creation or deletion of the memory table (for example, after the memory table is full of data, a new memory table will be created, and the original memory table will be switched to a read-only state. After the data in the memory table is persisted, the memory table will be deleted), resulting in a large overhead. In an embodiment of the present application, in order to avoid frequent creation or deletion of memory tables, a method for multiple leaf layers to reuse index layers is proposed. Specifically, the distributed index tree can include a global index layer and multiple leaf layers. A leaf layer can correspond to a memory table. For any memory table, when the state of the memory table switches, the leaf layer corresponding to the memory table will switch state, but the global index layer remains unchanged. Multiple leaf layers can reuse the global index layer. That is to say, when the state of the memory table switches, the creation or deletion process of the memory table can be simplified to the state switching of the leaf layer, while the global index layer remains unchanged. The leaf layer that undergoes state switching can still reuse the global index layer. In this way, frequent creation or deletion of memory tables can be avoided, greatly reducing the update overhead of the index tree.
[0033] In some embodiments, the memory table of the distributed index tree may include three states (that is, the leaf layer of the distributed index tree can support three states), namely, the first state, the second state, and the third state. The first state can be represented as the active state. The memory table supports read and write operations in the first state and can store the latest inserted data. The second state can be represented as the immutable state. The memory table supports read operations and persistence operations in the second state, but does not support write operations, that is, the memory table in the second state is a read-only table, and the data in the memory table can be persisted to the data file of the distributed index tree. The third state can be represented as the flushed state. The memory table supports read operations and clear operations in the third state, but does not support write operations, that is, the memory table in the third state is a read-only table, and the data in the memory table can be cleared.
[0034] Based on the three states of the memory table described above, in practical applications, to ensure that the distributed index tree can receive new data (i.e., perform write operations on the data) at any time, at least one memory table must be in the first state at any time. To ensure that at least one memory table is in the first state at any time, the number of leaf layers in the distributed index tree must be at least two, meaning at least two leaf layers must be created. Of course, in other implementations, if the business scenario requires higher read and write performance, more leaf layers can be created.
[0035] The above three states of the memory table can be switched between each other, as shown below: Figure 3 shown.
[0036] Figure 3 In the example, if a memory table in the first state meets the conditions for switching to the second state, the memory table can switch from the first state to the second state. The conditions for switching to the second state can be, for example, that the data in the memory table in the first state reaches an upper threshold. If the data in the memory table in the second state is persisted, the memory table can switch from the second state to the third state. If the data in the memory table in the third state is cleared, the memory table can switch from the third state to the first state.
[0037] By switching the states between memory tables, it can be ensured that the memory tables can normally receive new data and persist the data to the data files.
[0038] Based on the system for implementing a distributed index tree provided by the embodiment of the present application, when the distributed index tree is actually applied, the system needs to be started first. After the system is successfully started, multiple computing nodes in the distributed computing cluster can perform write operations and read operations on data based on the distributed index tree. Among them, when starting the system, that is, in the system startup phase, the following three initialization phases (1) to (3) can be included:
[0039] (1) Multiple memory nodes in the distributed memory cluster register their local memory with the high-performance network. After multiple memory nodes complete the registration, the memory sharing service is started and the memory layer is initialized.
[0040] This phase is the initialization phase of the memory layer. During the initialization phase of the memory layer, the shared memory service can be started, and multiple memory nodes in the distributed memory cluster complete shared memory registration. The high-performance network can be, for example, a Remote Direct Memory Access (RDMA) network or a Compute Express Link (CXL) network, which are not specifically limited here. Shared memory can use, for example, PM, Field Programmable Gate Array (FPGA), Random Access Memory (RAM), and other types of storage, which are not specifically limited here.
[0041] (2) The first computing node in the distributed computing cluster creates a memory table and records the root node of the index tree in a fixed global memory address. The remaining computing nodes obtain the root node of the index tree from the fixed global memory address and initialize the memory cluster of the index tree.
[0042] This phase initializes the memory portion of the index tree. The first compute node can be the first compute node started in the distributed computing cluster. After starting, the first compute node can take over the management of the global shared memory space, create an in-memory index table (which can be in skip list or B+ tree format), and then record the root node of the index at a fixed global memory address. When other compute nodes start up, they can obtain the root node of the index tree from the fixed address, thereby completing the initialization of the index tree in memory.
[0043] (3) The computing nodes in the distributed computing cluster read the file information of the data file and initialize the storage cluster of the index tree based on the file information.
[0044] This phase is the initialization phase of the index tree storage part. During this phase, the computing nodes in the distributed computing cluster can continue to read the data file information in the underlying shared storage to complete the initialization of the storage part of the index tree, and then complete the initialization of the entire index tree.
[0045] After the system is successfully started, multiple computing nodes in the distributed computing cluster can perform write and read operations on data based on the distributed index tree. Taking a computing node in the distributed computing cluster as an example, when performing a write operation, the computing node may include the following steps:
[0046] Write write operation logs to the log file;
[0047] When the write operation log is successfully written to the log file, the data to be written is written to the memory table.
[0048] See Figure 4 When a computing node performs a write operation, it first performs a log write operation (i.e. Figure 4 Step 1) shown in the figure, after writing the log, execute the operation of writing the memory table (i.e. Figure 4 Step 2 shown). When writing a log, the computing node can write the write operation log into the log file of the distributed storage cluster through the file system interface. In some embodiments, in order to ensure sequential writing, the log file may only support serial writing, that is, when there are multiple log writing operations, the multiple log writing operations need to be queued and executed sequentially. When writing a memory table, the computing node can write the newly written data into the memory table in the distributed memory cluster through the remote KV interface. In some embodiments, in order to ensure sequential writing, at any time, only one computing node can be allowed to write the data to be written to the memory table. In order to ensure that only one computing node writes data to the memory table at any time, a global lock mechanism can be used to resolve write conflicts in multi-node concurrent scenarios, or a one-write-multiple-read mode can be selected according to demand.
[0049] When a computing node writes data to be written to a memory table, it may specifically write the data to the memory table in a first state. After writing the data to the memory table in the first state, it may determine whether the memory table meets the conditions for switching to the second state. This condition may be, for example, that the data in the memory table reaches an upper threshold. If the memory table meets the conditions for switching to the second state, the computing node may switch the memory table from the first state to the second state, that is, switch the leaf layer corresponding to the memory table from the first state to the second state.
[0050] When the computing node switches the memory table from the first state to the second state, it can further determine whether the memory table meets the persistence condition. The persistence condition can be determined according to the actual scenario and is not specifically limited here. If the memory table meets the persistence condition, the computing node can persist the data in the memory table to the data file in the distributed storage cluster. For details, please refer to Figure 5 .
[0051] Figure 5 The process of persisting the data in the memory table to the data file shown can be expressed as a Minor merge process. After the computing node switches the memory table from the first state to the second state, if the memory table meets the persistence conditions, the data in the memory table can be persisted to the data file. When persisting, the data in the memory table and the data in the first-layer data file (the data file includes multiple layers, and the data in the first-layer data file is the latest) can be read into the merge cache for merging, and the merge result is written to the first-layer data file. At the same time, the memory table is switched from the second state to the third state. In this way, the persistence of the data in the memory table can be achieved.
[0052] After persisting the data in the memory table to the first-layer data file, the computing node can further determine whether the first-layer data file meets the merge condition. The merge condition can be, for example, that the data in the first-layer data file reaches an upper threshold. If the first-layer data file meets the merge condition, the computing node can read the data in the first-layer data file and the data in the second-layer data file into the merge cache for merging, and write the merge result to the second-layer data file. After writing the merge result to the second-layer data file, the computing node can further determine whether the second-layer data file meets the merge condition. The merge condition can be, for example, that the data in the second-layer data file reaches an upper threshold. If the second-layer data file meets the merge condition, the computing node can read the data in the second-layer data file and the data in the third-layer data file into the merge cache for merging, and write the merge result to the third-layer data file. By analogy, the computing node can implement the merging of multiple layers of data files. The merging of data files can be expressed as a major merge.
[0053] After the computing node persists the data in the memory table and switches the memory table from the second state to the third state, it can immediately clear the data in the memory table, or it can adopt a delayed clearing strategy to select an appropriate clearing time, that is, to determine whether the memory table meets the clearing conditions, and clear the data in the memory table if the clearing conditions are met. The clearing condition can be, for example, that the number of memory tables in the first state is less than the lower limit threshold. At this time, in order to ensure that new data can be received, the data in the memory table can be cleared. After clearing the data in the memory table, the computing node can switch the memory table from the third state to the first state. After switching the memory table from the third state to the first state, the computing node can write new data to the memory table when performing a write operation on the data.
[0054] It should be noted that when performing a read operation on data, the memory table will be queried first. Only when the target data is not found in the memory table will the data file be queried (please refer to the subsequent read operation process for details), and the efficiency of data query in the memory table is higher than the efficiency of data query in the data file. Therefore, when clearing the memory table, it is preferable to adopt a delayed clearing strategy, that is, the memory table is cleared only when necessary and the memory table is switched to the first state. In this way, not only can write operation blocking be avoided, but also query efficiency can be improved (for example, the target data happens to be in the memory table in the third state. If the memory table is not cleared, the target data can be queried in the memory table, and the query efficiency is high. If the memory table is cleared, the target data needs to be queried from the data file, and the query efficiency is low. Therefore, adopting a delayed clearing strategy can improve query efficiency).
[0055] When a compute node performs a read operation, it can include the following steps:
[0056] Query the memory table of the first state, the memory table of the second state, and the memory table of the third state in sequence according to the key;
[0057] When the value corresponding to the key is found, the value is returned;
[0058] If the value corresponding to the key is not found, query the data file based on the key.
[0059] See Figure 6 When a computing node performs a read operation, it first performs a memory read operation (i.e. Figure 6 If the result is a hit, the value (i.e. target data) is returned. If the result is not a hit, the operation of reading the data file (i.e. Figure 6Step 2 shown). When reading memory, the computing node can query whether there is a value corresponding to the key in the memory table through the KV interface. When searching the memory table, it can first search the memory table in the first state, then search the memory table in the second state, and finally search the memory table in the third state until the value corresponding to the key is found (because the data in the memory table in the first state is the latest, followed by the data in the memory table in the second state, and finally the data in the memory table in the third state). When reading data files, the computing node can read the data file layer by layer through the file interface until the value corresponding to the key is found. Among them, when the computing node reads the data file layer by layer, it can use the traditional index tree optimization read operation method to optimize the read operation to improve query efficiency.
[0060] It should be noted that when executing a read operation, multiple computing nodes can execute the read operation in parallel, that is, the distributed index tree supports the parallel execution of multiple read operations.
[0061] In some implementations, to improve write performance, multiple computing nodes in a distributed computing cluster can be deployed in a one-write, multiple-read mode. In this mode, the write and read nodes within the multiple computing nodes need to synchronize metadata about the data files in a timely manner to prevent read operations from accessing illegal files. Specifically, before performing a read operation on a data file, a computing node can perform the following operations:
[0062] Acquire first metadata information of the data file from the memory layer or the storage layer;
[0063] Determining whether the first metadata information is consistent with the locally stored second metadata information;
[0064] In case of inconsistency, the second metadata information is synchronized according to the first metadata information, and after the synchronization is completed, a read operation is performed on the data file.
[0065] See Figure 7. In order to synchronize the metadata information of data files and avoid read operations accessing illegal files, two methods can be used. The first method is to synchronize metadata through shared memory in the memory layer. Specifically, a separate space can be opened up in the global shared memory, and the write node can write the first metadata information of the data file (such as the update log, version information, etc. of the data file) in this space. Before reading the data file, the read node can obtain the first metadata information from the memory layer. The first metadata information is the global version of the metadata information, and then determine whether the local metadata information (i.e., the second metadata information) is consistent with the first metadata information. If inconsistent, the local metadata information is synchronized according to the first metadata information. After the synchronization is completed, the read operation of the data file can be performed. If consistent, there is no need to synchronize the metadata, and the read operation can be performed directly on the data file. The second method is to synchronize metadata through the storage layer. Specifically, the write node can record the first metadata information of the data file (such as the update log, version information, etc. of the data file) in a fixed file of the storage layer (such as the manifest file). Before the read node reads the data file, it can first read the first metadata information from the fixed file, and then determine whether the local metadata information (i.e., the second metadata information) is consistent with the first metadata information. If inconsistent, the local metadata information is synchronized according to the first metadata information. After the synchronization is completed, the read operation of the data file can be performed. If consistent, there is no need to synchronize the metadata, and the read operation can be performed directly on the data file.
[0066] In summary, the key innovation of the technical solution of the embodiment of the present application is that the memory table of the index tree is extended from the computing node host memory to the distributed shared memory pool space, and memory semantic interfaces such as unified addressing and dynamic application release are implemented. The memory interface is used to implement the insertion and deletion functions of the memory table. The memory table supports multiple index forms such as skip lists and B+ trees. The high-performance network is used to provide remote read and write services to the outside world, and KV read and write capabilities are provided to the computing layer, and concurrent read and write capabilities of multiple requests are supported. The embodiment of the present application uses memory decoupling technology to construct a distributed shared memory cluster that supports elastic expansion. The memory table of the index tree is extended to the distributed shared memory space through a high-performance network, and an efficient remote KV read and write interface is provided to the outside world. The computing node implements the initialization, reading, writing, and merging operations of the memory table through the KV interface, and persists the memory table data to the data files of the distributed storage cluster as needed. In a specific implementation, the distributed computing cluster can be configured with a small amount of memory resources, the distributed memory cluster can be configured with a large amount of RAM memory and a small amount of computing power, and the memory nodes in the distributed memory cluster can register local memory to high-performance networks such as RDMA to form a large-capacity distributed shared memory pool. Compute nodes remotely access shared memory through RDMA read and write instructions, and access distributed storage clusters through high-speed networks or Ethernet as needed.
[0067] The embodiment of the present application implements a distributed index tree structure based on a distributed resource cluster with three layers of decoupling: computing, memory, and storage. The memory layer and the storage layer jointly store the entire index tree structure, and the computing node implements the reading, writing, and management functions of the index tree through a high-performance KV interface. The index tree based on distributed shared memory supports elastic expansion of computing, memory, and storage resources, and can independently expand any resource, which not only solves the problem of single-node capacity limitation, but also does not lead to a decrease in resource utilization. In addition, the distributed index tree of the embodiment of the present application provides a single index tree structure to the outside world, and does not require data partitioning, thereby avoiding the problem of server load imbalance caused by hot data.
[0068] The foregoing description describes specific embodiments of the present application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims or specification may be performed in an order different from that described in the embodiments and still achieve the desired results. Furthermore, the processes depicted in the accompanying drawings do not necessarily require the specific order shown or the sequential order to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0069] In short, the above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application shall be included in the scope of protection of the present application.
[0070] The systems, devices, modules, or units described in the above embodiments may be implemented by computer chips or entities, or by products having certain functions. A typical implementation device is a computer. Specifically, the computer may be, for example, a personal computer, a laptop computer, a cellular phone, a camera phone, a smartphone, a personal digital assistant, a media player, a navigation device, an email device, a game console, a tablet computer, a wearable device, or a combination of any of these devices.
[0071] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be implemented by any method or technology to store information. The information can be computer-readable instructions, data structures, program modules or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technology, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic tape, magnetic disk storage or other magnetic storage devices or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media (transitory media), such as modulated data signals and carrier waves.
[0072] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.
[0073] The various embodiments in this application are described in a progressive manner. Similar parts between the various embodiments can be referred to in conjunction with each other. Each embodiment focuses on the differences between the other embodiments. In particular, the system embodiment is generally similar to the method embodiment, so the description is relatively simple. For relevant parts, refer to the partial description of the method embodiment.
Claims
1. A system for implementing a distributed index tree, comprising a computing layer, a memory layer, and a storage layer, wherein: The memory layer includes a distributed memory cluster, and the distributed memory cluster is used to store a memory table of the index tree; The storage layer includes a distributed storage cluster, and the distributed storage cluster is used to store log files and data files of the index tree; The computing layer includes a distributed computing cluster, and computing nodes in the distributed computing cluster share the memory table, the log file, and the data file.
2. The system according to claim 1, wherein the index tree comprises a global index layer and multiple leaf layers, and each leaf layer corresponds to one of the memory tables; When the state of the memory table switches, the leaf layer switches its state, the index layer remains unchanged, and the multiple leaf layers reuse the index layer.
3. The system of claim 2, wherein the memory table includes a first state, a second state, and a third state. The memory table supports read and write operations in the first state, supports read operations and persist operations in the second state, and supports read operations and clear operations in the third state. in, At any time, there is at least one memory table in the first state.
4. The system of claim 3, wherein when the memory table in the first state meets a condition for switching to the second state, the memory table switches from the first state to the second state; When the data in the memory table in the second state is persisted, the memory table is switched from the second state to the third state; When the data in the memory table in the third state is cleared, the memory table is switched from the third state to the first state. 5 . The system according to claim 1 , wherein the computing node accesses the memory table through a KV interface and accesses the log file and the data file through a file interface.
6. The system according to any one of claims 1 to 5, wherein during the startup phase of the system: The multiple memory nodes in the distributed memory cluster register local memories with the high-performance network, and when the multiple memory nodes complete the registration, start the memory sharing service and initialize the memory layer; The first computing node in the distributed computing cluster creates the memory table and records the root node of the index tree in a fixed global memory address, and the remaining computing nodes obtain the root node of the index tree from the fixed global memory address and initialize the memory cluster of the index tree; The computing nodes in the distributed computing cluster read the file information of the data file and initialize the storage cluster of the index tree according to the file information.
7. The system according to claim 6, wherein when the computing node performs a write operation: Writing the write operation log into the log file; When the write operation log is successfully written into the log file, writing the data to be written into the memory table; in, The log file supports serial writing, and at any time, one computing node is allowed to write the data to be written into the memory table.
8. The system according to claim 7, wherein the computing node writes the data to be written into the memory table, comprising: The computing node writes the data to be written into the memory table in the first state; Wherein, when the memory table in the first state meets the condition for switching to the second state, the computing node switches the memory table from the first state to the second state.
9. The system according to claim 8, wherein after the computing node switches the memory table from the first state to the second state, the computing node determines whether the memory table meets the persistence condition; When determining that the memory table meets the persistence condition, the computing node reads the data in the memory table and the data in the first-layer data file into a merge cache for merging, writes the merged result into the first-layer data file, and switches the memory table from the second state to the third state.
10. The system according to claim 9, wherein after the computing node switches the memory table from the second state to the third state, the computing node determines whether the memory table satisfies a clearing condition; When determining that the memory table meets the clearing condition, the computing node clears the data in the memory table and switches the memory table from the third state to the first state.
11. The system according to claim 6, wherein when the computing node performs a read operation: Query the memory table of the first state, the memory table of the second state, and the memory table of the third state in sequence according to the key; When the value corresponding to the key is found, the value is returned; If the value is not found, the data file is queried according to the key.
12. The system according to claim 11, wherein before the computing node performs a read operation on the data file: Acquire first metadata information of the data file from the memory layer or the storage layer; Determining whether the first metadata information is consistent with the locally stored second metadata information; In case of inconsistency, the second metadata information is synchronized according to the first metadata information, and after the synchronization is completed, a read operation is performed on the data file.