ZNS SSD-based B + tree index construction method for dynamic placement of cold and hot data
By adopting the B+ tree index construction method of dynamic placement of hot and cold data based on ZNS SSD in the storage system, the problems of low performance, insufficient resource utilization and low data management efficiency in existing storage technologies are solved, and efficient data storage and access are achieved.
Patent Information
- Application Number
- CN202510273029.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2045-03-10
AI Technical Summary
Existing storage technologies have many problems in performance, storage resource utilization and data management, and are difficult to meet the growing data storage and processing needs.
The B+ tree index construction method based on ZNS SSD is adopted to dynamically place hot and cold data. By storing the internal nodes of the B+ tree in DRAM, hot leaf nodes in PM, and cold leaf nodes in ZNS SSD, efficient and dynamic management of data is achieved.
It significantly improves random read and write performance, reduces latency, improves data access efficiency, realizes efficient utilization of storage resources, and balances storage costs and performance.
Smart Images

Figure CN120215825A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data storage, and particularly relates to a method for constructing a B+ tree index for dynamic placement of hot and cold data based on ZNS SSDs, which is applicable to constructing a high-performance and low-power storage system. Background Art
[0002] With the rapid development of information technology, the amount of data has grown exponentially, and various applications have put forward higher requirements for the performance, cost, and management efficiency of storage systems. The traditional storage architecture gradually reveals its deficiencies when dealing with these challenges.
[0003] In terms of storage performance, due to the limitations of the mechanical structure, the read and write speeds of traditional hard disks are relatively slow. Especially when facing random read and write operations, the high seek time and latency seriously affect the system's response speed. For example, in a database system, frequent random read and write operations are common scenarios. The performance bottleneck of traditional hard disks will lead to low efficiency of data query and writing, and cannot meet the requirements of real-time services. Although the emergence of solid-state drives (SSDs) has significantly improved storage performance, with the continuous increase in the amount of data and the continuous improvement of application requirements for read and write speeds, single SSD storage is also difficult to fully meet high-performance requirements. In addition, in a large-scale data storage environment, the overall read and write performance of the system not only depends on the storage device itself, but also is closely related to the organization and distribution method of data in the storage system. For the widely used B+ tree index structure, if all nodes are stored on traditional storage devices, when performing data retrieval and update, due to the large difference in access frequencies of different nodes, the overall performance cannot reach the optimal.
[0004] The utilization of storage resources is also a key issue. There are huge differences in cost and performance among different types of storage media. Dynamic random access memory (DRAM) and persistent memory (PM) have excellent read and write performance and can quickly respond to CPU access requests, but their costs are relatively high, and the cost of capacity expansion is also large. While ZNS SSD hard disks (zoned namespace solid-state drives) have relatively low costs and large storage capacities, their performance is somewhat inferior to that of DRAM and PM. In practical applications, if data is not reasonably stored according to its characteristics and all data is stored in high-performance DRAM or PM, it will cause a significant increase in storage costs; conversely, if all data is stored on low-performance storage media such as ZNS SSD hard disks, it cannot meet the performance requirements for rapid access to hot data, resulting in waste of storage resources and a decline in the overall system performance. Therefore, how to achieve efficient utilization of storage resources, balance storage costs and performance while ensuring system performance has become an urgent problem to be solved.
[0005] From a data management perspective, the access patterns of data are dynamically changing in practical applications. In many business scenarios, some data is frequently accessed within a certain period of time, while the access frequency of other data is relatively low. For example, in the order data of an e-commerce platform, recent order data is frequently queried and processed, belonging to hot data; while the order data from a long time ago has a relatively low access frequency and belongs to cold data. Traditional storage management methods often lack an effective response mechanism for the dynamically changing data access frequency and cannot automatically adjust the storage strategy according to the change in data heat. This leads to difficulties in meeting the fast access requirements of hot data and in reasonably utilizing storage resources to store cold data during the data management process. In addition, for the widely used B+ tree index data structure, under the conditions of dynamically changing data volume and continuously changing data access patterns, the traditional index maintenance method is inefficient, and during index update, insertion, and deletion operations, it is prone to causing imbalance in the index structure, thereby affecting the data access performance.
[0006] In summary, there are many problems in existing storage technologies in terms of performance, storage resource utilization, and data management. There is an urgent need for an innovative storage architecture and data management method to solve these problems in order to meet the growing data storage and processing requirements. Summary of the Invention
[0007] To solve the problems in the background technology, achieve optimized data storage and efficient access, and solve the problems of low read performance, large write overhead, and poor adaptability in the prior art, the present invention provides a method for constructing a B+ tree index based on dynamic placement of hot and cold data on ZNS SSDs, including: persistent memory PM, dynamic random access memory DRAM, and ZNS SSD hard disks;
[0008] Both the persistent memory PM and the dynamic random access memory DRAM are directly connected to the CPU memory bus and have the ability to be accessed by the CPU byte by byte;
[0009] The internal nodes of the B+ tree are stored in the dynamic random access memory DRAM, where the internal nodes only store the index information for locating the leaf nodes;
[0010] The pseudo-leaf nodes of the B+ tree are stored in the persistent memory PM, where the pseudo-leaf nodes only store the key index information of the leaf nodes;
[0011] The leaf nodes that actually store key-value pairs in the B+ tree are divided into hot leaf nodes and cold leaf nodes according to the data access frequency. Among them, the hot leaf nodes are stored in the persistent memory PM, and the cold leaf nodes are stored in the ZNS SSD hard disks.
[0012] The present invention has at least the following beneficial effects
[0013] In the present invention, the internal nodes of the B+ tree are stored in the dynamic random access memory (DRAM). Due to its fast read and write speed, it can quickly locate the leaf nodes, significantly reducing the time overhead during data query. The hot leaf nodes are stored in the persistent memory (PM), enabling frequent accessed data to be quickly read and written, effectively meeting the stringent requirements of real-time services for data read and write speed, such as the fast response to frequently queried and updated data in a database system. Traditional storage devices are restricted by mechanical structures during random read and write, resulting in high latency. However, this solution stores hot data in high-speed DRAM and PM, greatly reducing the latency during random read and write, effectively improving the random read and write performance, enabling the system to operate efficiently even in the face of a large number of random read and write operations, for example, performing well in big data analysis scenarios with frequent random read and write operations. Description of the Drawings
[0014] Figure 1 It is the overall architecture diagram of the DPZB+tree in the present invention;
[0015] Figure 2 It is the identification schematic diagram of the hot leaf node in the present invention;
[0016] Figure 3 It is the replacement schematic diagram of the hot leaf node in the present invention;
[0017] Figure 4 It is the node splitting schematic diagram in the present invention. Detailed Embodiments
[0018] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0019] Please refer to Figure 1 , the present invention provides a method for constructing a B+ tree index for dynamic placement of hot and cold data based on ZNS SSD, including: persistent memory PM, dynamic random access memory DRAM, and ZNS SSD hard disk;
[0020] Both the persistent memory PM and the dynamic random access memory DRAM are directly connected to the CPU memory bus and have the ability to be accessed by the CPU byte by byte;
[0021] The internal nodes of the B+ tree are stored in the dynamic random access memory DRAM. Among them, the internal nodes only store the index information for locating the leaf nodes;
[0022] The pseudo-leaf nodes of the B+ tree are stored in the persistent memory PM. Among them, the pseudo-leaf nodes only store the key index information of the leaf nodes;
[0023] The leaf nodes that actually store key-value pairs in the B+ tree are divided into hot leaf nodes and cold leaf nodes according to the data access frequency. Among them, the hot leaf nodes are stored in the persistent memory PM, and the cold leaf nodes are stored in the ZNS SSD hard disk.
[0024] In this embodiment, the overall system architecture of the B+ tree index with dynamic placement of hot and cold data based on the DRAM-PM-ZNS SSD hybrid storage architecture is as Figure 1 shown. In our design, the DPZB+tree is divided into three layers: the internal nodes stored in DRAM, the pseudo-leaf nodes (leaf node index information) stored in PM, and the hot leaf nodes stored in PM and the cold leaf nodes stored in ZNS SSD; in this storage architecture, the persistent memory is directly connected to the CPU memory bus and can be accessed by the CPU byte by byte; in a computer system, the ZNS SSD is generally connected to a storage controller (such as the storage controller integrated in the south bridge chip on the motherboard), and the storage controller then communicates with the CPU. The small-capacity persistent memory is mainly used to cache the key metadata and hot access data of the storage system. In this way, the characteristics of byte-addressable of the persistent memory can be effectively utilized to perform differential modification on the metadata, reduce write amplification, and at the same time use its in-place update characteristics to block the chain recursive update between the nodes of each layer of the B+ tree, while the ZNS SSD is used as the main medium to store most of the infrequently accessed data.
[0025] Three-layer structure of DPZB+tree
[0026] Internal node layer: The internal nodes are stored in DRAM. The relatively fast read and write speeds of DRAM make the efficiency extremely high when searching for leaf nodes. At the same time, the in-place update characteristic of DRAM can block cascade updates and significantly reduce the write amplification rate. The internal nodes only store the index information for locating the leaf nodes and play a key guiding role in the data search process. For example, in a B+ tree with a large amount of data, the internal nodes can quickly narrow down the search range to a specific leaf node branch, greatly improving the search efficiency.
[0027] Pseudo-leaf node layer: Only the key index information of the leaf nodes, that is, the pseudo-leaf nodes, is stored in PM. The advantage of this design is that when the system power is off, the entire tree can be reconstructed based on the relevant information of the head leaf nodes, and the usage of PM can be effectively controlled.
[0028] Hot and cold data storage layer: The leaf nodes that actually store key-value pairs are stored in PM and ZNS SSD respectively according to their hot and cold attributes. Hot nodes are stored in PM to prevent frequent small-granularity updates from causing write amplification and garbage collection problems on ZNS SSD, ensuring efficient reading and writing of hot data. Cold nodes are stored in ZNS SSD because of their lower requirements for write operation frequency, giving full play to the advantages of large-capacity storage of ZNS SSD while reducing storage costs.
[0029] During the data access process, there is close data interaction among the three storage hardware. When the CPU needs to access data, it first looks up the relevant internal node index information in DRAM to determine the range of leaf nodes where the data may be located. Then, based on the index information stored in PM for the pseudo-leaf nodes, it further determines whether the data is stored in the hot nodes in PM or the cold nodes in ZNS SSD. If it is hot node data, it is directly read from PM; if it is cold node data, it is read from ZNS SSD through the storage controller.
[0030] When writing data, hot data is also written to PM and cold data is written to ZNS SSD according to the hot and cold attributes of the data. At the same time, during data operations such as B+ tree structure adjustment (splitting, merging, etc.), DRAM, PM, and ZNS SSD also need to work together to ensure data consistency and integrity. For example, when a leaf node splits, it may be necessary to update the index information of the internal node in DRAM, update the relevant information of the pseudo-leaf node in PM, and store the new data in the corresponding PM or ZNS SSD. These three types of hardware in the DRAM-PM-ZNS SSD hybrid storage architecture achieve efficient data storage and access through different connection methods with the CPU, clear storage function division of labor, and close data interaction and cooperation, meeting the comprehensive requirements of modern computer systems for storage performance, capacity, and cost.
[0031] Preferably, the key index information of the leaf node includes: node minimum value, hot node flag bit, parent node pointer, number of keywords, left sibling node pointer, right sibling node pointer, pointer to persistent memory, pointer to ZNS SSD, and access timestamp; among them, the node minimum value is used to reconstruct the B+ tree structure; the hot node flag bit is used to indicate whether the leaf node is a hot leaf node. If it is marked as a hot leaf node, when accessing, the pointer to the persistent memory PM is used to query data from the persistent memory PM, otherwise the pointer to the ZNS SSD hard disk is used to query data from the ZNS SSD hard disk; the parent node pointer is used to record the position information of the leaf node in the parent node of the B+ tree. Through the parent node pointer, the structure of the B+ tree can be traced back and adjusted upward conveniently; the number of keywords is used to record the number of keywords contained in the leaf node; the left sibling node pointer and the right sibling node pointer are used to connect adjacent leaf nodes at the same level, so that the leaf nodes form an ordered linked list structure; the access timestamp is used to record the time when the leaf node was last read or written.
[0032] In this embodiment, the pseudo-leaf node is stored in the PM (persistent memory), mainly including the key index information of the leaf node, which specifically includes the following parts:
[0033] Node minimum value: It is used to reconstruct the tree structure based on this value when the system restarts normally, avoiding full disk scanning and improving the system recovery efficiency. Through the node minimum value, the approximate position and order relationship of the node in the tree can be quickly determined.
[0034] Hot node flag bit: This flag bit is used to identify whether the corresponding leaf node is a hot node. When the system accesses the leaf node, according to this flag bit, it is decided whether to use the pointer to the PM (if it is a hot node) or the pointer to the ZNS SSD (if it is a cold node), so as to realize the quick judgment and accurate access of the storage positions of hot and cold data.
[0035] Parent node pointer: Records the position of the parent node of the pseudo-leaf node, which helps to maintain the hierarchical structure of the B+ tree. When performing operations such as data insertion and deletion, it is convenient to trace back and adjust the tree structure upward.
[0036] Number of keywords: Records the number of keywords (keys in key-value pairs) contained in the leaf node, providing important information for data management and operations. For example, it is used when judging whether the node is full and whether split or merge operations are needed.
[0037] Left sibling node pointer, right sibling node pointer: Used to connect adjacent leaf nodes at the same level to form an ordered linked list structure among leaf nodes. This is very important for operations such as range queries. Through these pointers, adjacent leaf nodes can be quickly traversed to obtain data that meets the conditions.
[0038] Pointer to persistent memory: When the leaf node is a hot node, this pointer points to the actual key-value pair data stored in the PM for fast reading and writing of hot data.
[0039] Pointer to ZNS SSD: When the leaf node is a cold node, this pointer points to the actual key-value pair data stored in the ZNS SSD for accessing cold data stored on the large-capacity ZNS SSD.
[0040] Access timestamp: Records the time of the last write operation on this node. When the PM space is insufficient, the system replaces the hot node with the farthest write time (i.e., the least active) based on the access timestamp to ensure that the PM always caches the current hottest data. Read operations do not modify the access timestamp.
[0041] The hot and cold data storage layer is the leaf node layer that actually stores the key-value pairs, and is stored in the PM and ZNS SSD respectively according to the hot and cold attributes of the data:
[0042] Data stored in the PM: mainly hot node data, that is, data that is frequently accessed (frequently read or written). Storing hot data in the PM can utilize the fast read and write speeds and byte-addressing characteristics of the PM to quickly respond to data access requests and reduce latency. At the same time, it avoids the write amplification and excessive garbage collection problems that may be caused by frequent small-grained updates of hot data on the ZNS SSD, improving the overall performance of the storage system.
[0043] Data stored in the ZNS SSD: mainly cold node data, that is, data that is rarely accessed. The ZNS SSD has the characteristics of large capacity and relatively low cost, and is suitable for storing a large amount of infrequently accessed data. Storing cold data in the ZNS SSD gives full play to its large-capacity storage advantage and reduces the storage cost.
[0044] In summary, the pseudo-leaf node layer mainly stores the key information for indexing and managing the leaf nodes, while the hot and cold data storage layer stores the actual key-value pair data in the appropriate storage media according to the access frequency characteristics of the data.
[0045] Please refer to Figure 2, The write operations of B+ trees are quite random. Therefore, hot nodes exhibit timeliness and dynamics. If the currently written hot nodes are not identified and placed reasonably, it will cause serious write amplification and reduce the performance of the storage system. In read-write intensive applications, the request operations within a certain period often follow specific distribution characteristics, such as Zipfian distribution, that is, most operations are concentrated on a small part of the data. Therefore, the DPZB+tree proposed in the present invention adopts a clustering algorithm. Before the read-write operations are executed, the key value information key corresponding to all current read-write operation requests is clustered to find the key at the clustering center. At this time, the average distance from this key to the remaining keys is the shortest, that is, this key is the hot key of the current read-write sequence, and the node corresponding to this key will bear most of the read-write requests in the subsequent read-write operations. The specific operations are as follows:
[0046] Preferably, a hot and cold recognition strategy for read-write requests is set to identify the hot leaf nodes corresponding to the read-write request sequence. The hot and cold recognition strategy for read-write requests includes:
[0047] S101: Obtain the current M read-write request sequences. Among them, each request in the read-write request sequence contains the key value information of the key-value pair data to be processed;
[0048] S102: Based on the principle of temporal locality, select the key values of K requests from the M read-write request sequences at a step size of M / K as the initial central points;
[0049] S103: According to the principle of proximity, assign each request in the read-write request sequence to the nearest central point according to the key value difference to form K clusters;
[0050] d(k i , k j ) = |k i - k j |, i = 1, 2,..., M, j = 1, 2,..., K)
[0051] Among them, d(k i , k j ) represents the difference between the key value k i of the i-th request in the read-write request sequence and the key value k j of the j-th central point;
[0052] S104: For each cluster, calculate the new central point of each current cluster according to the formula , where G j represents the j-th cluster, and k mj represents the key value of the m-th request in the j-th cluster G j , represents the j-th cluster G jCalculate the average value of the key values of all requests;
[0053] S105: Continuously repeat steps S103 - S105 until the center point converges, then stop clustering; at this time, the key values of the K center points obtained are the key values of the hot leaf nodes of the current read / write request sequence.
[0054] In this embodiment, after finding the K hot keys of the current write through the above steps, the leaf nodes where these K hot keys are located are the hot nodes of the current write sequence. Before performing specific operations, it is necessary to ensure that these leaf nodes have been stored as hot nodes in the persistent memory; if the B+ tree is empty, directly create the first leaf node in the persistent memory and write the current request sequence into the persistent memory. If the tree is not empty, first find the hot keys through the clustering algorithm, and then find the leaf nodes corresponding to the hot keys by searching the B+ tree. After that, judge whether it has been stored as a hot node in the persistent memory through the hot node flag bit bool hot of the leaf node. If it is not, it needs to be migrated from the ZNS where cold data is placed to the persistent memory where hot data is placed. At this time, first judge whether there is space in the persistent memory for the leaf nodes that need to be migrated. If the space is insufficient, it is necessary to first execute the hot node replacement algorithm to select the node with the farthest access timestamp among the hot nodes and write it back to the ZNS SSD. By identifying the hot keys and reasonably storing the corresponding hot nodes in the persistent memory, the write amplification problem caused by hot nodes on the ZNS SSD is avoided, the processing efficiency of the storage system for frequently accessed data is improved, the high-speed read / write characteristics of the persistent memory are fully utilized, and the performance of the overall storage system is improved.
[0055] Please refer to Figure 3 , preferably, after identifying the hot leaf nodes, process the current read / write request sequence, and the process is as follows:
[0056] Find the corresponding leaf node in the B+ tree by searching the key values of the hot leaf nodes of the current read / write request sequence, and judge whether the leaf node is a cold node. If it is a cold node and there is no space in the persistent memory PM for storage, select the node with the farthest access timestamp among the hot leaf nodes and write it back to the ZNS SSD hard disk to free up space to write the leaf node into the persistent memory PM; if it is a cold node but there is storage space in the persistent memory PM, migrate the leaf node from the ZNS SSD hard disk to the persistent memory PM; if the leaf node is not a cold node, no operation is performed.
[0057] In this embodiment, since the small-capacity persistent memory is only used to cache leaf node index information and store some hot nodes, the number of nodes that can be cached simultaneously has an upper limit. When the number of cached nodes reaches the upper limit, the persistent cache area cannot accommodate new nodes, which will have a greater impact on the performance of the storage system. Therefore, we save the node access timestamp information in the pseudo-leaf nodes as the criterion for judging the hot and cold of the leaf nodes. In the initial state of the persistent cache area, there are no cached nodes, that is, the cached node set H is an empty set. Whenever a new cached node n is added to the set H or a read / write operation is performed on the cached node n, the access timestamp Ta of n is set to the current timestamp Tc. When the number of elements in H reaches the upper limit, the persistent cache area cannot accommodate new nodes. At this time, the node with the farthest access time from the current timestamp is calculated as the candidate cold node. Finally, the cold node is flushed back to the ZNS SSD, and its cache space is replaced by a new hot node. Through the above method, the reasonable organization of the B+ tree nodes stored in the persistent memory is ensured, and the node write amplification factor is reduced.
[0058] Specifically, for leaf nodes, we give K write hot requests through the K-means clustering algorithm, and the K nodes corresponding to these K hot requests will undertake most of the writes in the subsequent writes. Therefore, in order to reduce the cascading updates and garbage collection caused by frequent writes in the ZNS SSD, we need to dynamically place the non-hot nodes among the K nodes into the persistent memory as hot nodes. When the persistent memory space is insufficient, we need to formulate a node replacement strategy to replace the least active hot node among them. As Figure 3 shown, according to the algorithm, we get the hot key as 8. By searching, we find that the leaf node 2 where the hot key is located is currently stored as a cold node in the ZNS SSD. Therefore, we need to migrate node 2 as a whole to the persistent memory as a hot node before writing. At this time, the persistent memory is already full. Therefore, we need to traverse the dynamic array of hot node pointers to find the coldest leaf node among the hot nodes, that is, the leaf node node 1 with the farthest access timestamp from the current timestamp, copy node 2 to the ZNS SSD, modify the access timestamp to the current access timestamp, and at the same time copy node 1 to the ZNS SSD.
[0059] Preferably, the internal nodes are made to be an integer multiple of 64 bytes in size by adding an appropriate amount of reserved fields to ensure that they can fully adapt to the cache read / write granularity.
[0060] Preferably, the data structure sizes of the hot leaf nodes and the cold leaf nodes are uniformly adjusted to 4096 bytes, so that whether the leaf nodes are stored as hot nodes in the persistent memory or as cold nodes in the ZNS SSD, they can avoid the generation of space fragmentation and reduce additional read / write operations, improving the performance of the entire storage system.
[0061] In this embodiment, the structural design of the leaf nodes in the B+ tree is crucial, which directly affects query efficiency, space utilization, the complexity of insertion and deletion operations, the performance of range queries, and the optimization of memory and disk. A reasonable design can significantly improve the overall performance of the B+ tree. Therefore, in combination with the characteristics of each hardware, the present invention designs a B+ tree node applicable to the DRAM-PM-ZNS SSD storage architecture.
[0062] Considering that the read / write granularity of the cache is 64 bytes, if the size of the data structure is not an integer multiple of 64 bytes, free space fragmentation may occur in the cache, which will affect the utilization and efficiency of the cache. In addition, a data structure with a size that does not match the cache read / write granularity may require multiple read / write operations to complete, because a cache line (64 bytes) may contain parts of other data structures at the same time, which requires additional read / write operations, thus reducing the read / write efficiency. For this reason, we add reserved fields to the internal nodes stored in DRAM to align their sizes with the read / write granularity of the cache, thus effectively avoiding the generation of space fragmentation and improving the read / write efficiency. The data structure of the pseudo-leaf node is set to 64 bytes, and the read / write granularity of the cache is 64 bytes. Setting the data structure of the pseudo-leaf node to 64 bytes enables it to better adapt to the read / write granularity of the cache. When the data of the pseudo-leaf node needs to be read and written in the cache, since its size is exactly the same as the read / write granularity of the cache, it can be completely read and written by the cache, avoiding free space fragmentation in the cache and improving the utilization and read / write efficiency of the cache. At the same time, it also reduces the multiple read / write operations caused by the mismatch between the data size and the cache read / write granularity, reducing the system overhead.
[0063] At the same time, the read / write granularity of persistent memory is 256 bytes, and the read / write granularity of ZNS SSD is 4096 bytes. To optimize performance, we adjust the data structure size of the leaf nodes to 4096 bytes, so that whether they are stored as hot nodes in persistent memory or as cold nodes in ZNS SSD, space fragmentation and additional read / write operations can be avoided, thereby improving the performance of the entire storage system.
[0064] While optimizing the storage structure, how to efficiently handle the splitting and merging operations of nodes is also a key issue. The splitting and merging operations of the leaf nodes in the B+ tree play a crucial role in ensuring the balance of the tree, optimizing space utilization, improving query and write efficiency, supporting efficient range queries, reducing the number of disk I / Os, and enhancing system stability. A good splitting and merging strategy can significantly improve the overall performance of the storage system, especially in application scenarios that need to process large-scale data, frequent updates, and queries.
[0065] Please refer to Figure 4 , while optimizing the storage structure, how to efficiently handle the split and merge operations of nodes is also a key issue. The split and merge operations of the leaf nodes of the B+ tree play a crucial role in ensuring the balance of the tree, optimizing the storage space utilization rate, improving the query and write efficiency, supporting efficient range queries, reducing the number of disk I / Os, and enhancing the system stability. A good split and merge strategy can significantly improve the overall performance of the storage system, especially in application scenarios that need to process large-scale data, frequent updates, and queries.
[0066] Preferably, when node splitting occurs and the persistent memory PM has enough space to store the new hot leaf node, regardless of whether the leaf node to be split is a cold node or a hot node, the new node after splitting is stored on the persistent memory PM; when the storage space of the persistent memory is full and a hot node splits, find the hot node with the farthest access timestamp in the persistent memory, flush it back to the ZNS SSD to release the persistent memory space, and then store the new node after splitting on the persistent memory; when the storage space of the persistent memory is full and a cold node splits, find the two leaf nodes with the farthest and the second farthest access timestamps in the persistent memory, flush them back to the ZNS SSD, and after releasing enough space, store the new node after splitting on the persistent memory.
[0067] In this embodiment, for node splitting, the present invention takes into account the principle of spatial locality and ensures that regardless of whether the leaf node to be split is a cold node or a hot node, it should be stored as a hot node in the persistent memory after splitting. This is because node splitting usually means that a large number of write operations have occurred. Therefore, after splitting, all nodes should be stored as hot nodes in the persistent memory to maximize the support for data access. The specific process is as Figure 4 shown. In the initial situation, the node corresponding to the first piece of data we write will be stored as a hot node in the persistent memory. At the same time, when the persistent memory has enough space to store a new hot node and node splitting occurs, the new nodes after splitting should be stored on the persistent memory to improve the access efficiency. When the storage space of the persistent memory is full and node splitting occurs, for the splitting of a hot node, we only need to find a node with the farthest access timestamp and flush it back to the ZNS SSD to release the persistent memory space. For the splitting of a cold node, we need to find two nodes (the farthest and the second farthest access timestamps) and flush them back to the ZNS SSD.
[0068] Preferably, when the number of key-value pairs in a leaf node is less than N / 2 of its maximum storage capacity after a deletion operation, a node merging operation is triggered; if the leaf node has no left sibling node, it merges with the right sibling node; if the leaf node has no right sibling node, it merges with the left sibling node; if the leaf node has both a left sibling node and a right sibling node, and either the left sibling node or the right sibling node is a hot leaf node, the leaf node preferentially merges with the hot leaf node; if the leaf node has both a left sibling node and a right sibling node, and both the left sibling node and the right sibling node are hot leaf nodes or cold leaf nodes, the leaf node defaults to merging with the left sibling node.
[0069] In this embodiment, for the node merging problem, the core idea of the DPZB+tree is to maximize the utilization efficiency of persistent memory. Specifically, when the number of key-value pairs in a leaf node is less than N / 2 of its maximum storage capacity after a deletion operation, node merging is triggered. At this time, preference is given to merging with the hot node among the left and right sibling nodes to make full use of the hot nodes in persistent memory and improve the access efficiency and storage performance after merging.
[0070] In summary, the DPZB+tree significantly improves the performance based on ZNS SSDs through a series of optimization measures, reduces tail latency, and decreases write amplification. First, the DPZB+tree adopts a B+tree index structure, and under read-intensive workloads, its performance is superior to that of the LSM-tree. Since the leaf nodes of the DPZB+tree are ordered and connected by a linked list, the Scan operation can efficiently traverse the data in order, thus quickly completing data retrieval and processing. Second, the DPZB+tree avoids the way of the LSM-tree relying on merge operations to free up space and improve query speed, avoiding the performance overhead caused by multiple merges. Therefore, in terms of long-tail latency, the DPZB+tree is far superior to the LSM-tree. In addition, the DPZB+tree also identifies hot and cold data by adopting the K-means algorithm. Under sequential write workloads, the read and write speeds of the DPZB+tree are significantly better than those of other data structures such as the LSM-tree. The DPZB+tree also saves pseudo-leaf nodes through persistent memory, achieving fast recovery operations. When a system failure occurs, the DPZB+tree can complete the recovery process within milliseconds because it reduces the time and I / O overhead required for recovery by saving important data in persistent memory.
[0071] The DPZB+tree of the present invention adopts a DRAM-PM-ZNS SSD hybrid storage architecture. By saving the index information of leaf nodes in the persistent memory (PM) layer, it effectively solves the additional overhead brought by the cascading update of the B+tree. Secondly, DPZB+tree designs a low-overhead one-dimensional K-means algorithm to identify the hot request nodes currently being written. By storing the hot nodes in PM, it reduces the write overhead caused by frequent access to hot nodes, thereby improving the overall performance of the system. At the same time, aiming at the limitation of PM capacity, DPZB+tree designs a hot node replacement algorithm. By dynamically based on the access timestamp of leaf nodes, it reasonably replaces hot and cold nodes to optimize the utilization efficiency of PM. Finally, DPZB+tree optimizes the design of the leaf node structure and the splitting and merging algorithms, making full use of the characteristics of small-capacity PM, and effectively improving the performance of the storage system.
[0072] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the purpose and scope of the present technical solution, and they should all be covered within the scope of the claims of the present invention.
Claims
1. A B+ tree index construction method for dynamic placement of hot and cold data based on ZNS SSD, characterized in that: include: Persistent memory PM, dynamic random access memory DRAM and ZNS SSD hard drive; Both the persistent memory PM and the dynamic random access memory DRAM are directly connected to the CPU memory bus and can be accessed by the CPU byte by byte. The internal nodes of the B+ tree are stored in the dynamic random access memory DRAM, where the internal nodes only store index information for locating leaf nodes; The pseudo leaf nodes of the B+ tree are stored in the persistent memory PM, where the pseudo leaf nodes only store the key index information of the leaf nodes; The leaf nodes that actually store key-value pairs in the B+ tree are divided into hot leaf nodes and cold leaf nodes according to the frequency of data access. The hot leaf nodes are stored in the persistent memory PM, and the cold leaf nodes are stored in the ZNS SSD hard disk.
2. According to claim 1, a B+ tree index construction method for dynamic placement of hot and cold data based on ZNS SSD is characterized in that: The key index information of the leaf node includes: node minimum value, hot node flag, parent node pointer, keyword quantity, left sibling node pointer, right sibling node pointer, pointer to persistent memory, pointer to ZNS SSD and access timestamp; wherein, the node minimum value is used to reconstruct the B+ tree structure; the hot node flag is used to mark whether the leaf node is a hot leaf node, if it is marked as a hot leaf node, when accessing, the pointer to persistent memory PM is used to query data from persistent memory PM, otherwise, the pointer to ZNS SSD hard disk is used to query data from ZNS SSD hard disk; the parent node pointer is used to record the parent node position information of the leaf node in the B+ tree, and the structure of the B+ tree can be easily traced back and adjusted through the parent node pointer; the keyword quantity is used to record the number of keywords contained in the leaf node; the left sibling node pointer and the right sibling node pointer are used to connect adjacent leaf nodes at the same level, so that an ordered linked list structure is formed between the leaf nodes; the access timestamp is used to record the time when the leaf node was most recently read and written.
3. According to claim 1, a B+ tree index construction method for dynamic placement of hot and cold data based on ZNS SSD is characterized in that: Set a read / write request hot / cold identification strategy to identify hot leaf nodes corresponding to the read / write request sequence. The read / write request hot / cold identification strategy includes: S101: Acquire the current M read / write request sequences, wherein each request in the read / write request sequence contains key-value information of key-value pair data to be processed; S102: Based on the principle of temporal locality, select K requested key values from the M read / write request sequences with a step size of M / K as the initial center point; S103: According to the proximity principle, each request in the read / write request sequence is assigned to the nearest center point according to the key value difference, so as to form K clusters; d(k i ,k j )=|k i -k j |,i=1,2,…,M,j=1,2,…,K) Among them, d(k i , k j ) represents the key value k of the i-th request in the read / write request sequence i and the key value k of the jth center point j the gap; S104: For each cluster, according to the formula Calculate the new center point of each current cluster, where G j represents the jth cluster, k mj represents the jth cluster G j The key value of the mth request in Represents the jth cluster G j Calculate the average value of all requested key values; S105: Steps S103 to S105 are continuously repeated until the center points converge, and clustering is stopped; the key values of the K center points obtained at this time are the key values of the hot leaf nodes of the current read and write request sequence.
4. According to claim 3, a B+ tree index construction method for dynamic placement of hot and cold data based on ZNS SSD is characterized in that: After identifying the hot leaf node, the current read and write request sequence is processed as follows: By searching the key value of the hot leaf node of the current read and write request sequence, the corresponding leaf node is found in the B+ tree to determine whether the leaf node is a cold node. If it is a cold node and the persistent memory PM has no space to store it, the node with the furthest most recent access timestamp in the hot leaf node is selected and written back to the ZNS SSD hard disk to make room to write the leaf node to the persistent memory PM; if it is a cold node but the persistent memory PM has storage space, the leaf node is migrated from the ZNS SSD hard disk to the persistent memory PM; if the leaf node is not a cold node, no operation is performed.
5. According to claim 3, a B+ tree index construction method for dynamic placement of hot and cold data based on ZNS SSD is characterized in that: The internal node is made to have an integer multiple of 64 bytes by adding an appropriate amount of reserved fields, thereby ensuring that it can fully adapt to the cache read and write granularity.
6. According to claim 3, a B+ tree index construction method for dynamic placement of hot and cold data based on ZNS SSD is characterized in that: The data structure sizes of the hot leaf nodes and cold leaf nodes are uniformly adjusted to 4096 bytes, so that whether the leaf nodes are stored in persistent memory as hot nodes or in ZNS SSD as cold nodes, space fragmentation can be avoided and additional read and write operations can be reduced, thereby improving the performance of the entire storage system.
7. According to claim 3, a B+ tree index construction method for dynamic placement of hot and cold data based on ZNS SSD is characterized in that: When a node split occurs and the persistent memory PM is sufficient to store the new hot leaf node, the new node after the split is stored in the persistent memory PM regardless of whether the leaf node to be split is a cold node or a hot node; When the persistent memory storage space is full and the hot node splits, find the hot node with the farthest access timestamp in the persistent memory and flush it back to the ZNS SSD to free up the persistent memory space, and then store the new node after the split on the persistent memory; when the persistent memory storage space is full and the cold node splits, find the two leaf nodes with the farthest and second farthest access timestamps in the persistent memory and flush them back to the ZNS SSD. After freeing up enough space, store the new node after the split on the persistent memory.
8. According to claim 3, a B+ tree index construction method for dynamic placement of hot and cold data based on ZNS SSD is characterized in that: When the number of key-value pairs of a leaf node is less than N / 2 of its maximum storage capacity after the deletion operation, the node merge operation is triggered; If the leaf node has no left sibling node, it is merged with the right sibling node; if the leaf node has no right sibling node, it is merged with the left sibling node; if the leaf node has both a left sibling node and a right sibling node, and the left sibling node or the right sibling node is a hot leaf node, the leaf node is merged with the hot leaf node first; if the leaf node has both a left sibling node and a right sibling node, and both the left sibling node and the right sibling node are hot leaf nodes or cold leaf nodes, the leaf node is merged with the left sibling node by default.
Citation Information
Patent Citations
Read-write performance optimization method based on persistent memory B + tree index
CN115422182A
Method for reducing LSM write blocking and read-write amplification based on NVM and B + tree
CN115982169A
Key value storage method and system based on SSD-SMR hybrid storage
CN116880750A
SSD-HDD hybrid storage method for collaborative optimization of Prometheus performance and cost of time sequence database
CN118295591A
Separated memory system, hybrid indexing method, controller and distributed storage system
CN118568091A
Cited By
Simplified volume resource management method in storage system, electronic equipment, medium and product
CN120743804A
Thin volume resource management methods, electronic devices, media and products in storage systems
CN120743804B
Hybrid memory index management method and device, equipment and medium
CN120804104A
Anti-cache indexing method for mixed storage of nonvolatile memory and solid state disk
CN121501801A
A method for anti-caching index of hybrid storage of non-volatile memory and solid state disk
CN121501801B