NVM-based read optimization method for LSM hierarchical cold and hot partition index

By introducing NVM hierarchical hot and cold partition indexes into the LSM tree, the query path and data update mechanism of the LSM tree are optimized, solving the problems of read latency and write amplification in the LSM tree structure, and achieving faster read request responses and precise hot data management.

CN120723809APending Publication Date: 2025-09-30QINGHAI NORMAL UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510635359.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-16
Publication Date
2025-09-30

AI Technical Summary

Technical Problem

The existing LSM tree structure has problems with reads, such as long query paths, increased read latency, and duplicate writes when updating hot and cold data. Especially in big data storage environments, the write amplification effect caused by query latency and hot and cold data migration is severe.

Method used

An NVM-based LSM hierarchical hot and cold partition indexing method is adopted. The L0 to L2 layers of the LSM tree are placed in NVM, and hot and cold partition indexes are added. The query path is optimized through the hash skip table and LRU access time list, and the hot and cold index areas are dynamically adjusted to reduce invalid queries and write amplification.

Benefits of technology

It achieves faster read request response time, reduces write amplification of hot and cold data migration, and accurately filters hot data, improving query efficiency and the granularity of data updates.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120723809A_ABST
    Figure CN120723809A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of cold and hot partition indexing, and particularly discloses an NVM-based read optimization method for LSM hierarchical cold and hot partition indexing, an LSM tree hierarchical structure is divided into cold and hot index regions, the hot index regions can be queried firstly during query, and the cold index regions are searched under the condition of no hit. And L0-L2 are placed in an NVM layer, and data are stored by using a skip list, so that the efficiency of point request and range query is improved. According to the method, each PMTable is subjected to reading access time recording, each record represents one reading heat degree, and the judgment standard of each time of cold and hot data is the access time record in the latest period of time; the NVM has persistence capability, the original LSM tree is persisted in a disk or an SSD, and the L0-L2 layers of the LSM tree are put into the NVM, so that data loss caused by power failure and the like does not need to be worried about.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hot and cold partition indexes, and in particular to a read optimization method for LSM hierarchical hot and cold partition indexes based on NVM. Background Art

[0002] With the increasing number of e-commerce users, data volumes are rapidly increasing annually. LSM-tree-based data structures are widely used in commercial big data storage due to their superior write performance. For example, Google's LevelDB, Metadata's RocksDB, and Apache Cassandra all use LSM trees as their underlying data storage engine. LSM trees use a skip list data structure to store data in memory. When the storage size reaches a threshold, a batch of data is written to disk for persistence. The LSM tree's memory structure ensures that the written data is always ordered. However, because LSM trees use an append-only write mechanism at the disk's L0 layer, data in the L0 layer may be duplicated. LSM trees use a compaction strategy to merge and compress data from the Ln layer on disk and write it to the Ln+1 layer. When the Ln layer of the LSM tree reaches a threshold, the compaction strategy selects SSTable files from the Ln layer and any overlapping files in the Ln+1 layer for compaction and writes them to the Ln+1 layer. Deduplication and sorting are performed during this process. Therefore, all underlying data, except for the L0 layer, is organized as SSTables, effectively ensuring query efficiency.

[0003] However, existing LSM tree structures have several issues with read operations. First, due to the hierarchical structure of disks, lower levels have larger capacities and can store more data. This means that each query must go from memory to disk L0, then to L1, L2, and finally Ln. This results in a long query path and increases read latency. Second, due to the principle of locality, frequently accessed data may be accessed frequently within a short period of time. Repeated queries across multiple components can lead to unnecessary query delays due to invalid data searches. Finally, some existing hot and cold partitioning schemes divide the data into hot and cold data, which results in unnecessary duplicate writes each time the hot and cold data are updated. Summary of the Invention

[0004] To solve the problems existing in the prior art, the present invention provides a read optimization method for LSM hierarchical hot and cold partition indexes based on NVM, which solves the problems mentioned in the above background technology.

[0005] To achieve the above objectives, the present invention provides the following technical solution: a read optimization method for LSM hierarchical hot and cold partition indexes based on NVM, comprising the following steps:

[0006] S1. Initialize the storage system;

[0007] S2: First, write the WAL log file to persist the data. Then, compare the new data with the cached data, update the cached data, hash it to the target bucket based on the first few digits of the write key, and then update or insert it based on the skip table in the bucket. Then, start storing a batch of data in the in-memory Memtable, convert the Memtable to an Immutable Memtable, and then directly use the functions provided by PMDK to copy the Immutable Memtable's skip table to the L0 layer in NVM.

[0008] S3: When the L0 layer reaches two PMtables, the L0 data is compressed and merged into the L1 area;

[0009] S4. Build an index data structure for each PMtable of L1 layer data;

[0010] S5. Define and maintain a global heat record table;

[0011] S6. When the L1 layer data is full, the L1 layer data is compressed and merged into the L2 area;

[0012] S7: When the L2 layer is full, the data is compressed into the SSD using the same compression strategy as the L1 layer.

[0013] S8, query cache: Routes to a specified bucket based on the hash value of the first few digits of the query key, and then performs a query in the bucket;

[0014] S9: If the cache misses, the query goes to the MemTable and Immutable Memtable in memory. If there is no hit, the query goes to the L0 layer of the LSM tree and searches all the data in L0 for the latest data.

[0015] S10. When the L0 layer data is not hit, the L1-HI hot index and L1-CI cold index areas in the L1 layer are queried. The index areas are sorted by the largest key and the target key is searched using a binary search method. If the search is successful, the current time is added to the LRU access time list of the corresponding PMtable.

[0016] S11. Determine whether the L1 layer needs to transfer the hot and cold indexes based on the hot and cold index transfer strategy. If so, update the hot and cold index areas and dynamically change the hot and cold index transfer strategy based on the read frequency of the layer.

[0017] S12. If the L1 layer data does not hit, query the L2-HI hot index and L2-CI cold index area in the L2 layer; if the L2 layer data does not hit, query the L3-HI hot index and L3-CI cold index area in the L3 layer. This ends here.

[0018] Preferably, in step S1, the storage system uses NVM as the intermediate layer between memory and disk, and stores a three-layer LSM tree structure, placing the L0, L1 and L2 layers of the LSM tree into the NVM, and the L3 layer into the underlying SSD; in the NVM, the data of the L1 to L3 layers except L0 are added with hot and cold partition indexes, the hot zone index HI points to the hot skip table in the NVM, and the cold zone index CL points to the NVM cold skip table; the data structure of the hot and cold index areas consists of a pointer to the skip table and the maximum key of the skip table, and the two areas are sorted in key order.

[0019] Preferably, in step S4, the index data structure consists of a *value pointer and a maxKey; the *value pointer points to the PMtable of the current layer, and the maxKey represents the maximum key of the PMtable.

[0020] Preferably, in step S5, the global heat statistics table records each access time of each PMtable or SSTable, and each PMtable corresponds to an LRU access time list, which records each access time of the index; after each read hit, a new record of the current time will be added to the LRU access time list corresponding to the hit PMTable or SSTable, and when the LRU list reaches the threshold, the earliest time record will be eliminated first according to the LRU strategy.

[0021] Preferably, in step S6, the LRU access time list corresponding to all PMtables participating in the compression will be checked during the compression process. The more time records a PMtable has, the higher its popularity. After the merger, the new PMtable uses the PMtable with the highest popularity among those participating in the compression as the LRU access time of the new PMtable.

[0022] Preferably, in step S11, the hot and cold index transfer strategy is specifically: updating the hot and cold index areas according to the hot and cold transfer time dynamically configured according to the reading frequency of the layer, specifically reading out the LRU access time list of the PMTable or SSTable of the layer, and then selecting the access records within the configured transfer time, one record is regarded as an access heat, and then sorting and updating the hot and cold index areas according to the heat, and taking the most frequently accessed index as the hot index and the rest as the cold indexes.

[0023] The beneficial effects of the present invention are:

[0024] (1) Faster read request response time: This invention divides the LSM tree hierarchy into hot and cold index regions. When querying, the hot index region can be searched first, and the cold index region can be searched if no hit is found. Furthermore, L0-L2 are placed in the NVM layer and a skip table is used to store data, which speeds up the efficiency of point requests and range queries.

[0025] (2) Fine-grained in-place data updates: Traditional LSM tree structures require the merged data to be read out in batches to the memory at the SSD layer, then completely de-reordered and written back to the SSD. The byte addressability of NVM provides the opportunity to directly update the data on the storage without reading it into the memory buffer or writing it in batches.

[0026] (3) Reducing write amplification during hot and cold data migration: The present invention uses an index mechanism to reduce write amplification caused by hot and cold data migration. The original data will not be frequently rewritten due to the hot and cold migration mechanism. When the hot and cold migration mechanism is triggered, the present invention only needs to transfer the hot and cold indexes, and the size of the index is much smaller than the size of the actual data.

[0027] (4) More accurate hot data screening: Traditional hot data statistics simply count each read operation as a hot data, without real-time statistics of the access frequency within the most recent time period, resulting in errors in the definition of hot and cold data. This invention records the read access time of each PMTable, and each record represents a read hot data. The criterion for determining hot and cold data is the access time record within the most recent period. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 A flow chart for designing read optimization for hot and cold partition indexes on LSM layers based on NVM.

[0029] Figure 2 Design structure diagram for read optimization of NVM-based LSM hierarchical hot and cold partition indexes. DETAILED DESCRIPTION

[0030] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0031] The present invention provides a technical solution: a read optimization method for LSM hierarchical hot and cold partition indexes based on NVM, in which the MemTable and ImmuTable in memory have the same design structure as the original LSM tree in memory. In the LSM tree structure, NVM is used as an intermediate layer between memory and disk and stores a three-layer LSM tree structure. The L0, L1, and L2 layers of the LSM tree are placed in NVM, and the L3 layer is placed in the underlying SSD. This has several advantages: 1. NVM has byte addressing capabilities, so the data structure of the memory can be directly copied to NVM, without the need for similar operations such as serialization to disk or deserialization to memory as in the relationship between memory and disk. The data structure in NVM is organized into a jump table, called PMtable, so that the ImmuTable can be refreshed directly to NVM to maintain a consistent data structure. 2. NVM has persistence capabilities. The original LSM tree is persisted on a disk or SSD. Placing the L0 to L2 layers of the LSM tree in NVM does not require worrying about data loss due to power outages and other situations.

[0032] Hot and cold partition indexes are added to the data in the L1 to L3 layers except L0 in the NVM. The hot zone index HI (HotIndex) points to the hot skip table in the NVM, and the cold zone index CL (ColdIndex) points to the cold skip table in the NVM. The data structure of the hot and cold index areas consists of a pointer to the skip table and the maximum key of the skip table, and the two areas are sorted in key order. The L3 layer in the SSD also adds hot and cold partition indexes, and the hot and cold indexes are placed in the NVM. The data file of the L3 layer is composed of SSTables, just like the original LSM tree. The L0 layer compression adopts a small amount of data compression method. When there are two PMtables in the L0 layer, they are compressed and merged and written to the L1 layer. The compression and merging methods of the remaining layers are the same as the original LSM tree. The present invention maintains a global heat statistics table, which records the time of each access to each PMtable or SSTable. Each PMtable corresponds to an LRU access time list, which records the time of each access to the index. After each read hit, a new record of the current time will be added to the LRU access time list corresponding to the hit PMTable or SSTable. When the LRU list reaches the threshold, the oldest record will be eliminated first. Then the hot and cold index transfer operation of the layer where the hit data is located will be performed. The transfer strategy is to update the hot and cold index areas according to the hot and cold transfer time dynamically configured according to the reading frequency of the layer. Specifically, the LRU access time list of the PMTable or SSTable of the layer is read out, and then the access records within the configured transfer time are selected. One record is regarded as an access heat, and then sorted according to the heat and the hot and cold index areas are updated. The most frequently accessed index is used as the hot index, and the rest are used as cold indexes.

[0033] When a certain layer of the LSM tree reaches the threshold, Ln and Ln+1 will be compressed and merged. When compressing and merging multiple jump tables, the PMTable with the highest access frequency within a period of time is selected as the access time list of the new PMTable. The advantage of this design is that the search of the traditional LSM tree is a hierarchical sequential search, while the design of the hot and cold index partitions can prioritize the search of hot areas in each search, and then switch to searching the cold areas if there is no hit. In addition, the present invention adds an access time list to each PMTable, and does not rely solely on the frequency of heat to determine hot and cold data, but instead uses the most frequently accessed data within a timestamp as the hot data within a time period.

[0034] Finally, the present invention uses a hash skip table as the system cache. Each time the cache is written, the first few bits of the key to be cached are extracted and then hashed into a bucket. Each bucket is a skip table, which can support both point queries and range queries.

[0035] The present invention is based on the read optimization method of the LSM level cold and hot partition index of NVM, such as Figure 1 As shown, the design operation includes the following implementation steps:

[0036] S1. Initialize the storage system;

[0037] S2, such as Figure 2 As shown, Figure 2 The black arrows and numerical sequence in the figure represent the write operation process. First, the data is written to the WAL log file to persist it. Next, the new data is compared with the cached data to update the cached data. The first few digits of the write key are hashed to the target bucket, and then the data is updated or inserted according to the skip table in the bucket. The data structure of the node in the skip table consists of value and isContinuous. IsContinuous indicates a range query. If isContinuous is true, it indicates a range value from the previous range query, and the keys are continuous in the underlying storage. If it is false, it indicates the cached content of a single point query. If there is a duplicate key, the update operation is performed. If the inserted key value is within a range, the range is inserted and the isContinuous value of the current key is set to true. The write operation then starts by storing a batch of data in the in-memory Memtable, converting the Memtable to an Immutable Memtable, and then directly copying the ImmutableMemtable's skip table to the L0 layer in the NVM using functions provided by the PMDK.

[0038] S3: When the L0 layer reaches two PMtables, the L0 data is compressed and merged into the L1 area.

[0039] S4. Build an index data structure for each PMtable of L1 layer data; the structure consists of a *value pointer and a maxKey. The *value pointer points to the PMtable of the current layer, and the maxKey represents the maximum key of the PMtable.

[0040] S5. Define and maintain a global heat record table; the global heat record table records a list of the least-recently-used (LRU) access times of the PMtable. After each access, a record of the current read time is added to the LRU access time list corresponding to the PMtable. When the LRU access time list is full, the oldest time record (i.e., the one furthest from the current time) is eliminated according to the LRU policy.

[0041] S6. When the L1 layer data is full, the L1 layer data is compressed and merged into the L2 area. During the compression process, the LRU access time list corresponding to all PMtables participating in the compression is checked. The more time records a PMtable has, the higher its popularity. Therefore, after the merger, the new PMtable uses the PMtable with the highest popularity among those participating in the compression as the LRU access time of the new PMtable.

[0042] S7. When the L2 layer data is full, it is compressed into the SSD using the same compression strategy as the L1 layer; the same index structure is also constructed in the SSD to index the SSTable in the SSD.

[0043] S8, Figure 2 The red arrows and alphabetical order in the middle represent the flow of read operations. First, the cache is queried: the hash value of the first few digits of the query key is used to route to the specified bucket, and then the query is performed in the bucket.

[0044] S9. If the cache misses, the query goes to the MemTable and Immutable Memtable in memory. If there is no hit, the query goes to the L0 layer of the LSM tree. Since the L0 layer is unordered, all data in L0 must be queried to find the latest data.

[0045] S10: If the L0 data is not hit, the L1-HI hot index and L1-CI cold index areas in the L1 layer are queried. The index areas are sorted by the largest key, so a binary search can be used to find the target key. If the search is successful, the current time is added to the LRU access time list of the corresponding PMtable. When the LRU access time list is full, the node pointed to by the next pointer of the head node (i.e., the node with the earliest access time) is eliminated. After the query is completed, the searched key value is cached.

[0046] S11. According to the hot and cold index transfer strategy, determine whether the L1 layer needs to transfer the hot and cold indexes; if transfer is required, update the hot and cold index areas, and dynamically change the hot and cold index transfer strategy according to the reading frequency of the layer.

[0047] S12. If the L1 layer data does not hit, query the L2-HI hot index and L2-CI cold index area in the L2 layer; if the L2 layer data does not hit, query the L3-HI hot index and L3-CI cold index area in the L3 layer. The search method is the same as steps S10 and S11, and the process ends here.

[0048] It should be noted that, in this document, the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or apparatus comprising the element.

[0049] The terms used in the embodiments of the present invention are only for the purpose of describing specific embodiments and are not intended to limit the present invention. The singular forms "a", "an", "the" and "the" used in the embodiments of the present invention and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.

[0050] It should be understood that the term "and / or" as used herein is merely a description of the relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A exists alone, A and B exist simultaneously, or B exists alone. Furthermore, the character " / " in this document generally indicates that the associated objects are in an "or" relationship.

[0051] The word "if," as used herein, may be interpreted as "at the time of" or "when" or "in response to determining" or "in response to detecting," depending on the context. Similarly, the phrases "if it is determined" or "if (stated condition or event) is detected" may be interpreted as "when it is determined" or "in response to the determination" or "when detecting (stated condition or event)" or "in response to detecting (stated condition or event)," depending on the context.

[0052] The references to "first" and "second" in the embodiments merely distinguish similar objects and do not represent a specific ordering of the objects. It is understood that the specific order or precedence of "first" and "second" can be interchanged where appropriate. It should be understood that the objects distinguished by "first" and "second" can be interchanged where appropriate, so that the embodiments described herein can be implemented in an order other than that illustrated or described herein.

[0053] Although the present invention has been described in detail with reference to the aforementioned embodiments, it is still possible for those skilled in the art to modify the technical solutions described in the aforementioned embodiments, or to make equivalent substitutions for some of the technical features therein. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A read optimization method for LSM hierarchical hot and cold partition indexes based on NVM, characterized in that: The steps include: S1. Initialize the storage system; S2: First, write the WAL log file to persist the data. Then, compare the new data with the cached data, update the cached data, hash it to the target bucket based on the first few digits of the write key, and then update or insert it based on the skip table in the bucket. Then, start storing a batch of data in the in-memory Memtable, convert the Memtable to an Immutable Memtable, and then directly use the functions provided by PMDK to copy the Immutable Memtable's skip table to the L0 layer in NVM. S3: When the L0 layer reaches two PMtables, the L0 data is compressed and merged into the L1 area; S4. Build an index data structure for each PMtable of L1 layer data; S5. Define and maintain a global heat record table; S6. When the L1 layer data is full, the L1 layer data is compressed and merged into the L2 area; S7: When the L2 layer is full, the data is compressed into the SSD using the same compression strategy as the L1 layer. S8, query cache: Routes to a specified bucket based on the hash value of the first few digits of the query key, and then performs a query in the bucket; S9: If the cache misses, the query goes to the MemTable and Immutable Memtable in memory. If there is no hit, the query goes to the L0 layer of the LSM tree and searches all the data in L0 for the latest data. S10. When the L0 layer data is not hit, the L1-HI hot index and L1-CI cold index areas in the L1 layer are queried. The index areas are sorted by the largest key and the target key is searched using a binary search method. If the search is successful, the current time is added to the LRU access time list of the corresponding PMtable. S11. Determine whether the L1 layer needs to transfer the hot and cold indexes based on the hot and cold index transfer strategy. If so, update the hot and cold index areas and dynamically change the hot and cold index transfer strategy based on the read frequency of the layer. S12. If the L1 layer data does not hit, query the L2-HI hot index and L2-CI cold index area in the L2 layer; if the L2 layer data does not hit, query the L3-HI hot index and L3-CI cold index area in the L3 layer. This ends here.

2. The read optimization method for the NVM-based LSM hierarchical hot and cold partition index according to claim 1 is characterized by: In step S1, the storage system uses NVM as the intermediate layer between memory and disk, and stores a three-layer LSM tree structure. The L0, L1 and L2 layers of the LSM tree are placed in NVM, and the L3 layer is placed in the underlying SSD; in NVM, hot and cold partition indexes are added to the L1 to L3 layer data except L0, the hot zone index HI points to the hot skip table in NVM, and the cold zone index CL points to the NVM cold skip table; the data structure of the hot and cold index areas consists of a pointer to the skip table and the maximum key of the skip table, and the two areas are sorted in key order.

3. The read optimization method for the NVM-based LSM hierarchical hot and cold partition index according to claim 1 is characterized in that: In step S4, the index data structure consists of a *value pointer and a maxKey; the *value pointer points to the PMtable of the current layer, and the maxKey represents the maximum key of the PMtable.

4. The read optimization method for the NVM-based LSM hierarchical hot and cold partition index according to claim 1 is characterized in that: In step S5, the global heat statistics table records each access time of each PMtable or SSTable. Each PMtable corresponds to an LRU access time list, which records each access time of the index. After each read hit, a new record of the current time will be added to the LRU access time list corresponding to the hit PMTable or SSTable. When the LRU list reaches the threshold, the earliest time record will be eliminated first according to the LRU strategy.

5. The read optimization method for the NVM-based LSM hierarchical hot and cold partition index according to claim 1 is characterized in that: In step S6, the LRU access time list corresponding to all PMtables participating in the compression will be checked during the compression process. The more time records a PMtable has, the higher its popularity. After the merger, the new PMtable uses the PMtable with the highest popularity among those participating in the compression as the LRU access time of the new PMtable.

6. The read optimization method for the NVM-based LSM hierarchical cold and hot partition index according to claim 1 is characterized in that: In step S11, the hot and cold index transfer strategy is specifically: updating the hot and cold index areas according to the hot and cold transfer time dynamically configured according to the reading frequency of the layer, specifically reading out the LRU access time list of the PMTable or SSTable of the layer, and then selecting the access records within the configured transfer time, one record is regarded as an access heat, and then sorting and updating the hot and cold index areas according to the heat, and taking the most frequently accessed index as the hot index and the rest as the cold indexes.