Data processing method and device of Cassandra key value storage system based on interval tree layering

By introducing a design based on interval tree hierarchy in the Cassandra key-value storage system, the problem of frequent construction and update of interval trees consumes resources, and more efficient data writing and reading performance is achieved.

CN119961267AActive Publication Date: 2025-05-09HUAQIAO UNIVERSITY

Patent Information

Application Number
CN202510428250.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-08
Publication Date
2025-05-09
Estimated Expiration
2045-04-08

AI Technical Summary

Technical Problem

In the existing Cassandra key-value storage system, frequent construction and update of interval trees consume system resources, affecting read and write performance, and the space for improvement of interval trees is limited.

Method used

The Cassandra key-value storage system based on interval tree hierarchy is adopted, which is divided into low-level and high-level interval trees, which are used to manage the key range of ordered string table files of low-level and high-level external memory layers respectively. By reducing the frequent reconstruction of interval trees and optimizing data query paths, the system performance is improved.

Benefits of technology

It effectively reduces the resource consumption of interval tree reconstruction, improves write performance, and reduces computational burden and memory consumption by flexibly selecting query levels, improves read efficiency, and significantly improves Cassandra's overall read and write performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119961267A_ABST
    Figure CN119961267A_ABST
Patent Text Reader

Abstract

The invention discloses a data processing method and device of a Cassandra key value storage system based on interval tree layering, and relates to the field of data storage.The method includes the steps that an external storage layer and an interval tree are subjected to hierarchical design, the external storage layer is divided into a low-level external storage layer and a high-level external storage layer according to the hierarchical characteristics of an LSM tree, and the low-level external storage layer and the high-level external storage layer are divided into the low-level external storage layer and the high-level external storage layer; and the interval tree is correspondingly divided into a low-level interval tree and a high-level interval tree. Constructing a low-level interval tree to realize quick response aiming at the data which is frequently changed in the low-level external storage layer; and for relatively stable data of a high-level external storage layer, a high-level interval tree is constructed, and efficient data management and query are realized. According to the method, a self-adaptive construction strategy is adopted, an original interval tree is preferentially subjected to incremental updating when data is written every time, reconstruction is triggered only when the interval tree is in an unbalanced state, and unnecessary calculation overhead is reduced. According to the method, the problems of high performance loss and low overall reading and storage efficiency caused by frequent reconstruction of the interval tree in a high-frequency reading and writing scene of Cassandra are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data storage, and in particular to a data processing method and device for a Cassandra key-value storage system based on interval tree stratification. Background Art

[0002] Cassandra is an open source distributed NoSQL database, widely used in scenarios that require high availability, scalability and fault tolerance, especially in applications that require large-scale data storage and high concurrent access. Its design concept is based on the LSM tree (Log-Structured Merge-Tree) data structure, providing efficient data writing, reading and updating capabilities. Cassandra uses a distributed architecture to ensure that data is evenly distributed among multiple nodes through partitioning, load balancing and replication mechanisms, thereby improving the scalability and fault tolerance of the system.

[0003] LSM tree is the core data structure of Cassandra, and its biggest advantage is that it can optimize write performance. Data is first written to the mutable memory table (MemTable) in memory. When the mutable memory table reaches the capacity limit, it will be converted into a read-only immutable memory table and asynchronously refreshed to disk.

[0004] On disk, data is initially stored in the L0 layer of the LSM tree in the form of Sorted String Table files. As the number of L0 layer files continues to grow, the system will periodically trigger inter-layer merging (Compaction) operations to migrate data to higher layers in an orderly manner. During the merging process, the system will select an ordered string table file from the i layer and find a set of files in the i+1 layer whose key ranges overlap with it. These files will then be loaded into memory for unified merging and sorting. During this process, the system will check the validity of the key-value pairs, remove duplicate or expired data, and finally generate a new ordered string table file and write it back to the i+1 layer.

[0005] Through this hierarchical, progressive, and continuously organized data structure, the LSM tree not only achieves efficient batch writing, but also ensures the orderliness and efficient reading performance of data on the disk, which makes Cassandra particularly suitable for write-intensive application scenarios. The distributed architecture enables Cassandra to support large-scale, geographically distributed clusters, and evenly distribute data to multiple nodes through a consistent hashing algorithm. Cassandra's partitioning and replication mechanisms ensure high data availability and fault tolerance. When a node fails, the system can automatically initiate a data transfer request and ensure that data is not lost, and can still provide continuous services. However, as the amount of data increases, the scalability and performance of the system will also face bottlenecks.

[0006] Although Cassandra has powerful distributed storage capabilities, it still has some shortcomings. In order to optimize read performance, Cassandra introduced an interval tree structure in metadata management to improve search efficiency. However, the improvement space of the interval tree is limited, and its frequent construction and update will consume system resources, thus affecting the overall read and write performance. Therefore, how to improve the overall performance of the system by improving Cassandra's interval tree structure has become an important direction of current research. Summary of the invention

[0007] The purpose of this application is to propose a data processing method and device for a Cassandra key-value storage system based on interval tree hierarchy in response to the above-mentioned technical problems.

[0008] In a first aspect, the present invention provides a data processing method for a Cassandra key-value storage system based on interval tree layering, wherein the external storage layer of the LSM tree of each storage node in the Cassandra key-value storage system is divided into a low-level external storage layer and a high-level external storage layer, and the corresponding interval tree is divided into a low-level interval tree and a high-level interval tree, which are respectively used to record and manage the key range of each ordered string table file of the low-level external storage layer and the key range of each ordered string table file of the high-level external storage layer, and the data processing method includes a data writing process, and the steps include:

[0009] Obtain a data write request and obtain the primary key of the data to be written, perform hash calculation on the primary key of the data to be written, obtain a hash value corresponding to the data to be written, and determine the target storage node where the data to be written is to be stored according to the hash value corresponding to the data to be written;

[0010] If an immutable memory table is generated in the process of writing the data to be written to the target storage node, and the immutable memory table is written to the low-level external storage layer in the LSM tree of the target storage node and an ordered string table file of the low-level external storage layer is generated, then determine whether the inter-layer merge operation of the low-level external storage layer in the LSM tree will be triggered. If not, check whether the left subtree and right subtree of each node in the low-level interval tree are balanced. If not, trigger the reconstruction operation of the low-level interval tree; if it is triggered, after completing the inter-layer merge operation of the low-level external storage layer in the LSM tree, trigger the reconstruction operation of the low-level interval tree, and determine Whether the ordered string table file of the low-level external storage layer needs to be written to the high-level external storage layer in the LSM tree and generate an ordered string table file of the high-level external storage layer; if so, determine whether the inter-layer merge operation of the high-level external storage layer in the LSM tree will be triggered after the ordered string table file of the low-level external storage layer is written to the high-level external storage layer in the LSM tree; if so, after completing the inter-layer merge operation of the high-level external storage layer in the LSM tree, trigger the reconstruction operation of the high-level interval tree; otherwise, check whether the left subtree and right subtree of each node of the high-level interval tree are balanced; if not, trigger the reconstruction operation of the high-level interval tree.

[0011] As a preference, it also includes:

[0012] In response to determining that the left subtree and the right subtree of each node of the lower-level interval tree are balanced, the left subtree and the right subtree of each node of the higher-level interval tree are balanced, and / or the ordered string table file of the lower-level external memory layer does not need to be written to the higher-level external memory layer in the LSM tree, the existing structure is maintained.

[0013] As a preference, it also includes:

[0014] After completing the rebuild operation of the high-level interval tree, the existing structure is maintained.

[0015] Preferably, in the target storage node, it is first determined whether the data to be written can be written into the variable memory table of the memory layer of the LSM tree of the target storage node. If so, the data to be written is written into the variable memory table of the memory layer of the LSM tree of the target storage node, and the existing structure is maintained; otherwise, the variable memory table is converted into an immutable memory table, a new variable memory table is created, the data to be written is written into the new variable memory table, and the immutable memory table is written into the external memory layer of the LSM tree of the target storage node.

[0016] As a preference, the LSM tree has a total of N external storage layers, with the corresponding levels increasing from top to bottom, where the top The first layer is set as the low-level external memory layer, and the remaining layers are high-level external memory layers.

[0017] Preferably, the process further includes a data reading process, the steps of which include:

[0018] Obtain the data read request and parse it to obtain the primary key of the data to be read and calculate the corresponding hash value, and determine the target storage node where the data to be read is located according to the hash value;

[0019] According to the primary key of the data to be read, the data to be read is first searched in the memory layer of the target storage node. If the data to be read cannot be found in the memory layer, the ordered string table file where the data to be read is located is searched in the low-level interval tree; if the ordered string table file where the data to be read is located is found in the low-level interval tree, the reading result is obtained through the ordered string table file where the data to be read is located; if the ordered string table file where the data to be read is located is not found in the low-level interval tree, the ordered string table file where the data to be read is located is searched in the high-level interval tree; if the ordered string table file where the data to be read is located is found in the high-level interval tree, the reading result is obtained through the ordered string table file where the data to be read is located; if the ordered string table file where the data to be read is located is found in the high-level interval tree, the reading result is obtained through the ordered string table file where the data to be read is located, if the ordered string table file where the data to be read is located is not found in the high-level interval tree, the reading failure is returned.

[0020] In a second aspect, the present invention provides a data processing device for a Cassandra key-value storage system based on interval tree layering, wherein the external storage layer of the LSM tree of each storage node in the Cassandra key-value storage system is divided into a low-level external storage layer and a high-level external storage layer, and the corresponding interval tree is divided into a low-level interval tree and a high-level interval tree, which are respectively used to record and manage the key range of each ordered string table file of the low-level external storage layer and the key range of each ordered string table file of the high-level external storage layer, and the data processing device includes a data writing module, including:

[0021] The node determination module is configured to obtain a data write request and obtain a primary key of the data to be written, perform a hash calculation on the primary key of the data to be written, obtain a hash value corresponding to the data to be written, and determine a target storage node where the data to be written is to be stored according to the hash value corresponding to the data to be written;

[0022] The write module is configured to generate an immutable memory table during the process of writing the data to be written to the target storage node, write the immutable memory table to the low-level external storage layer in the LSM tree of the target storage node and generate an ordered string table file of the low-level external storage layer, then determine whether the inter-layer merge operation of the low-level external storage layer in the LSM tree will be triggered. If it will not be triggered, check whether the left subtree and the right subtree of each node in the low-level interval tree are balanced. If they are not balanced, trigger the reconstruction operation of the low-level interval tree; if it will be triggered, after completing the inter-layer merge operation of the low-level external storage layer in the LSM tree, trigger the reconstruction operation of the low-level interval tree. , and determine whether the ordered string table file of the low-level external storage layer needs to be written to the high-level external storage layer in the LSM tree and generate an ordered string table file of the high-level external storage layer. If so, determine whether the inter-layer merge operation of the high-level external storage layer in the LSM tree will be triggered after the ordered string table file of the low-level external storage layer is written to the high-level external storage layer in the LSM tree. If so, after completing the inter-layer merge operation of the high-level external storage layer in the LSM tree, trigger the reconstruction operation of the high-level interval tree. Otherwise, check whether the left subtree and right subtree of each node of the high-level interval tree are balanced. If not, trigger the reconstruction operation of the high-level interval tree.

[0023] In a third aspect, the present invention provides an electronic device comprising one or more processors; a storage device for storing one or more programs, wherein when the one or more programs are executed by the one or more processors, the one or more processors implement the method described in any implementation manner in the first aspect.

[0024] In a fourth aspect, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.

[0025] In a fifth aspect, the present invention provides a computer program product, comprising a computer program, which, when executed by a processor, implements the method described in any implementation manner in the first aspect.

[0026] Compared with the prior art, the present invention has the following beneficial effects:

[0027] (1) The data processing method of the Cassandra key-value storage system based on interval tree layering proposed in the present invention reduces the waste of system resources by reducing the need for frequent interval tree construction during data writing. Traditional methods usually require the entire interval tree to be rebuilt each time, while the present invention uses interval tree layering design so that each time reconstruction usually only requires reconstruction of the low-level interval tree, and the reconstruction frequency of the high-level interval tree is low, thereby effectively reducing the reconstruction overhead and improving the writing performance.

[0028] (2) The data processing method of the Cassandra key-value storage system based on interval tree layering proposed in the present invention can flexibly select the interval tree of the corresponding level to be accessed during reading by dividing the interval tree into a low-level interval tree and a high-level interval tree during the data reading process. In some cases, only querying the low-level interval tree can meet the demand, thereby reducing the computational burden and memory consumption during the query process and improving the query efficiency.

[0029] (3) The data processing method of the Cassandra key-value storage system based on interval tree layering proposed in the present invention can reduce the resource consumption caused by interval tree construction, and utilize the characteristics of interval tree layering to greatly improve the read and write performance of Cassandra; through more precise interval tree management, high efficiency in the data writing and reading process is ensured, especially when processing large-scale data, it can better balance performance and resource utilization. This optimization not only improves the response speed of the system, but also enhances its adaptability in high concurrency and high throughput scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0031] Figure 1 A flow chart of a data processing method for a Cassandra key-value storage system based on interval tree hierarchy according to an embodiment of the present application;

[0032] Figure 2 A schematic diagram of the structure of the external storage layer and interval tree of the data processing method of the Cassandra key-value storage system based on interval tree layering in an embodiment of the present application;

[0033] Figure 3 A schematic diagram of a data processing device of a Cassandra key-value storage system based on interval tree hierarchy according to an embodiment of the present application;

[0034] Figure 4 A schematic diagram of the hardware structure of an electronic device provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0035] In order to make the purpose, technical scheme and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0036] Figure 1 A data processing method for a Cassandra key-value storage system based on interval tree layering provided by an embodiment of the present application is shown. The external storage layer of the LSM tree of each storage node in the Cassandra key-value storage system is divided into a low-level external storage layer and a high-level external storage layer, and the corresponding interval tree is divided into a low-level interval tree and a high-level interval tree, which are respectively used to record and manage the key range of each ordered string table file of the low-level external storage layer and the key range of each ordered string table file of the high-level external storage layer. The data processing method includes a data writing process, and the steps include:

[0037] S1, obtain a data write request and obtain the primary key of the data to be written, perform hash calculation on the primary key of the data to be written, obtain the hash value corresponding to the data to be written, and determine the target storage node where the data to be written is to be stored according to the hash value corresponding to the data to be written.

[0038] In a specific embodiment, the LSM tree has a total of N external storage layers, and the corresponding levels increase from top to bottom, where the front layer at the top The first layer is set as the low-level external memory layer, and the remaining layers are high-level external memory layers.

[0039] Specifically, the overall structure of the LSM tree (Log-Structured Merge-Tree) consists of a memory layer and an external memory layer. The memory layer contains mutable memory tables and immutable memory tables. When the mutable memory table is full, it can be converted into an immutable memory table and prepared to be refreshed to the external memory layer. The external memory layer consists of multiple levels, each of which contains several sorted string table files as the persistence form of data on the disk.

[0040] The interval tree is a metadata information structure used to record and manage the key range (i.e., data interval) of each ordered string table file in the external memory layer. Each ordered string table file stores an ordered key range on the disk, and the interval tree constructs a tree structure through these key ranges, which is convenient for fast positioning and retrieval. Each node of the interval tree usually contains the minimum key (Low), maximum key (High), and center key (Center) information corresponding to the ordered string table, which is used for efficient interval search and interval overlap judgment. The interval tree is stored in the memory and the key range in the interval tree is synchronously updated according to the key range stored in the external memory layer.

[0041] The construction of the interval tree follows the attribution rule based on the center key: taking the center key as the dividing point, all intervals covering the center key are placed in the node; the intervals not covering the center key are recursively placed in the left subtree or right subtree of the node according to their positions. Figure 2 In one embodiment, it is assumed that there are two ordered string table files with key interval ranges of 6-12 and 9-15, and the central key of the current node is 12. Since both intervals cover the central key 12, that is, their interval ranges include 12, they will be mounted on the current node with the central key 12 at the same time. The intervals with a key interval range less than 12 or greater than 12 will continue to be recursively placed in the left subtree or right subtree corresponding to the current node. When the key range of the data written gradually becomes larger, the left subtree and / or right subtree of the current node can be further extended as the next node and generate the left subtree and / or right subtree of the next node.

[0042] The embodiment of the present application divides the external memory layer into layers, which are divided into a low-level external memory layer and a high-level external memory layer. Figure 2 In one embodiment, there are 7 layers in the external memory layer, of which L0 to L3 are low-level external memory layers, and a low-level interval tree is constructed for them to quickly respond to frequent changes in data; while L4 to L6 are high-level external memory layers, whose data are relatively stable, and another high-level interval tree is constructed separately to achieve efficient management and query. Therefore, two interval trees, a low-level interval tree and a high-level interval tree, are used to record and manage the key ranges of the ordered string table files in the low-level external memory layer and the high-level external memory layer respectively. In a preferred embodiment, the topmost layer in the external memory layer is The layers are used as the low-level external storage layer, and the remaining layers at the bottom are used as the high-level external storage layer. In other embodiments, other layers can also be selected as the boundary between the low-level external storage layer and the high-level external storage layer.

[0043] During the data writing process, the acquired data writing request is first parsed to obtain the primary key of the data to be written, and a hash calculation is performed on the primary key of the data to be written. According to the corresponding hash value, the target storage node where the data to be written will be stored is determined according to the hash value.

[0044] In a specific embodiment, in the target storage node, it is first determined whether the data to be written can be written into the variable memory table of the memory layer of the LSM tree of the target storage node. If so, the data to be written is written into the variable memory table of the memory layer of the LSM tree of the target storage node, and the existing structure is maintained; otherwise, the variable memory table is converted into an immutable memory table, a new variable memory table is created, the data to be written is written into the new variable memory table, and the immutable memory table is written into the external memory layer of the LSM tree of the target storage node.

[0045] Specifically, on the target storage node, determine whether the data to be written can be written into the variable memory table. If the amount of data to be written is less than or equal to the remaining space of the variable memory table, the data to be written can be written into the variable memory table and the existing structure is maintained. If the amount of data to be written is greater than the remaining space of the variable memory table, the data to be written cannot be written into the variable memory table. It is necessary to convert the variable memory table into an immutable memory table, create a new variable memory table, and write the data to be written into the new variable memory table. The immutable memory table is written to the external storage layer in the LSM tree of the target storage node and an ordered string table file is generated.

[0046] S2, if an immutable memory table is generated in the process of writing the data to be written to the target storage node, and the immutable memory table is written to the low-level external storage layer in the LSM tree of the target storage node and an ordered string table file of the low-level external storage layer is generated, then it is determined whether the inter-layer merge operation of the low-level external storage layer in the LSM tree will be triggered. If it will not be triggered, check whether the left subtree and right subtree of each node in the low-level interval tree are balanced. If they are not balanced, the reconstruction operation of the low-level interval tree is triggered; if it will be triggered, after completing the inter-layer merge operation of the low-level external storage layer in the LSM tree, the reconstruction operation of the low-level interval tree is triggered, and it is determined Determine whether the ordered string table file of the low-level external storage layer needs to be written to the high-level external storage layer in the LSM tree and generate an ordered string table file of the high-level external storage layer. If so, determine whether the inter-layer merge operation of the high-level external storage layer in the LSM tree will be triggered after the ordered string table file of the low-level external storage layer is written to the high-level external storage layer in the LSM tree. If so, after completing the inter-layer merge operation of the high-level external storage layer in the LSM tree, trigger the reconstruction operation of the high-level interval tree. Otherwise, check whether the left subtree and right subtree of each node of the high-level interval tree are balanced. If not, trigger the reconstruction operation of the high-level interval tree.

[0047] In a specific embodiment, it also includes:

[0048] In response to determining that the left subtree and the right subtree of each node of the lower-level interval tree are balanced, the left subtree and the right subtree of each node of the higher-level interval tree are balanced, and / or the ordered string table file of the lower-level external memory layer does not need to be written to the higher-level external memory layer in the LSM tree, the existing structure is maintained.

[0049] In a specific embodiment, it also includes:

[0050] After completing the rebuild operation of the high-level interval tree, the existing structure is maintained.

[0051] Specifically, when the immutable memory table is written to the low-level external memory layer in the LSM tree of the target storage node, the immutable memory table is first written to the L0 layer of the low-level external memory layer, and it is determined whether the inter-layer merge operation of the low-level external memory layer in the LSM tree will be triggered after the immutable memory table is written to the L0 layer of the low-level external memory layer, that is, whether the inter-layer merge operation will occur in the L0-L3 layers. If it is not triggered, check whether the left subtree and right subtree of each node in the low-level interval tree are balanced. If balanced, continue to maintain the existing structure. If unbalanced, trigger the reconstruction operation of the low-level interval tree. When the low-level interval tree is balanced, maintain the existing structure. If it is triggered, after completing the inter-layer merge operation of the low-level external memory layer, the reconstruction operation of the low-level interval tree is triggered to restore its balance, and further determine whether the ordered string table file of the low-level external memory layer needs to be written into the high-level external memory layer and generate the ordered string table file of the high-level external memory layer. If not, continue to maintain the existing structure; if necessary, after the ordered string table file of the low-level external memory layer is written into the high-level external memory layer, determine whether the inter-layer merge operation of the high-level external memory layer will be triggered, that is, whether the inter-layer merge operation will occur in the L4-L6 layers. If it is triggered, the high-level interval tree is rebuilt, the inter-layer merge operation of the high-level external memory layer is completed to adapt to the new data distribution, and then the existing structure is maintained; if it is not triggered, check whether the left subtree and the right subtree of each node of the high-level interval tree are balanced. If balanced, continue to maintain the existing structure without any modification; if unbalanced, trigger the reconstruction operation of the high-level interval tree, and then maintain the existing structure.

[0052] Furthermore, the inter-layer merge operation is to write the ordered string table file of the previous level's external memory layer into the next level's external memory layer. Whether the inter-layer merge operation will be triggered depends on whether the amount of data in the L0 layer exceeds the data amount threshold of the L0 layer after the data to be written is written into the L0 layer. If so, the ordered string table file of the L0 layer is written to the L1 layer, and so on.

[0053] Furthermore, each interval tree in the embodiments of the present application adopts an adaptive construction strategy, and gives priority to incremental updates when writing data to be written, that is, only the local interval structure is adjusted for the newly added data to reduce computing overhead and improve writing efficiency. In one of the embodiments, it can be set that when the left subtree and the right subtree of each node of the interval tree differ by five layers or more, it means that the left subtree and the right subtree of each node of the interval tree are unbalanced, which is used as a basis for checking whether the left subtree and the right subtree of each node of the high-level interval tree are balanced. Only when the interval tree reaches an unbalanced state, such as when the left and right subtrees of any subtree differ by five layers or more, is reconstruction triggered, and the interval tree is rebuilt according to the latest data distribution, so that the interval tree is restored to balance, thereby effectively avoiding performance losses caused by frequent reconstruction.

[0054] When writing data, the data to be written is first written to the memory layer. When the memory layer is full, the data in the memory layer will be refreshed to the ordered string table file in the disk. In the traditional way, each time a new ordered string table file is written, the interval tree will be rebuilt. After adopting the hierarchical design of the interval tree of the embodiment of the present application, most of the reconstruction operations of the interval tree are aimed at the ordered string table files of the low-level external memory layer, which reduces the resource overhead required in the reconstruction process. In addition, with the help of the adaptive construction method, there is no need to rebuild the interval tree every time, and reconstruction is only performed when the interval tree reaches an extremely unbalanced state, thereby avoiding the performance degradation of the system due to frequent reconstruction.

[0055] In a specific embodiment, a data reading process is also included, and the steps include:

[0056] Obtain the data read request and parse it to obtain the primary key of the data to be read and calculate the corresponding hash value, and determine the target storage node where the data to be read is located according to the hash value;

[0057] According to the primary key of the data to be read, the data to be read is first searched in the memory layer of the target storage node. If the data to be read cannot be found in the memory layer, the ordered string table file where the data to be read is located is searched in the low-level interval tree; if the ordered string table file where the data to be read is located is found in the low-level interval tree, the reading result is obtained through the ordered string table file where the data to be read is located; if the ordered string table file where the data to be read is located is not found in the low-level interval tree, the ordered string table file where the data to be read is located is searched in the high-level interval tree; if the ordered string table file where the data to be read is located is found in the high-level interval tree, the reading result is obtained through the ordered string table file where the data to be read is located; if the ordered string table file where the data to be read is located is found in the high-level interval tree, the reading result is obtained through the ordered string table file where the data to be read is located, if the ordered string table file where the data to be read is located is not found in the high-level interval tree, the reading failure is returned.

[0058] Specifically, the embodiments of the present application may also propose a data reading method for a Cassandra key-value storage system based on an interval tree hierarchy, the steps of which are as follows: first, obtain a data read request, parse the data read request to obtain the primary key of the data to be read and calculate the corresponding hash value, and the coordination node determines the target storage node where the data to be read is located according to the hash value corresponding to the primary key of the data to be read. The coordination node sends the data read request to the corresponding target storage node. After receiving the data read request, the target storage node first searches for the data to be read in the memory layer. If the data to be read is found in the memory layer, the reading result is returned to the coordination node, and the user's read request is finally completed. If the data to be read cannot be found in the memory layer, the target storage node first searches for the ordered string table file where the data to be read is located in the low-level interval tree according to the primary key of the data to be read. If the ordered string table file where the data to be read is located can be found in the low-level interval tree, the target storage node will read the data in the ordered string table file where the data to be read is located, obtain the reading result, and return the reading result to the coordination node, and finally complete the user's read request; if the ordered string table file where the data to be read is located cannot be found in the low-level interval tree, the target storage node will search for the ordered string table file where the data to be read is located in the high-level interval tree according to the primary key of the data to be read. If the ordered string table file where the data to be read is located can be found in the high-level interval tree, the target storage node will read the data in the ordered string table file where the data to be read is located, obtain the reading result, and return the reading result to the coordination node, and finally complete the user's read request; if the ordered string table file where the data to be read is located cannot be found in the high-level interval tree, the result of read failure is returned.

[0059] During the data search process, the required data will first be searched in the memory layer. If it is not found, the corresponding ordered string table file will be located through the interval tree. Since most read operations are concentrated in the low-level external memory layer of the LSM tree, only the low-level interval tree needs to be searched, without traversing all levels of the interval tree, which greatly reduces the overhead of the search process and improves the reading efficiency.

[0060] Further references Figure 3 As an implementation of the methods shown in the above figures, the present application provides an embodiment of a data processing device for a Cassandra key-value storage system based on interval tree hierarchy. Figure 1 Corresponding to the method embodiment shown, the device can be specifically applied to various electronic devices.

[0061] The embodiment of the present application provides a data processing device for a Cassandra key-value storage system based on interval tree layering. The external storage layer of the LSM tree of each storage node in the Cassandra key-value storage system is divided into a low-level external storage layer and a high-level external storage layer, and the corresponding interval tree is divided into a low-level interval tree and a high-level interval tree, which are respectively used to record and manage the key range of each ordered string table file of the low-level external storage layer and the key range of each ordered string table file of the high-level external storage layer. The data processing device includes a data writing module, including:

[0062] The node determination module 1 is configured to obtain a data write request and obtain a primary key of the data to be written, perform a hash calculation on the primary key of the data to be written, obtain a hash value corresponding to the data to be written, and determine a target storage node where the data to be written is to be stored according to the hash value corresponding to the data to be written;

[0063] The writing module 2 is configured to generate an immutable memory table during the process of writing the data to be written into the target storage node, write the immutable memory table into the low-level external storage layer in the LSM tree of the target storage node and generate an ordered string table file of the low-level external storage layer, then determine whether the inter-layer merge operation of the low-level external storage layer in the LSM tree will be triggered. If it will not be triggered, check whether the left subtree and the right subtree of each node in the low-level interval tree are balanced. If they are not balanced, trigger the reconstruction operation of the low-level interval tree; if it will be triggered, after completing the inter-layer merge operation of the low-level external storage layer in the LSM tree, trigger the reconstruction operation of the low-level interval tree. It determines whether the ordered string table file of the low-level external storage layer needs to be written into the high-level external storage layer in the LSM tree and generates an ordered string table file of the high-level external storage layer. If so, it determines whether the inter-layer merge operation of the high-level external storage layer in the LSM tree will be triggered after the ordered string table file of the low-level external storage layer is written into the high-level external storage layer in the LSM tree. If so, after completing the inter-layer merge operation of the high-level external storage layer in the LSM tree, the reconstruction operation of the high-level interval tree is triggered. Otherwise, it checks whether the left subtree and the right subtree of each node of the high-level interval tree are balanced. If not, the reconstruction operation of the high-level interval tree is triggered.

[0064] Figure 4 Schematic diagram of the hardware structure of the electronic device provided by the embodiment of the present invention. Figure 4 As shown, the electronic device of this embodiment includes: a processor 401 and a memory 402; wherein the memory 402 is used to store computer-executable instructions; the processor 401 is used to execute the computer-executable instructions stored in the memory to implement the various steps performed by the electronic device in the above embodiment. For details, please refer to the relevant description in the above method embodiment.

[0065] Optionally, the memory 402 may be independent or integrated with the processor 401 .

[0066] When the memory 402 is independently provided, the electronic device further includes a bus 403 for connecting the memory 402 and the processor 401 .

[0067] The embodiment of the present invention further provides a computer storage medium, in which computer execution instructions are stored. When the processor 401 executes the computer execution instructions, the above method is implemented.

[0068] The embodiment of the present invention further provides a computer program product, including a computer program. When the computer program is executed by the processor 401, the above method is implemented.

[0069] In the embodiments provided by the present invention, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of modules is only a logical function division, and there may be other division methods in actual implementation, such as multiple modules can be combined or integrated into another system, or some features can be ignored or not executed. Another point, the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or modules, which can be electrical, mechanical or other forms.

[0070] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the modules may be selected according to actual needs to implement the solution of this embodiment.

[0071] In addition, each functional module in each embodiment of the present invention may be integrated into one processing unit, each module may exist physically separately, or two or more modules may be integrated into one unit. The unit formed by the above modules may be implemented in the form of hardware or in the form of hardware plus software functional units.

[0072] The above-mentioned integrated module implemented in the form of a software function module can be stored in a computer-readable storage medium. The above-mentioned software function module is stored in a storage medium, including a number of instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) or a processor 401 to perform some steps of the methods of various embodiments of the present application.

[0073] It should be understood that the processor 401 may be a central processing unit (CPU), or other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), etc. A general-purpose processor may be a microprocessor or the processor 401 may be any conventional processor 401, etc. The steps of the method disclosed in the invention may be directly embodied in the hardware processor 401 for execution, or may be executed by a combination of hardware and software modules in the processor 401.

[0074] The memory 402 may include a high-speed RAM memory, and may also include a non-volatile storage NVM, such as at least one disk memory, and may also be a USB flash drive, a mobile hard disk, a read-only memory, a magnetic disk, or an optical disk.

[0075] The bus 403 may be an Industry Standard Architecture (ISA), a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. The bus 403 may be divided into an address bus, a data bus, a control bus, etc. For ease of representation, the bus 403 in the drawings of the present application is not limited to only one bus 403 or one type of bus 403.

[0076] The above storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk. The storage medium can be any available medium that can be accessed by a general or special purpose computer.

[0077] An exemplary storage medium is coupled to the processor 401, so that the processor 401 can read information from the storage medium and write information to the storage medium. Of course, the storage medium can also be a component of the processor 401. The processor 401 and the storage medium can be located in an application specific integrated circuit (ASIC). Of course, the processor 401 and the storage medium can also exist as discrete components in an electronic device or a main control device.

[0078] Those skilled in the art can understand that all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, disk or optical disk and other media that can store program codes.

[0079] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or replace some or all of the technical features therein with equivalents. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A data processing method for a Cassandra key-value storage system based on interval tree stratification, characterized in that: The external memory layer of the LSM tree of each storage node in the Cassandra key-value storage system is divided into a low-level external memory layer and a high-level external memory layer, and the corresponding interval tree is divided into a low-level interval tree and a high-level interval tree, which are respectively used to record and manage the key range of each ordered string table file of the low-level external memory layer and the key range of each ordered string table file of the high-level external memory layer. The data processing method includes a data writing process, and the steps include: Obtain a data write request and obtain a primary key of the data to be written, perform a hash calculation on the primary key of the data to be written to obtain a hash value corresponding to the data to be written, and determine a target storage node where the data to be written is to be stored according to the hash value corresponding to the data to be written; If an immutable memory table is generated during the process of writing the data to be written into the target storage node, and the immutable memory table is written into the low-level external storage layer in the LSM tree of the target storage node and an ordered string table file of the low-level external storage layer is generated, it is determined whether the inter-layer merge operation of the low-level external storage layer in the LSM tree will be triggered. If it will not be triggered, it is checked whether the left subtree and the right subtree of each node of the low-level interval tree are balanced. If they are unbalanced, the reconstruction operation of the low-level interval tree is triggered; if it will be triggered, after completing the inter-layer merge operation of the low-level external storage layer in the LSM tree, the reconstruction operation of the low-level interval tree is triggered, and Determine whether the ordered string table file of the low-level external storage layer needs to be written into the high-level external storage layer in the LSM tree and generate an ordered string table file of the high-level external storage layer; if so, determine whether an inter-layer merge operation of the high-level external storage layer in the LSM tree will be triggered after the ordered string table file of the low-level external storage layer is written into the high-level external storage layer in the LSM tree; if so, after completing the inter-layer merge operation of the high-level external storage layer in the LSM tree, trigger the reconstruction operation of the high-level interval tree; otherwise, check whether the left subtree and the right subtree of each node of the high-level interval tree are balanced; if not, trigger the reconstruction operation of the high-level interval tree.

2. The data processing method of the Cassandra key-value storage system based on interval tree layering according to claim 1 is characterized in that: Also includes: In response to determining that the left subtree and the right subtree of each node of the lower-level interval tree are balanced, the left subtree and the right subtree of each node of the higher-level interval tree are balanced, and / or the ordered string table file of the lower-level external memory layer does not need to be written to the higher-level external memory layer in the LSM tree, the existing structure is maintained.

3. The data processing method of the Cassandra key-value storage system based on interval tree layering according to claim 1 is characterized in that: Also includes: After completing the rebuild operation of the high-level interval tree, the existing structure is maintained.

4. The data processing method of the Cassandra key-value storage system based on interval tree layering according to claim 1 is characterized in that: In the target storage node, it is first determined whether the data to be written can be written into the variable memory table of the memory layer of the LSM tree of the target storage node. If so, the data to be written is written into the variable memory table of the memory layer of the LSM tree of the target storage node, and the existing structure is maintained; otherwise, the variable memory table is converted into an immutable memory table, a new variable memory table is created, the data to be written is written into the new variable memory table, and the immutable memory table is written into the external memory layer of the LSM tree of the target storage node.

5. The data processing method of the Cassandra key-value storage system based on interval tree layering according to claim 1 is characterized in that: The LSM tree has N layers of external storage, and the corresponding levels increase from top to bottom. The first layer is set as the low-level external memory layer, and the remaining layers are high-level external memory layers.

6. The data processing method of the Cassandra key-value storage system based on interval tree layering according to claim 1 is characterized in that: It also includes a data reading process, the steps of which include: Obtain a data read request and parse it to obtain the primary key of the data to be read and calculate the corresponding hash value, and determine the target storage node where the data to be read is located according to the hash value; According to the primary key of the data to be read, the data to be read is first searched in the memory layer of the target storage node. If the data to be read cannot be found in the memory layer, the ordered string table file where the data to be read is located is searched in the low-level interval tree; if the ordered string table file where the data to be read is located is found in the low-level interval tree, the reading result is obtained through the ordered string table file where the data to be read is located; if the ordered string table file where the data to be read is located cannot be found in the low-level interval tree, the ordered string table file where the data to be read is located is searched in the high-level interval tree; if the ordered string table file where the data to be read is located is found in the high-level interval tree, the reading result is obtained through the ordered string table file where the data to be read is located, and if the ordered string table file where the data to be read is located cannot be found in the high-level interval tree, the reading failure is returned.

7. A data processing device for a Cassandra key-value storage system based on interval tree hierarchy, characterized in that: The external memory layer of the LSM tree of each storage node in the Cassandra key-value storage system is divided into a low-level external memory layer and a high-level external memory layer, and the corresponding interval tree is divided into a low-level interval tree and a high-level interval tree, which are respectively used to record and manage the key range of each ordered string table file of the low-level external memory layer and the key range of each ordered string table file of the high-level external memory layer, and the data processing device includes a data writing module, including: The node determination module is configured to obtain a data write request and obtain a primary key of the data to be written, perform a hash calculation on the primary key of the data to be written to obtain a hash value corresponding to the data to be written, and determine a target storage node where the data to be written is to be stored according to the hash value corresponding to the data to be written; The writing module is configured to determine whether an inter-layer merge operation of the low-level external storage layer in the LSM tree will be triggered if an immutable memory table is generated during the process of writing the data to be written into the target storage node, and the immutable memory table is written into the low-level external storage layer in the LSM tree of the target storage node and an ordered string table file of the low-level external storage layer is generated. If it is not triggered, check whether the left subtree and the right subtree of each node of the low-level interval tree are balanced. If they are not balanced, trigger the reconstruction operation of the low-level interval tree; if it is triggered, after completing the inter-layer merge operation of the low-level external storage layer in the LSM tree, trigger the reconstruction of the low-level interval tree. Operation, and determine whether the ordered string table file of the low-level external storage layer needs to be written into the high-level external storage layer in the LSM tree and generate an ordered string table file of the high-level external storage layer; if so, determine whether the inter-layer merge operation of the high-level external storage layer in the LSM tree will be triggered after the ordered string table file of the low-level external storage layer is written into the high-level external storage layer in the LSM tree; if so, after completing the inter-layer merge operation of the high-level external storage layer in the LSM tree, trigger the reconstruction operation of the high-level interval tree; otherwise, check whether the left subtree and the right subtree of each node of the high-level interval tree are balanced, and if unbalanced, trigger the reconstruction operation of the high-level interval tree.

8. An electronic device, comprising: one or more processors; a storage device for storing one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 6.

9. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

10. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, the method according to any one of claims 1 to 6 is implemented.

Citation Information

Patent Citations

  • Hotness perception local updating method applied to key value storage system

    CN114969069A

  • Key value storage method based on multi-tree conversion mechanism

    CN114996275A

  • Method for storing a data page in a data storage device using similarity based data reduction

    WO2022139626A1

Cited By

  • Log structure merge tree compression state scanning method based on order-preserving dictionary

    CN121008756A