Data storage method, apparatus and device, and computer-readable storage medium

By monitoring the resource utilization of B+ tree nodes and delaying node splitting, the problem of low utilization after B+ tree node splitting is solved, and the storage resource utilization and write efficiency are improved.

WO2025112505A1PCT designated stage expired Publication Date: 2025-06-05INSPUR SUZHOU INTELLIGENT TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/101728
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-11-28
Filing Date
2024-06-26
Publication Date
2025-06-05

AI Technical Summary

Technical Problem

When existing B+ trees are stored in data, the utilization rate after node split may be only 50%, resulting in wasted storage resources, especially when data is updated frequently.

Method used

By monitoring the node resource utilization of B+ tree, when the resource utilization of the sibling nodes of the node to be split meets the migration conditions, the node split task is prohibited and the data is migrated to the sibling nodes to delay and avoid unnecessary node splitting.

Benefits of technology

The node resource utilization rate of B+ tree storage is improved, the consumption caused by node splitting tasks is reduced, and the writing efficiency and overall system throughput is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024101728_05062025_PF_FP_ABST
    Figure CN2024101728_05062025_PF_FP_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of storage. Specifically disclosed are a data storage method, apparatus and device, and a computer-readable non-volatile storage medium. On the basis of the size of a data block of a device in which a B+ tree is located and the size of a unit of data to be stored, the order of the B+ tree to be created is determined, so as to call a B+ tree creation script to create a B+ tree of a reasonable order. The utilization of node resources of the B+ tree is monitored, and when the B+ tree meets a node splitting condition, whether a sibling node of a node to be split meets a migration condition for receiving data of the node to be split is queried first. If the migration condition is met, the B+ tree creation script is prohibited from executing a node splitting task, and the data of the node to be split is migrated to the sibling node; and if the migration condition is not met, the B+ tree is allowed to execute the node splitting task, such that the utilization rate of node resources stored in the B+ tree is improved by means of delaying the node splitting of the B+ tree and preventing unnecessary node splitting of the B+ tree, thereby improving the utilization rate of storage resources when the B+ tree is used for data storage.
Need to check novelty before this filing date? Find Prior Art

Description

Data storage method, device, equipment and computer-readable storage medium

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to the Chinese patent application filed with the China Patent Office on November 28, 2023, with application number 202311598779.7, and application name “A data storage method, device, equipment and computer-readable storage medium”, all contents of which are incorporated by reference into this application. Technical Field

[0003] The present application relates to the field of storage technology, and in particular to a data storage method, apparatus, device, and computer non-volatile readable storage medium. Background Art

[0004] In data storage systems, for metadata and other data with relatively small data units but a large amount of data, there are challenges with high-volume and highly concurrent data access and querying. To effectively manage the storage of these small data units, the current storage strategy uses a B+ tree (Balanced Plus Tree) operation algorithm for data storage. When a B+ tree node is about to be full, the current strategy is to immediately split the node to maintain the tree's balanced structure. However, this results in the utilization rate of the split node being as low as 50%, which is a significant waste of storage resources for storage systems with large and frequent data updates.

[0005] Summary of the Invention

[0006] The purpose of this application is to provide a data storage method, apparatus, device and computer non-volatile readable storage medium for improving storage resource utilization when using B+ tree for data storage.

[0007] To solve the above technical problems, the present application provides a data storage method, comprising:

[0008] Calculate the order of the B+ tree to be created based on the data block size of the device and the size of the unit data to be stored;

[0009] Writing the order of the B+ tree to be created into the B+ tree creation script of the device to call the B+ tree creation script to create a B+ tree for storing the data to be stored;

[0010] Start the node resource monitoring program for the B+ tree;

[0011] When the B+ tree is detected to meet the node splitting conditions, the node resource utilization of the sibling nodes of the node to be split is queried;

[0012] If the node resource utilization of the sibling node meets the migration conditions for receiving the data of the node to be split, the B+ tree creation script is prohibited from executing the node splitting task and the information of the sibling node and the information of the node to be split are written into the B+ tree creation script. After calling the B+ tree creation script to migrate the data of the node to be split to the sibling node, the parent node is updated according to the data information of the sibling node and the data information of the node to be split.

[0013] If the node resource utilization of all sibling nodes of the node to be split meets the migration conditions, the information that does not meet the migration conditions is written into the B+ tree creation script, so that the B+ tree creation script executes the node splitting task for the node to be split.

[0014] In some embodiments, the order of the B+ tree to be created is calculated based on the data block size of the device and the size of the unit data to be stored, including:

[0015] Run the node resource monitoring process and divide the device's data block size by the unit data size to be stored to obtain the order of the B+ tree to be created.

[0016] The B+ tree is detected to meet the node splitting conditions, including:

[0017] When it is detected that the B+ tree creation script is called to execute a write command, it is checked to determine whether the node to be written to the data to be written in the B+ tree meets the node splitting condition; the node resource utilization of each leaf node of the B+ tree is monitored regularly, and it is detected that there is a leaf node that meets the node splitting condition.

[0018] In some embodiments, the node resource utilization of the sibling node satisfies the migration conditions for receiving data of the node to be split, including:

[0019] If the sibling node is an adjacent node of the node to be split and the node resource utilization of the sibling node does not reach the first utilization threshold, it is determined that the sibling node meets the migration condition.

[0020] In some embodiments, the node resource utilization of the sibling node satisfies the migration conditions for receiving data of the node to be split, including:

[0021] If the distance between the sibling node and the node to be split does not exceed one node and the node resource utilization of the sibling node does not reach the first utilization threshold, it is determined that the sibling node meets the migration condition.

[0022] In some embodiments, it further includes:

[0023] Monitor the frequency of writing to each leaf node of the B+ tree by calling the B+ tree creation script;

[0024] Calculate a first utilization threshold corresponding to the leaf node according to the write frequency;

[0025] The first utilization threshold is negatively correlated with the write frequency.

[0026] In some embodiments, the node resource utilization of the sibling node satisfies the migration conditions for receiving data of the node to be split, including:

[0027] If the average node space utilization of all sibling nodes of the node to be split does not exceed the second utilization threshold, it is determined that the sibling node meets the migration condition.

[0028] In some embodiments, monitoring that the B+ tree satisfies a node splitting condition includes:

[0029] It is monitored that the call to the B+ tree creation script calculates that the node to which the data to be written is to be written meets the node splitting condition, and the node to which the data to be written is to be written is determined to be the node to be split.

[0030] In some embodiments, the node to which the data to be written is to be written satisfies a node splitting condition, including:

[0031] The amount of data already written to the node to which the data to be written is the order of the B+ tree minus one.

[0032] In some embodiments, the node to which the data to be written is to be written satisfies a node splitting condition, including:

[0033] The amount of data already written to the node to which the data to be written is to be written reaches the order of the B+ tree.

[0034] In some embodiments, querying node resource utilization of sibling nodes of the node to be split includes:

[0035] Query the area where the data to be written is located in the node to be split;

[0036] If the location to be written is located in the left half of the node to be split, query the node resource utilization of the sibling node to the left of the node to be split;

[0037] If the location to be written is located in the right half of the node to be split, query the node resource utilization of the sibling node to the right of the node to be split;

[0038] The node resource utilization of the sibling node meets the migration conditions for receiving data from the node to be split, including:

[0039] The node resource utilization of the sibling nodes on the side where the right half area is located meets the migration conditions.

[0040] In some embodiments, querying node resource utilization of sibling nodes of the node to be split includes:

[0041] Query the area where the data to be written is located in the node to be split;

[0042] If the position to be written is located in the left half of the node to be split, the node resource utilization of the sibling node on the left side of the node to be split is queried. If the node resource utilization of the sibling node on the left side meets the migration conditions, the information of the sibling node on the left side and the information of the node to be split are written into the B+ tree creation script; if the node resource utilization of the sibling node on the left side does not meet the migration conditions, the node resource utilization of the sibling node on the right side of the node to be split is queried. If the node resource utilization of the sibling node on the right side meets the migration conditions, the information of the sibling node on the right side and the information of the node to be split are written into the B+ tree creation script; if both the node resource utilization of the sibling node on the left side of the node to be split and the node resource utilization of the sibling node on the right side of the node to be split do not meet the migration conditions, then the step of writing the information that does not meet the migration conditions into the B+ tree creation script is entered;

[0043] If the position to be written is located in the right half of the node to be split, the node resource utilization of the sibling node on the right side of the node to be split is queried. If the node resource utilization of the sibling node on the right side meets the migration conditions, the information of the sibling node on the right and the information of the node to be split are written into the B+ tree creation script; if the node resource utilization of the sibling node on the right side does not meet the migration conditions, the node resource utilization of the sibling node on the left side of the node to be split is queried. If the node resource utilization of the sibling node on the left side meets the migration conditions, the information of the sibling node on the left and the information of the node to be split are written into the B+ tree creation script; if the node resource utilization of the sibling node on the left side of the node to be split and the node resource utilization of the sibling node on the right side of the node to be split both do not meet the migration conditions, then the step of writing the information that does not meet the migration conditions into the B+ tree creation script is entered.

[0044] In some embodiments, monitoring that the B+ tree satisfies a node splitting condition includes:

[0045] It is monitored that there is a node to be split whose node resource utilization reaches a third utilization threshold in the B+ tree.

[0046] In some embodiments, monitoring that a node to be split exists in the B+ tree and that its node resource utilization reaches a third utilization threshold includes:

[0047] It is detected that the number of nodes to be split in the B+ tree has reached the order of the B+ tree minus one.

[0048] In some embodiments, monitoring that a node to be split exists in the B+ tree and that its node resource utilization reaches a third utilization threshold includes:

[0049] It is detected that the B+ tree has a node to be split whose number of written data has reached the order of the B+ tree.

[0050] In some embodiments, querying node resource utilization of sibling nodes of the node to be split includes:

[0051] Query the node resource distribution of all sibling nodes to the left of the node to be split and the node resource distribution of all sibling nodes to the right of the node to be split respectively;

[0052] If the node resource utilization of the sibling node meets the migration conditions for receiving the data of the node to be split, the information of the sibling node and the information of the node to be split are written into the B+ tree creation script, so that the B+ tree creation script is called to migrate the data of the node to be split to the sibling node, and then the parent node is updated according to the data information of the sibling node and the data information of the node to be split, including:

[0053] If the node resource distribution of a sibling node on one side meets the migration conditions, the node resource balancing algorithm is used to write the information of the node to be split and the information of the sibling node on one side that meets the migration conditions into the B+ tree creation script, so as to call the B+ tree creation script to perform node resource balancing processing on the node to be split and the sibling node on one side that meets the migration conditions.

[0054] In some embodiments, further comprising:

[0055] Monitor the node resource distribution of each leaf node of the B+ tree;

[0056] A node resource balancing algorithm is used to balance the node resources of each leaf node according to the node resource distribution of each leaf node.

[0057] In some embodiments, information about the sibling nodes and the node to be split is written into a B+ tree creation script, and after calling the B+ tree creation script to migrate the data of the node to be split to the sibling nodes, the parent node is updated according to the data information of the sibling nodes and the data information of the node to be split, including:

[0058] Determine the data to be migrated for the node to be split so that after the migration, neither the node to be split nor the sibling node to which it is migrated meets the node splitting condition;

[0059] Write the identifier of the node to be split, the identifier of the sibling node to be migrated, and the information of the data to be migrated into the B+ tree creation script, so as to call the B+ tree creation script to migrate the data to be migrated from the storage location corresponding to the node to be split to the storage location corresponding to the sibling node, and update the parent node according to the first data of the node to be split and the first data of the sibling node after migration.

[0060] In some embodiments, information that does not meet the migration condition is written into the B+ tree creation script, so that the B+ tree creation script performs a node splitting task for the node to be split, including:

[0061] The information that does not meet the migration conditions is written into the B+ tree creation script to call the B+ tree creation script to use the node splitting algorithm to calculate the data to be migrated of the node to be split and the position of the new node, migrate the data to be migrated to the temporary storage space and migrate the data to the new node after creating the new node, and then update the parent node according to the first data of the node to be split and the first data of the new node.

[0062] In some embodiments, applied to a distributed storage system, the data to be stored is metadata, and the B+ tree is stored separately from other data in the distributed storage system.

[0063] In some embodiments, the order of the B+ tree to be created is calculated based on the data block size of the device and the size of the unit data to be stored, including:

[0064] Run the node resource monitoring process and divide the device's data block size by the unit data size to be stored to obtain the order of the B+ tree to be created.

[0065] Write the order of the B+ tree to be created into the device's B+ tree creation script, and call the B+ tree creation script to create a B+ tree for storing the data to be stored, including:

[0066] Based on the node resource monitoring process, the order of the B+ tree to be created is written into the B+ tree creation process, the order of the B+ tree to be created is written into the B+ tree creation script, and the B+ tree creation process is run to call the B+ tree creation script to create the B+ tree;

[0067] When the B+ tree is detected to meet the node splitting conditions, the node resource utilization of the sibling nodes of the node to be split is queried, including:

[0068] It is detected that the node to which the data to be written is to be written satisfies the node splitting condition after calculation by calling the B+ tree creation script, and the node to which the data to be written is to be split is determined;

[0069] Query the node resource utilization of sibling nodes that are no more than one node away from the node to be split;

[0070] If the node resource utilization of the sibling node meets the migration conditions for receiving the data of the node to be split, the B+ tree creation script is prohibited from executing the node splitting task and the information of the sibling node and the information of the node to be split are written into the B+ tree creation script. After calling the B+ tree creation script to migrate the data of the node to be split to the sibling node, the parent node is updated according to the data information of the sibling node and the data information of the node to be split, including:

[0071] If the queried sibling node meets the migration conditions, the B+ tree creation script is prohibited from performing the node splitting task;

[0072] Determine the data to be migrated for the node to be split so that after the migration, neither the node to be split nor the sibling node to which it is migrated meets the node splitting condition;

[0073] Write the identifier of the node to be split, the identifier of the sibling node to be migrated, and the information of the data to be migrated into the B+ tree creation script, so as to call the B+ tree creation script to migrate the data to be migrated from the storage location corresponding to the node to be split to the storage location corresponding to the sibling node, and update the parent node according to the first data of the node to be split and the first data of the sibling node after the migration;

[0074] If the node resource utilization of all sibling nodes of the node to be split meets the migration conditions, the information that does not meet the migration conditions is written into the B+ tree creation script, so that the B+ tree creation script executes the node splitting task for the node to be split, including:

[0075] If none of the queried sibling nodes meet the migration conditions, the information that does not meet the migration conditions will be written into the B+ tree creation script, so as to call the B+ tree creation script to use the node splitting algorithm to calculate the data to be migrated of the node to be split and the position of the new node, migrate the data to be migrated to the temporary storage space and migrate the data to the new node after creating the new node, and then update the parent node according to the first data of the node to be split and the first data of the new node.

[0076] To solve the above technical problems, the present application further provides a data storage device, comprising:

[0077] A first calculation unit is used to calculate the order of the B+ tree to be created according to the data block size of the device and the size of the unit data to be stored;

[0078] A first writing unit is used to write the order of the B+ tree to be created into a B+ tree creation script of the device, so as to call the B+ tree creation script to create a B+ tree for storing the data to be stored;

[0079] Start the control unit, which is used to start the node resource monitoring program for the B+ tree;

[0080] The first query unit is configured to query node resource utilization of sibling nodes of the node to be split when it is detected that the B+ tree meets the node splitting condition;

[0081] The second writing unit is used to prohibit the B+ tree creation script from executing the node splitting task and write the information of the sibling nodes and the information of the node to be split into the B+ tree creation script if the node resource utilization of the sibling nodes meets the migration conditions for receiving the data of the node to be split, so as to call the B+ tree creation script to migrate the data of the node to be split to the sibling nodes, and then update the parent node according to the data information of the sibling nodes and the data information of the node to be split; if the node resource utilization of all the sibling nodes of the node to be split meets the migration conditions, the information that does not meet the migration conditions is written into the B+ tree creation script, so that the B+ tree creation script executes the node splitting task for the node to be split.

[0082] To solve the above technical problems, the present application further provides a data storage device, comprising:

[0083] memory for storing computer programs;

[0084] The processor is used to execute a computer program, and when the computer program is executed by the processor, the steps of any one of the above data storage methods are implemented.

[0085] In order to solve the above technical problems, the present application also provides a computer non-volatile readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of any of the above data storage methods are implemented.

[0086] The data storage method provided by the present application determines the order of the B+ tree to be created according to the data block size of the device and the size of the unit data to be stored, so as to call the B+ tree creation script to create a B+ tree with a reasonable order, and monitors the node resource utilization of the B+ tree. When the B+ tree meets the node splitting condition, it first queries whether the sibling nodes of the node to be split meet the migration condition for receiving the data of the node to be split. If so, the B+ tree creation script is prohibited from executing the node splitting task and the data of the node to be split is migrated to the sibling node. If not, the B+ tree is allowed to execute the node splitting task. Therefore, by delaying the node splitting of the B+ tree and preventing unnecessary node splitting of the B+ tree, the node resource utilization of the B+ tree storage is improved, thereby improving the storage resource utilization when the B+ tree is used for data storage.

[0087] This application predicts node splitting when writing data to a B+ tree and the node to be written meets the node splitting condition. At this time, the node splitting of the writing process of the B+ tree is delayed by monitoring whether the sibling nodes on the left and right sides of the node to be written meet the migration condition. While improving the node resource utilization of the B+ tree, the consumption of applying for new nodes caused by the node splitting task is reduced, thereby improving the writing efficiency.

[0088] This application predicts node splitting by monitoring whether the node resource utilization of each node of the B+ tree meets the node splitting condition. At this time, the B+ tree node split that will be triggered is delayed by monitoring whether the sibling nodes on the left and right sides of the node meet the migration condition, further avoiding write delays caused by triggering data migration or node splitting when new data is written to the B+ tree.

[0089] This application executes a node resource balancing algorithm by monitoring the data distribution of each leaf node of the B+ tree, and reduces the triggering of B+ tree node splitting by balancing the node resources of each leaf node.

[0090] The present application also provides a data storage device, equipment and computer non-volatile readable storage medium, which have the above-mentioned beneficial effects and are not described in detail here. BRIEF DESCRIPTION OF THE DRAWINGS

[0091] In order to more clearly illustrate the embodiments of the present application or the technical solutions of the prior art, the following is a brief introduction to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.

[0092] FIG1 is a flow chart of a data storage method provided in an embodiment of the present application;

[0093] FIG2 is a schematic diagram of an original B+ tree insertion operation;

[0094] FIG3 is a schematic diagram of an original B+ tree splitting operation;

[0095] FIG4 is a schematic diagram of a B+ tree suspended splitting process during an insertion operation provided by an embodiment of the present application;

[0096] FIG5 is a schematic structural diagram of a data storage device provided in an embodiment of the present application;

[0097] FIG6 is a schematic structural diagram of a data storage device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0098] The core of this application is to provide a data storage method, apparatus, device and computer non-volatile readable storage medium for improving storage resource utilization when using B+ tree for data storage.

[0099] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0100] For ease of understanding, the system architecture applicable to this application is first introduced. The specific implementation methods provided in the embodiments of this application can be applied to a single storage device or storage system, such as a distributed storage system. The data stored in the B+ tree used in the embodiments of this application can be metadata or other small data.

[0101] When using B+ trees to store metadata, dedicated space and disk arrays (RAID) are allocated in the storage system to separate B+ tree metadata storage from data reading and writing. This reduces metadata write amplification and bandwidth preemption through dedicated space, helping to improve performance.

[0102] Based on the above architecture, the data storage method provided in the embodiment of the present application is described below with reference to the accompanying drawings.

[0103] FIG1 is a flow chart of a data storage method provided in an embodiment of the present application.

[0104] As shown in FIG1 , the data storage method provided in the embodiment of the present application includes:

[0105] S101: Calculate the order of the B+ tree to be created according to the data block size of the device and the size of the unit data to be stored.

[0106] S102: Writing the order of the B+ tree to be created into a B+ tree creation script of the device, so as to call the B+ tree creation script to create a B+ tree for storing the data to be stored.

[0107] S103: Start a node resource monitoring program for the B+ tree.

[0108] S104: When it is detected that the B+ tree meets the node splitting condition, query the node resource utilization of the sibling nodes of the node to be split.

[0109] S105: If the node resource utilization of the sibling node meets the migration conditions for receiving the data of the node to be split, the B+ tree creation script is prohibited from executing the node splitting task and the information of the sibling node and the information of the node to be split are written into the B+ tree creation script, so as to call the B+ tree creation script to migrate the data of the node to be split to the sibling node, and then update the parent node according to the data information of the sibling node and the data information of the node to be split.

[0110] S106: If the node resource utilization of all sibling nodes of the node to be split meets the migration condition, the information that does not meet the migration condition is written into the B+ tree creation script, so that the B+ tree creation script executes the node splitting task for the node to be split.

[0111] In specific implementation, the data storage method provided by the embodiment of the present application can be applied to a single storage device or storage system, and can be specifically executed based on the central processing unit of the storage device. When a B+ tree is used to store data in a storage device, a read-write process is used to call a B+ tree creation script to create and manage the B+ tree. The data storage method provided by the embodiment of the present application does not require any changes to the original B+ tree creation script of the device, that is, there is no need to change the creation of the B+ tree, the node splitting algorithm, etc. Instead, the B+ tree creation and management process is monitored by pre-deploying a node resource monitoring program for the B+ tree, thereby preventing the original B+ tree creation script from performing excessive node splitting operations on the B+ tree.

[0112] If the data storage method provided in the embodiment of the present application is applied to a distributed storage system, and the data to be stored is metadata, the B+ tree can be stored separately from other data in the distributed storage system, that is, dedicated space and disk array (RAID) are allocated to the B+ tree, and the B+ tree metadata storage is separated from data reading and writing, so as to reduce metadata write amplification, bandwidth preemption, etc. through dedicated space, which helps to improve performance.

[0113] For S101 and S102, the B+ tree is the preferred data organization structure for block structure storage and is irreplaceable in the storage field. However, the selection of the order of the B+ tree (node ​​size) is also very important. The node size has a very large impact on write amplification and performance. Therefore, in the data storage method provided in the embodiment of the present application, the order of the B+ tree to be created is calculated based on the data block size of the device and the size of the unit data to be stored to provide the basis for the B+ tree creation script to create the B+ tree, so as to make rational use of the storage space of the block structure storage.

[0114] For S103, start the node resource monitoring program for the B+ tree, specifically list the created B+ tree into the monitoring object of the node resource monitoring program, and trigger the reading of the node resource utilization of the B+ tree or the periodic reading of the node resource utilization of the B+ tree when new data is written to the B+ tree.

[0115] In S104, monitoring that the B+ tree meets the node splitting condition may include: when monitoring that the B+ tree creation script is called to execute a write command, checking to determine whether the node to be written to the data to be written to the B+ tree meets the node splitting condition; and regularly monitoring the node resource utilization of each leaf node of the B+ tree, and detecting the existence of a leaf node that meets the node splitting condition. Wherein, the leaf node meeting the node resource utilization condition may be, based on the configuration of the B+ tree creation script, that the amount of data written to the leaf node is equal to the order of the B+ tree, or that the amount of data written to the leaf node is equal to the order of the B+ tree minus one.

[0116] When it is detected that the B+ tree meets the node splitting condition and the node splitting task is to be executed on the B+ tree according to the algorithm of the B+ tree creation script, the B+ tree creation script is temporarily suspended from executing the node splitting task, and the node resource utilization of the sibling nodes of the node to be split is first queried.

[0117] For S105, if the node resource utilization of the brother node allows receiving the data of the node to be split, that is, the migration conditions of one or more data of the node to be split are migrated to the brother node, then the node resource monitoring program sends a prohibition on the B+ tree creation script to execute the node splitting task to the selected brother node, and writes the information of the selected brother node and the information of the node to be split into the B+ tree creation script, so as to call the B+ tree creation script to migrate the data of the node to be split to the brother node, and then update the parent node according to the data information of the brother node and the data information of the node to be split.

[0118] According to the B+ tree creation rules, all data is stored in leaf nodes, and the data in each node is arranged from small to large. The parent node stores the first keyword of all its child nodes. Therefore, after performing data migration on the node to be split, it is also necessary to update the keyword in its corresponding parent node based on the first data of the node to be split and the first data of the migrated sibling node. If the first data of the parent node is involved, it is also necessary to recursively update it upward in the B+ tree until the root node.

[0119] Then, the information of the sibling node and the information of the node to be split are written into the B+ tree creation script, and the B+ tree creation script is called to migrate the data of the node to be split to the sibling node, and then the parent node is updated according to the data information of the sibling node and the data information of the node to be split. This may include: determining the data to be migrated of the node to be split so that after migration, neither the node to be split nor the migrated sibling node meets the node splitting condition; writing the identifier of the node to be split, the identifier of the migrated sibling node and the information of the data to be migrated into the B+ tree creation script, and calling the B+ tree creation script to migrate the data to be migrated from the storage location corresponding to the node to be split to the storage location corresponding to the sibling node, and updating the parent node according to the first data of the node to be split and the first data of the sibling node after migration.

[0120] The node resource utilization of the sibling node satisfies the migration condition for receiving data from the node to be split. This may include: if the sibling node is a neighboring node of the node to be split and the node resource utilization of the sibling node does not reach a first utilization threshold, then determining that the sibling node meets the migration condition. The first utilization threshold may be 75%. That is, if a sibling node with a node resource utilization of less than 75% exists among the left and right neighboring sibling nodes of the node to be split, the data of the node to be split may be migrated to the sibling node.

[0121] Alternatively, the node resource utilization of the sibling node satisfies the migration condition for receiving the data of the node to be split, which may also include: if the distance between the sibling node and the node to be split does not exceed one node and the node resource utilization of the sibling node does not reach a first utilization threshold, then the sibling node is determined to meet the migration condition. The first utilization threshold may be 75%. That is, if the data of the node to be split is considered to be migrated to a sibling node one node away, two node data migrations are required. Nodes that are further apart require more migrations. If the B+ tree creation script is executing a data writing task, it will take a long time to wait. Therefore, if there is no sibling node with a node resource utilization rate of less than 75% among the left and right adjacent sibling nodes of the node to be split, then the data of the node to be split is considered to be migrated to the sibling node one node away. If there is no sibling node with a node resource utilization rate of less than 75% among the sibling nodes one node away, then the sibling node of the node to be split is considered not to meet the migration condition.

[0122] Regarding the method for setting the first utilization threshold, the data storage method provided in the embodiment of the present application may further include: monitoring the write frequency of calling the B+ tree creation script to each leaf node of the B+ tree; calculating the first utilization threshold corresponding to the leaf node based on the write frequency; wherein the first utilization threshold is negatively correlated with the write frequency. In other words, the node resource utilization threshold corresponding to the migration condition can be determined based on the data write frequency. The higher the write frequency, the lower the node resource utilization threshold, thereby preventing the data migration of the node to be split from occupying the more popular sibling node.

[0123] Alternatively, the node resource utilization of the sibling nodes satisfies the migration conditions for receiving data of the node to be split, and may further include: if the average node space utilization of all sibling nodes of the node to be split does not exceed the second utilization threshold, then it is determined that the sibling node meets the migration conditions. That is to say, when the B+ tree creation script does not execute the write task on the B+ tree, the node resource utilization of the B+ tree can be monitored periodically. If there is a node to be split, it is not necessarily that the left and right adjacent sibling nodes of the node to be split meet the migration conditions, but the average node resource utilization of all sibling nodes does not exceed the second utilization threshold. At this time, the data of the sibling nodes can be migrated one by one, so that the data of the node with higher node resource utilization is migrated to the node with lower node resource utilization, so that the B+ tree achieves the effect of balanced node resource utilization when the write task is not executed.

[0124] Regarding S106, if the resources of the sibling node do not allow receiving the data of the node to be split, and the B+ tree is allowed to perform node classification, the information that does not meet the migration conditions is written into the B+ tree creation script, so that the B+ tree creation script executes the node splitting task for the node to be split. Specifically, writing the information that does not meet the migration conditions into the B+ tree creation script, so that the B+ tree creation script executes the node splitting task for the node to be split, can include: writing the information that does not meet the migration conditions into the B+ tree creation script, calling the B+ tree creation script to use the node splitting algorithm to calculate the data to be migrated of the node to be split and the position of the new node, migrating the data to be migrated to a temporary storage space, and migrating the data to be migrated to the new node after creating the new node, and then updating the parent node according to the first data of the node to be split and the first data of the new node.

[0125] Since the node resource monitoring program detects that the B+ tree meets the node splitting conditions, it is not necessarily the node where the B+ tree creation script is executed. Therefore, after the information that does not meet the migration conditions is written into the B+ tree creation script, if the conditions for the B+ tree creation script to execute the node splitting task are met, the B+ tree creation script will directly execute the node splitting task.

[0126] The data storage method provided by the embodiment of the present application determines the order of the B+ tree to be created based on the data block size of the device and the size of the unit data to be stored to call the B+ tree creation script to create a B+ tree with a reasonable order, and monitors the node resource utilization of the B+ tree. When the B+ tree meets the node splitting condition, it first queries whether the brother node of the node to be split meets the migration condition for receiving the data of the node to be split. If so, the B+ tree creation script is prohibited from executing the node splitting task and the data of the node to be split is migrated to the brother node. If not, the B+ tree is allowed to execute the node splitting task, thereby improving the node resource utilization of the B+ tree storage by delaying the node splitting of the B+ tree and preventing unnecessary node splitting of the B+ tree, thereby improving the storage resource utilization when the B+ tree is used for data storage.

[0127] In the above embodiment, when new data is written to a B+ tree, if the node to be written is full or nearly full, node sorting of the B+ tree is triggered according to the node splitting algorithm of the B+ tree creation script. Monitoring that the B+ tree satisfies the node splitting condition in S104 may include: monitoring that the node to which the data to be written is to be written calculated by calling the B+ tree creation script satisfies the node splitting condition, and determining that the node to which the data to be written is to be written is the node to be split.

[0128] Specifically, the node to which the data to be written is to be written satisfies a node splitting condition, which may include: the amount of data already written to the node to which the data to be written is equal to the order of the B+ tree minus one. In other words, when writing new data to the B+ tree, if the location to be written is almost full, a node split of the B+ tree is triggered.

[0129] Alternatively, the node to which the data to be written is to be written satisfies the node splitting condition, further comprising: the amount of data already written to the node to which the data to be written is to be written reaches the order of the B+ tree. In other words, when writing new data to the B+ tree, if the location to be written is full, a node split of the B+ tree is triggered.

[0130] When a node split is triggered by writing data to a B+ tree, querying the node resource utilization of the sibling nodes of the node to be split in S104 may include: querying the area of ​​the node to be split where the data to be written is located; if the location to be written is in the left half of the node to be split, querying the node resource utilization of the sibling nodes to the left of the node to be split; if the location to be written is in the right half of the node to be split, querying the node resource utilization of the sibling nodes to the right of the node to be split; the node resource utilization of the sibling nodes meeting the migration conditions for receiving data from the node to be split includes: the node resource utilization of the sibling nodes on the side of the right half of the node to be split meeting the migration conditions. In other words, the location of the sibling nodes to be queried can be determined based on the location of the data to be written in the node to be split. By comparing the size of the data to be written with the intermediate data of the node to be split, if the data to be written is smaller than the intermediate data of the node to be split, the location to be written of the data to be written is determined to be in the left half of the node to be split; if the data to be written is larger than the intermediate data of the node to be split, the location to be written of the data to be written is determined to be in the right half of the node to be split.

[0131] Alternatively, querying the node resource utilization of the sibling nodes of the node to be split in S104 may include: querying the area where the to-be-written data is located in the to-be-split position in the node to be split; if the to-be-written position is located in the left half of the node to be split, querying the node resource utilization of the sibling nodes on the left side of the node to be split, and if the node resource utilization of the sibling nodes on the left side meets the migration conditions, writing the information of the sibling nodes on the left side and the information of the node to be split into the B+ tree creation script; if the node resource utilization of the sibling nodes on the left side does not meet the migration conditions, querying the node resource utilization of the sibling nodes on the right side of the node to be split, and if the node resource utilization of the sibling nodes on the right side meets the migration conditions, writing the information of the sibling nodes on the right side and the information of the node to be split into the B+ tree creation script; if the node resource utilization of the sibling nodes on the left side of the node to be split and the node resource utilization of the sibling nodes on the right side of the node to be split do not meet the migration conditions, If the node to be split does not meet the migration condition, the step of writing the information that does not meet the migration condition into the B+ tree creation script is entered. If the location to be written is located in the right half of the node to be split, the node resource utilization of the sibling node on the right side of the node to be split is queried. If the node resource utilization of the sibling node on the right side meets the migration condition, the information of the sibling node on the right side and the information of the node to be split are written into the B+ tree creation script. If the node resource utilization of the sibling node on the right side does not meet the migration condition, the node resource utilization of the sibling node on the left side of the node to be split is queried. If the node resource utilization of the sibling node on the left side meets the migration condition, the information of the sibling node on the left side and the information of the node to be split are written into the B+ tree creation script. If the node resource utilization of the sibling node on the left side of the node to be split and the node resource utilization of the sibling node on the right side of the node to be split do not meet the migration condition, the step of writing the information that does not meet the migration condition into the B+ tree creation script is entered. In other words, the location of the data to be written in the node to be written can be used to determine which sibling node on the side to be queried first. If the sibling node on that side does not meet the migration condition, the sibling node on the other side is queried.

[0132] When no new data is written to the B+ tree, the B+ tree creation script will not be triggered to execute the node splitting task. However, the node splitting task may be triggered the next time new data is written to the B+ tree. Therefore, monitoring that the B+ tree meets the node splitting condition in S104 may include: monitoring that the B+ tree has a node to be split whose node resource utilization reaches a third utilization threshold.

[0133] Among them, the third utilization threshold can be 75%, or calculated based on the write frequency of each leaf node of the B+ tree by calling the B+ tree creation script, so that the third utilization threshold is negatively correlated with the write frequency of the corresponding leaf node.

[0134] Alternatively, detecting that a node to be split exists in a B+ tree whose node resource utilization reaches a third utilization threshold may include detecting that the number of written data in the B+ tree reaches the order of the B+ tree minus one. Alternatively, detecting that a node to be split exists in a B+ tree whose node resource utilization reaches the third utilization threshold may include detecting that the number of written data in the B+ tree reaches the order of the B+ tree.

[0135] Based on this, querying the node resource utilization of the sibling nodes of the node to be split in S104 may include: querying the node resource distribution of all sibling nodes to the left of the node to be split and the node resource distribution of all sibling nodes to the right of the node to be split; if the node resource utilization of the sibling nodes meets the migration condition for receiving the data of the node to be split, writing the information of the sibling nodes and the information of the node to be split into a B+ tree creation script, invoking the B+ tree creation script to migrate the data of the node to be split to the sibling nodes, and then updating the parent node based on the data information of the sibling nodes and the data information of the node to be split. This may include: if the node resource utilization distribution of the sibling nodes on one side meets the migration condition, writing the information of the node to be split and the information of the sibling nodes on the one side that meet the migration condition into the B+ tree creation script using a node resource balancing algorithm, invoking the B+ tree creation script to perform node resource balancing processing on the node to be split and the sibling nodes on the one side that meet the migration condition. In other words, when no new data is written to the B+ tree, by monitoring the distribution of the sibling nodes on the left and right of the node to be split, it is determined whether to migrate the node to be split and its sibling nodes on the left or the node to be split and its sibling nodes on the right.

[0136] Alternatively, the data storage method provided in the embodiment of the present application may further include: monitoring the node resource distribution of each leaf node of the B+ tree; and using a node resource balancing algorithm to perform node resource balancing processing on each leaf node according to the node resource distribution of each leaf node.

[0137] That is to say, when no new data is written into the B+ tree and no node meets the node splitting conditions, the node resource distribution of each leaf node of the B+ tree can be monitored regularly, and the data of the nodes with higher node resource utilization can be migrated to the nodes with lower node resource utilization. After the migration, the parent node is updated accordingly until the root node.

[0138] FIG2 is a schematic diagram of an original B+ tree insertion operation; FIG3 is a schematic diagram of an original B+ tree splitting operation; and FIG4 is a schematic diagram of suspending splitting during a B+ tree insertion operation provided by an embodiment of the present application.

[0139] Based on the above embodiment, the embodiment of the present application provides an explanation of a data storage scenario when writing new data to a B+ tree.

[0140] In the data storage method provided in an embodiment of the present application, S101 calculates the order of the B+ tree to be created based on the data block size of the device and the size of the unit data to be stored, which can include: running a node resource monitoring process, dividing the data block size of the device by the size of the unit data to be stored, and obtaining the order of the B+ tree to be created.

[0141] In S102, the order of the B+ tree to be created is written into the B+ tree creation script of the device to call the B+ tree creation script to create a B+ tree for storing the data to be stored. This may include: based on the node resource monitoring process, writing the order of the B+ tree to be created into the B+ tree creation process to write the order of the B+ tree to be created into the B+ tree creation script, running the B+ tree creation process to call the B+ tree creation script to create the B+ tree.

[0142] In S104, when it is monitored that the B+ tree meets the node splitting condition, the node resource utilization of the sibling nodes of the node to be split is queried, which may include: monitoring that the node to which the data to be written is to be written calculated by calling the B+ tree creation script meets the node splitting condition, and determining that the node to which the data to be written is to be written is the node to be split; and querying the node resource utilization of the sibling nodes that are no more than one node away from the node to be split.

[0143] In S105, if the node resource utilization of the sibling node meets the migration condition for receiving the data of the node to be split, the B+ tree creation script is prohibited from executing the node splitting task and the information of the sibling node and the information of the node to be split is written into the B+ tree creation script, so as to call the B+ tree creation script to migrate the data of the node to be split to the sibling node, and then update the parent node according to the data information of the sibling node and the data information of the node to be split. It can include: if the queried sibling node meets the migration condition, prohibiting the B+ tree creation script from executing the node splitting task; determining the data to be migrated of the node to be split so that after migration, neither the node to be split nor the sibling node to which it is migrated meets the node splitting condition; writing the identifier of the node to be split, the identifier of the sibling node to which it is migrated, and the information of the data to be migrated into the B+ tree creation script, so as to call the B+ tree creation script to migrate the data to be migrated from the storage location corresponding to the node to be split to the storage location corresponding to the sibling node, and updating the parent node according to the first data of the node to be split and the first data of the sibling node after migration.

[0144] In S106, if the node resource utilization of all sibling nodes of the node to be split meets the migration conditions, the information that does not meet the migration conditions is written into the B+ tree creation script, so that the B+ tree creation script executes the node splitting task for the node to be split, which may include: if none of the queried sibling nodes meet the migration conditions, the information that does not meet the migration conditions is written into the B+ tree creation script, so as to call the B+ tree creation script to use the node splitting algorithm to calculate the data to be migrated of the node to be split and the position of the new node, migrate the data to be migrated to a temporary storage space and migrate the data to be migrated to the new node after creating the new node, and then update the parent node according to the first data of the node to be split and the first data of the new node.

[0145] Assume that the calculated B+ tree has a 6-order order, meaning each node can store a maximum of six data items. Before writing new data, assume that a portion of this 6-order B+ tree is shown in Figure 2, where node 1 is the parent node of nodes 2, 3, and 4. When new data 8 needs to be inserted, as shown in Figure 3, this new data 8 needs to be inserted into node 3. According to the node splitting conditions of the B+ tree creation script, node 3 must be split. At this point, some of the data in node 3 (such as 9, 11, and 12) must be migrated to temporary storage space. A new node 5 is then requested and initialized. 9, 11, and 12 are then migrated from the temporary storage space to node 5. The parent node must be updated, and the first data item in node 5 must be inserted into the parent node. Recursion is performed, if necessary, all the way back to the root node. As can be seen, if the B+ tree creation script is not intervened, two memory copies and one node creation operation are performed when writing a new node and triggering a node split. In extreme cases, node resource utilization can be as low as 50%, wasting valuable memory and disk space.

[0146] However, using the data storage method provided by the embodiment of the present application, when new data 8 needs to be inserted, as shown in Figure 4, the new data 8 needs to be inserted into node 3. According to the node splitting condition of the B+ tree creation script, node 3 needs to be split. At this time, the call to the B+ tree creation script to perform the node splitting task on node 3 is temporarily postponed. Since node 8 is located in the left half of node 3, the node 2 on the left side of node 3 is queried and it is found that node 2 meets the migration condition. Then, data 6 can be migrated from node 3 to the last position of the data written to node 2, and data 8 can be inserted into node 3. At this time, the first data of node 3 has changed, so its parent node needs to be updated. Therefore, only one memory copy is required, and no node creation operation needs to be performed, which not only speeds up the process of writing data to the B+ tree, but also saves node resources.

[0147] By applying the data storage method provided in the embodiment of the present application, by temporarily suspending the splitting of the B+ tree, the node resource utilization of the B+ tree can be significantly improved, and even the node resource utilization of the B+ tree can be increased to 100%, thereby improving the overall throughput of the system; this method can obtain efficient metadata access and increase the efficiency of concurrent queries. If the node resource balancing algorithm introduced in the above embodiment of the present application is not adopted, and only the data of the node to be split is migrated to the adjacent sibling node, according to the data distribution law of the B+ tree, the node resource utilization after improvement can be increased by an average of 15%. If the node resource balancing algorithm introduced in the above embodiment of the present application is adopted, the node resource utilization of the B+ tree can be further improved.

[0148] The above describes in detail various embodiments corresponding to the data storage method. On this basis, the present application also discloses data storage devices, equipment and computer non-volatile readable storage media corresponding to the above method.

[0149] FIG5 is a schematic structural diagram of a data storage device provided in an embodiment of the present application.

[0150] As shown in FIG5 , the data storage device provided in the embodiment of the present application includes:

[0151] The first calculation unit 501 is used to calculate the order of the B+ tree to be created according to the data block size of the device and the size of the unit data to be stored;

[0152] A first writing unit 502 is configured to write the order of the B+ tree to be created into a B+ tree creation script of the device, so as to call the B+ tree creation script to create a B+ tree for storing the data to be stored;

[0153] Start the control unit 503, which is used to start the node resource monitoring program for the B+ tree;

[0154] The first query unit 504 is configured to query the node resource utilization of the sibling nodes of the node to be split when it is detected that the B+ tree meets the node splitting condition;

[0155] The second writing unit 505 is used to prohibit the B+ tree creation script from executing the node splitting task and write the information of the sibling nodes and the information of the node to be split into the B+ tree creation script if the node resource utilization of the sibling nodes meets the migration conditions for receiving the data of the node to be split, so as to call the B+ tree creation script to migrate the data of the node to be split to the sibling nodes, and then update the parent node according to the data information of the sibling nodes and the data information of the node to be split; if the node resource utilization of all the sibling nodes of the node to be split meets the migration conditions, the information that does not meet the migration conditions is written into the B+ tree creation script, so that the B+ tree creation script executes the node splitting task for the node to be split.

[0156] In some embodiments, the node resource utilization of the sibling node satisfies the migration conditions for receiving data of the node to be split, including:

[0157] If the sibling node is an adjacent node of the node to be split and the node resource utilization of the sibling node does not reach the first utilization threshold, it is determined that the sibling node meets the migration condition.

[0158] In other implementations, the node resource utilization of the sibling node satisfies the migration conditions for receiving data of the node to be split, including:

[0159] If the distance between the sibling node and the node to be split does not exceed one node and the node resource utilization of the sibling node does not reach the first utilization threshold, it is determined that the sibling node meets the migration condition.

[0160] In some embodiments, the data storage device provided by the embodiments of the present application may further include:

[0161] The second query unit is used to monitor the writing frequency of calling the B+ tree creation script to each leaf node of the B+ tree;

[0162] A second calculation unit, configured to calculate a first utilization threshold corresponding to the leaf node according to the write frequency;

[0163] The first utilization threshold is negatively correlated with the write frequency.

[0164] In other implementations, the node resource utilization of the sibling node satisfies the migration conditions for receiving data of the node to be split, including:

[0165] If the average node space utilization of all sibling nodes of the node to be split does not exceed the second utilization threshold, it is determined that the sibling node meets the migration condition.

[0166] In some embodiments, monitoring that the B+ tree satisfies a node splitting condition includes:

[0167] It is monitored that the call to the B+ tree creation script calculates that the node to which the data to be written is to be written meets the node splitting condition, and the node to which the data to be written is to be written is determined to be the node to be split.

[0168] In some embodiments, the node to which the data to be written is to be written satisfies a node splitting condition, including:

[0169] The amount of data already written to the node to which the data to be written is the order of the B+ tree minus one.

[0170] In some other implementations, the node to which the data to be written is to be written satisfies a node split condition, including:

[0171] The amount of data already written to the node to which the data to be written is to be written reaches the order of the B+ tree.

[0172] In some embodiments, the first query unit 504 queries the node resource utilization of the sibling nodes of the node to be split, including:

[0173] Query the area where the data to be written is located in the node to be split;

[0174] If the location to be written is located in the left half of the node to be split, query the node resource utilization of the sibling node to the left of the node to be split;

[0175] If the location to be written is located in the right half of the node to be split, query the node resource utilization of the sibling node to the right of the node to be split;

[0176] The node resource utilization of the sibling node meets the migration conditions for receiving data from the node to be split, including:

[0177] The node resource utilization of the sibling nodes on the side where the right half area is located meets the migration conditions.

[0178] In some other implementations, the first query unit 504 queries the node resource utilization of the sibling nodes of the node to be split, including:

[0179] Query the area where the data to be written is located in the node to be split;

[0180] If the position to be written is located in the left half of the node to be split, the node resource utilization of the sibling node on the left side of the node to be split is queried. If the node resource utilization of the sibling node on the left side meets the migration conditions, the information of the sibling node on the left side and the information of the node to be split are written into the B+ tree creation script; if the node resource utilization of the sibling node on the left side does not meet the migration conditions, the node resource utilization of the sibling node on the right side of the node to be split is queried. If the node resource utilization of the sibling node on the right side meets the migration conditions, the information of the sibling node on the right side and the information of the node to be split are written into the B+ tree creation script; if both the node resource utilization of the sibling node on the left side of the node to be split and the node resource utilization of the sibling node on the right side of the node to be split do not meet the migration conditions, then the step of writing the information that does not meet the migration conditions into the B+ tree creation script is entered;

[0181] If the position to be written is located in the right half of the node to be split, the node resource utilization of the sibling node on the right side of the node to be split is queried. If the node resource utilization of the sibling node on the right side meets the migration conditions, the information of the sibling node on the right and the information of the node to be split are written into the B+ tree creation script; if the node resource utilization of the sibling node on the right side does not meet the migration conditions, the node resource utilization of the sibling node on the left side of the node to be split is queried. If the node resource utilization of the sibling node on the left side meets the migration conditions, the information of the sibling node on the left and the information of the node to be split are written into the B+ tree creation script; if the node resource utilization of the sibling node on the left side of the node to be split and the node resource utilization of the sibling node on the right side of the node to be split both do not meet the migration conditions, then the step of writing the information that does not meet the migration conditions into the B+ tree creation script is entered.

[0182] In some other implementations, the B+ tree is monitored to meet the node splitting condition, including:

[0183] It is monitored that there is a node to be split whose node resource utilization reaches a third utilization threshold in the B+ tree.

[0184] In some embodiments, monitoring that a node to be split exists in the B+ tree and that its node resource utilization reaches a third utilization threshold includes:

[0185] It is detected that the number of nodes to be split in the B+ tree has reached the order of the B+ tree minus one.

[0186] In some other implementations, monitoring that a node to be split exists in a B+ tree and that its node resource utilization reaches a third utilization threshold includes:

[0187] It is detected that the B+ tree has a node to be split whose number of written data has reached the order of the B+ tree.

[0188] In some embodiments, querying node resource utilization of sibling nodes of the node to be split includes:

[0189] Query the node resource distribution of all sibling nodes to the left of the node to be split and the node resource distribution of all sibling nodes to the right of the node to be split respectively;

[0190] If the node resource utilization of the sibling node meets the migration conditions for receiving the data of the node to be split, the information of the sibling node and the information of the node to be split are written into the B+ tree creation script, so that the B+ tree creation script is called to migrate the data of the node to be split to the sibling node, and then the parent node is updated according to the data information of the sibling node and the data information of the node to be split, including:

[0191] If the node resource distribution of a sibling node on one side meets the migration conditions, the node resource balancing algorithm is used to write the information of the node to be split and the information of the sibling node on one side that meets the migration conditions into the B+ tree creation script, so as to call the B+ tree creation script to perform node resource balancing processing on the node to be split and the sibling node on one side that meets the migration conditions.

[0192] In some embodiments, the data storage device provided by the embodiments of the present application further includes:

[0193] The third query unit is used to monitor the node resource distribution of each leaf node of the B+ tree;

[0194] The balancing processing unit is used to adopt a node resource balancing algorithm to perform node resource balancing processing on each leaf node according to the node resource distribution of each leaf node.

[0195] In some embodiments, the second writing unit 505 writes the information of the sibling nodes and the information of the node to be split into the B+ tree creation script, and calls the B+ tree creation script to migrate the data of the node to be split to the sibling nodes, and then updates the parent node according to the data information of the sibling nodes and the data information of the node to be split, including:

[0196] Determine the data to be migrated for the node to be split so that after the migration, neither the node to be split nor the sibling node to which it is migrated meets the node splitting condition;

[0197] Write the identifier of the node to be split, the identifier of the sibling node to be migrated, and the information of the data to be migrated into the B+ tree creation script, so as to call the B+ tree creation script to migrate the data to be migrated from the storage location corresponding to the node to be split to the storage location corresponding to the sibling node, and update the parent node according to the first data of the node to be split and the first data of the sibling node after migration.

[0198] In some embodiments, the second writing unit 505 writes the information that does not meet the migration condition into the B+ tree creation script, so that the B+ tree creation script performs the node splitting task for the node to be split, including:

[0199] The information that does not meet the migration conditions is written into the B+ tree creation script to call the B+ tree creation script to use the node splitting algorithm to calculate the data to be migrated of the node to be split and the position of the new node, migrate the data to be migrated to the temporary storage space and migrate the data to the new node after creating the new node, and then update the parent node according to the first data of the node to be split and the first data of the new node.

[0200] In some embodiments, the data storage device provided by the embodiments of the present application is applied to a distributed storage system, the data to be stored is metadata, and the B+ tree is stored separately from other data in the distributed storage system.

[0201] In some embodiments, the first calculation unit 501 calculates the order of the B+ tree to be created based on the data block size of the device and the size of the unit data to be stored, including:

[0202] Run the node resource monitoring process and divide the device's data block size by the unit data size to be stored to obtain the order of the B+ tree to be created.

[0203] The first writing unit 502 writes the order of the B+ tree to be created into the B+ tree creation script of the device, so as to call the B+ tree creation script to create a B+ tree for storing the data to be stored, including:

[0204] Based on the node resource monitoring process, the order of the B+ tree to be created is written into the B+ tree creation process, the order of the B+ tree to be created is written into the B+ tree creation script, and the B+ tree creation process is run to call the B+ tree creation script to create the B+ tree;

[0205] When the first query unit 504 detects that the B+ tree meets the node splitting condition, it queries the node resource utilization of the sibling nodes of the node to be split, including:

[0206] It is detected that the node to which the data to be written is to be written satisfies the node splitting condition after calculation by calling the B+ tree creation script, and the node to which the data to be written is to be split is determined;

[0207] Query the node resource utilization of sibling nodes that are no more than one node away from the node to be split;

[0208] If the node resource utilization of the sibling node meets the migration condition for receiving the data of the node to be split, the second writing unit 505 prohibits the B+ tree creation script from executing the node splitting task and writes the information of the sibling node and the information of the node to be split into the B+ tree creation script, so as to call the B+ tree creation script to migrate the data of the node to be split to the sibling node, and then update the parent node according to the data information of the sibling node and the data information of the node to be split, including:

[0209] If the queried sibling node meets the migration conditions, the B+ tree creation script is prohibited from performing the node splitting task;

[0210] Determine the data to be migrated for the node to be split so that after the migration, neither the node to be split nor the sibling node to which it is migrated meets the node splitting condition;

[0211] Write the identifier of the node to be split, the identifier of the sibling node to be migrated, and the information of the data to be migrated into the B+ tree creation script, so as to call the B+ tree creation script to migrate the data to be migrated from the storage location corresponding to the node to be split to the storage location corresponding to the sibling node, and update the parent node according to the first data of the node to be split and the first data of the sibling node after the migration;

[0212] If the node resource utilization of all sibling nodes of the node to be split meets the migration condition, the second writing unit 505 writes the information that does not meet the migration condition into the B+ tree creation script, so that the B+ tree creation script executes the node splitting task for the node to be split, including:

[0213] If none of the queried brother nodes meet the migration conditions, the information that does not meet the migration conditions will be written into the B+ tree creation script, so as to call the B+ tree creation script to use the node splitting algorithm to calculate the data to be migrated of the node to be split and the position of the new node, migrate the data to be migrated to the temporary storage space and migrate the data to the new node after creating the new node, and then update the parent node according to the first data of the node to be split and the first data of the new node.

[0214] Since the embodiments of the apparatus part correspond to the embodiments of the method part, please refer to the description of the embodiments of the method part for the embodiments of the apparatus part, and they will not be repeated here.

[0215] FIG6 is a schematic structural diagram of a data storage device provided in an embodiment of the present application.

[0216] As shown in FIG6 , the data storage device provided in the embodiment of the present application includes:

[0217] Memory 610, for storing computer programs 611;

[0218] The processor 620 is configured to execute the computer program 611 , which, when executed by the processor 620 , implements the steps of the data storage method according to any one of the above embodiments.

[0219] Among them, the processor 620 may include one or more processing cores, such as a 3-core processor, an 8-core processor, etc. The processor 620 can be implemented in at least one hardware form of digital signal processing DSP (Digital Signal Processing), field programmable gate array FPGA (Field-Programmable Gate Array), and programmable logic array PLA (Programmable Logic Array). The processor 620 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the awake state, also known as the central processing unit CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 620 may be integrated with a graphics processing unit GPU (Graphics Processing Unit), which is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 620 may also include an artificial intelligence AI (Artificial Intelligence) processor, which is used to process computing operations related to machine learning.

[0220] The memory 610 may include one or more computer non-volatile readable storage media, which may be non-transitory. The memory 610 may also include a high-speed random access memory, and a non-volatile memory, such as one or more disk storage devices, flash memory storage devices. In this embodiment, the memory 610 is at least used to store the following computer program 611, wherein, after the computer program 611 is loaded and executed by the processor 620, it can implement the relevant steps in the data storage method disclosed in any of the aforementioned embodiments. In addition, the resources stored in the memory 610 may also include an operating system 612 and data 613, etc., and the storage method may be temporary storage or permanent storage. Among them, the operating system 612 may be Windows. The data 613 may include but is not limited to the data involved in the above method.

[0221] In some embodiments, the data storage device may further include a display screen 630 , a power supply 640 , a communication interface 650 , an input / output interface 660 , a sensor 670 , and a communication bus 680 .

[0222] Those skilled in the art will appreciate that the structure shown in FIG. 6 does not limit the data storage device and may include more or fewer components than shown in the figure.

[0223] The data storage device provided in the embodiment of the present application includes a memory and a processor. When the processor executes the program stored in the memory, it can implement the above data storage method and the effect is the same as above.

[0224] It should be noted that the above-described embodiments of the apparatus and equipment are merely illustrative. For example, the division of modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces, indirect coupling or communication connection of the apparatus or module, which may be electrical, mechanical or other forms. The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical modules, that is, they may be located in one place, or they may be distributed on multiple network modules. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.

[0225] In addition, the functional modules in the various embodiments of the present application may be integrated into a processing module, or each module may exist physically separately, or two or more modules may be integrated into a single module. The above-mentioned integrated modules may be implemented in the form of hardware or software functional modules.

[0226] If the integrated module is implemented in the form of a software functional module and sold or used as an independent product, it can be stored in a non-volatile computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the existing technology, or all or part of the technical solution can be embodied in the form of a software product, which is stored in a storage medium and executes all or part of the steps of the various embodiments of the present application.

[0227] To this end, an embodiment of the present application further provides a computer non-volatile readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the data storage method are implemented.

[0228] The computer non-volatile readable storage medium may include: a USB flash drive, a mobile hard disk, a read-only memory ROM (Read-Only Memory), a random access memory RAM (Random Access Memory), a magnetic disk or an optical disk, and other media that can store program codes.

[0229] The computer program contained in the computer non-volatile readable storage medium provided in this embodiment can implement the steps of the above data storage method when executed by a processor, and the effect is the same as above.

[0230] The above is a detailed introduction to a data storage method, device, equipment and computer non-volatile readable storage medium provided by the present application. The various embodiments in the specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. The same and similar parts between the various embodiments can be referred to each other. For the devices, equipment and computer non-volatile readable storage media disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the description is relatively simple, and the relevant parts can be referred to the method part description. It should be pointed out that for ordinary technicians in this technical field, without departing from the principles of this application, several improvements and modifications can be made to this application, and these improvements and modifications also fall within the scope of protection of the claims of this application.

[0231] It should also be noted that, in this specification, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus comprising the element.

Claims

1. A data storage method, characterized in that: include: The order of the balanced multi-way search tree to be created is calculated according to the data block size of the device and the size of the unit data to be stored; Writing the order of the balanced multi-way search tree to be created into a balanced multi-way search tree creation script of the device, so as to call the balanced multi-way search tree creation script to create a balanced multi-way search tree for storing the data to be stored; Starting a node resource monitoring program for the balanced multi-way search tree; When it is monitored that the balanced multi-way search tree meets the node splitting condition, querying the node resource utilization of the brother nodes of the node to be split; If the node resource utilization of the brother node meets the migration condition of receiving the data of the node to be split, the balanced multi-way search tree creation script is prohibited from executing the node splitting task and the information of the brother node and the information of the node to be split are written into the balanced multi-way search tree creation script, so as to call the balanced multi-way search tree creation script to migrate the data of the node to be split to the brother node, and then update the parent node according to the data information of the brother node and the data information of the node to be split; If the node resource utilization of all the brother nodes of the node to be split does not meet the migration condition, the information that does not meet the migration condition is written into the balanced multi-way search tree creation script, so that the balanced multi-way search tree creation script executes the node splitting task for the node to be split.

2. The data storage method according to claim 1, characterized in that: The step of calculating the order of the balanced multi-way search tree to be created according to the data block size of the device and the size of the unit data to be stored includes: Running a node resource monitoring process, dividing the data block size of the device by the size of the unit data to be stored, to obtain the order of the balanced multi-way search tree to be created; Monitoring that the balanced multi-way search tree satisfies the node splitting condition includes: When it is detected that the balanced multi-way search tree creation script is called to execute a write command, it is checked to determine that the node to be written of the data to be written in the balanced multi-way search tree meets the node splitting condition; the node resource utilization of each leaf node of the balanced multi-way search tree is monitored at regular intervals, and the existence of the leaf node that meets the node splitting condition is detected.

3. The data storage method according to claim 1, characterized in that: The node resource utilization of the brother node satisfies the migration condition of receiving the data of the node to be split, including: If the brother node is an adjacent node of the node to be split and the node resource utilization of the brother node does not reach a first utilization threshold, it is determined that the brother node meets the migration condition.

4. The data storage method according to claim 1, characterized in that: The node resource utilization of the brother node satisfies the migration condition of receiving the data of the node to be split, including: If the distance between the brother node and the node to be split does not exceed one node and the node resource utilization of the brother node does not reach a first utilization threshold, it is determined that the brother node meets the migration condition.

5. The data storage method according to claim 3 or 4, characterized in that: Also includes: Monitoring the writing frequency of calling the balanced multi-way search tree creation script to each leaf node of the balanced multi-way search tree; Calculate the first utilization threshold corresponding to the leaf node according to the write frequency; The first utilization threshold is negatively correlated with the write frequency.

6. The data storage method according to claim 1, characterized in that: The node resource utilization of the brother node satisfies the migration condition of receiving the data of the node to be split, including: If the average node space utilization of all the sibling nodes of the node to be split does not exceed the second utilization threshold, it is determined that the sibling nodes meet the migration condition.

7. The data storage method according to claim 1, characterized in that: The monitoring that the balanced multi-way search tree satisfies a node splitting condition includes: It is monitored that the node into which the data to be written is calculated by calling the balanced multi-way search tree creation script satisfies the node splitting condition, and the node into which the data to be written is determined to be the node to be split.

8. The data storage method according to claim 7, characterized in that: The node to which the to-be-written data is to be written satisfies the node splitting condition, including: The amount of data already written into the node into which the data to be written is the order of the balanced multi-way search tree minus one.

9. The data storage method according to claim 7, characterized in that: The node to which the data to be written is to be written satisfies the node splitting condition, including: The amount of data already written into the node into which the data to be written is to be written reaches the order of the balanced multi-way search tree.

10. The data storage method according to claim 7, characterized in that: The querying of node resource utilization of the sibling nodes of the node to be split includes: Querying the area where the to-be-written data is located in the to-be-split node; If the position to be written is located in the left half of the node to be split, querying the node resource utilization of the brother node on the left side of the node to be split; If the position to be written is located in the right half of the node to be split, querying the node resource utilization of the brother node on the right side of the node to be split; The node resource utilization of the brother node satisfies the migration condition for receiving the data of the node to be split, including: The node resource utilization of the brother node on the side where the right half area is located meets the migration condition.

11. The data storage method according to claim 7, characterized in that: The querying of node resource utilization of the sibling nodes of the node to be split includes: Querying the area where the to-be-written data is located in the to-be-split node; If the position to be written is located in the left half area of ​​the node to be split, query the node resource utilization of the brother node on the left side of the node to be split, if the node resource utilization of the brother node on the left side meets the migration condition, write the information of the brother node on the left side and the information of the node to be split into the balanced multi-way search tree creation script; if the node resource utilization of the brother node on the left side does not meet the migration condition, query the node resource utilization of the brother node on the right side of the node to be split, if the node resource utilization of the brother node on the right side meets the migration condition, write the information of the brother node on the right side and the information of the node to be split into the balanced multi-way search tree creation script; if both the node resource utilization of the brother node on the left side of the node to be split and the node resource utilization of the brother node on the right side of the node to be split do not meet the migration condition, enter the step of writing the information that does not meet the migration condition into the balanced multi-way search tree creation script; If the position to be written is located in the right half of the node to be split, query the right half of the node to be split. The node resource utilization of the brother node on the side of the node to be split is queried. If the node resource utilization of the brother node on the right side meets the migration condition, the information of the brother node on the right side and the information of the node to be split are written into the balanced multi-way search tree creation script; if the node resource utilization of the brother node on the right side does not meet the migration condition, the node resource utilization of the brother node on the left side of the node to be split is queried. If the node resource utilization of the brother node on the left side meets the migration condition, the information of the brother node on the left side and the information of the node to be split are written into the balanced multi-way search tree creation script; if the node resource utilization of the brother node on the left side of the node to be split and the node resource utilization of the brother node on the right side of the node to be split do not meet the migration condition, then enter the step of writing the information that does not meet the migration condition into the balanced multi-way search tree creation script.

12. The data storage method according to claim 1, characterized in that: The monitoring that the balanced multi-way search tree satisfies a node splitting condition includes: It is monitored that the balanced multi-way search tree has the node to be split whose node resource utilization reaches a third utilization threshold.

13. The data storage method according to claim 12, characterized in that: The monitoring that the balanced multi-way search tree has the node to be split whose node resource utilization reaches a third utilization threshold includes: It is monitored that the node to be split appears in the balanced multi-way search tree, the amount of data written into which reaches the order of the balanced multi-way search tree minus one.

14. The data storage method according to claim 12, characterized in that: The monitoring that the balanced multi-way search tree has the node to be split whose node resource utilization reaches a third utilization threshold includes: It is monitored that the node to be split has the number of written data reaching the order of the balanced multi-way search tree.

15. The data storage method according to claim 12, characterized in that: The querying of node resource utilization of the sibling nodes of the node to be split includes: Respectively query the node resource distribution of all the sibling nodes on the left side of the node to be split and the node resource distribution of all the sibling nodes on the right side of the node to be split; If the node resource utilization of the brother node meets the migration condition of receiving the data of the node to be split, the information of the brother node and the information of the node to be split are written into the balanced multi-way search tree creation script, so as to call the balanced multi-way search tree creation script to migrate the data of the node to be split to the brother node, and then update the parent node according to the data information of the brother node and the data information of the node to be split, including: If the node resource distribution of the brother node on one side meets the migration condition, a node resource balancing algorithm is used to write the information of the node to be split and the information of the brother node on the side that meets the migration condition into the balanced multi-way search tree creation script, so as to call the balanced multi-way search tree creation script to perform node resource balancing processing on the node to be split and the brother node on the side that meets the migration condition.

16. The data storage method according to claim 1, characterized in that: Also includes: Monitoring the node resource distribution of each leaf node of the balanced multi-way search tree; A node resource balancing algorithm is used to perform node resource balancing processing on each leaf node according to the node resource distribution of each leaf node.

17. The data storage method according to claim 1, characterized in that: The step of writing the information of the brother node and the information of the node to be split into the balanced multi-way search tree creation script, calling the balanced multi-way search tree creation script to migrate the data of the node to be split to the brother node, and then updating the parent node according to the data information of the brother node and the data information of the node to be split, comprises: Determine the data to be migrated of the node to be split so that after the migration, both the node to be split and the sibling node to which it is migrated do not meet the node splitting condition; Write the identifier of the node to be split, the identifier of the sibling node to be migrated, and the information of the data to be migrated into the balanced multi-way search tree creation script, so as to call the balanced multi-way search tree creation script to migrate the data to be migrated from the storage location corresponding to the node to be split to the storage location corresponding to the sibling node, and update the parent node according to the first data of the node to be split and the first data of the sibling node after migration.

18. The data storage method according to claim 1, characterized in that: The step of writing the information that does not meet the migration condition into the balanced multi-way search tree creation script so that the balanced multi-way search tree creation script performs the node splitting task on the node to be split, comprises: The information that does not meet the migration conditions is written into the balanced multi-way search tree creation script, so as to call the balanced multi-way search tree creation script to adopt the node splitting algorithm to calculate the data to be migrated of the node to be split and the position of the new node, migrate the data to be migrated to a temporary storage space and migrate the data to the new node after creating the new node, and then update the parent node according to the first data of the node to be split and the first data of the new node.

19. The data storage method according to claim 1, characterized in that: Applied to a distributed storage system, the data to be stored is metadata, and the balanced multi-way search tree is stored separately from other data in the distributed storage system.

20. The data storage method according to claim 1, characterized in that: The step of calculating the order of the balanced multi-way search tree to be created according to the data block size of the device and the size of the unit data to be stored includes: Running a node resource monitoring process, dividing the data block size of the device by the size of the unit data to be stored, to obtain the order of the balanced multi-way search tree to be created; The step of writing the order of the balanced multi-way search tree to be created into a balanced multi-way search tree creation script of the device, so as to call the balanced multi-way search tree creation script to create a balanced multi-way search tree for storing the data to be stored, comprises: Based on the node resource monitoring process, the order of the balanced multi-way search tree to be created is written into the balanced multi-way search tree creation process, the order of the balanced multi-way search tree to be created is written into the balanced multi-way search tree creation script, and the balanced multi-way search tree creation process is run to call the balanced multi-way search tree creation script to create the balanced multi-way search tree; When it is monitored that the balanced multi-way search tree meets the node splitting condition, querying the node resource utilization of the brother nodes of the node to be split includes: It is monitored that the node to which the data to be written is to be written calculated by calling the balanced multi-way search tree creation script satisfies the node splitting condition, and it is determined that the node to which the data to be written is to be written is the node to be split; Querying node resource utilization of the brother node whose distance from the node to be split is no more than one node; If the node resource utilization of the brother node meets the migration condition of receiving the data of the node to be split, the balanced multi-way search tree creation script is prohibited from executing the node splitting task and the information of the brother node and the information of the node to be split are written into the balanced multi-way search tree creation script, so as to call the balanced multi-way search tree creation script to migrate the data of the node to be split to the brother node, and then update the parent node according to the data information of the brother node and the data information of the node to be split, including: If the queried brother node meets the migration condition, the balanced multi-way search tree creation script is prohibited. The execution node splitting task is described; Determine the data to be migrated of the node to be split so that after the migration, both the node to be split and the sibling node to which it is migrated do not meet the node splitting condition; Writing the identifier of the node to be split, the identifier of the sibling node to be migrated, and the information of the data to be migrated into the balanced multi-way search tree creation script, so as to call the balanced multi-way search tree creation script to migrate the data to be migrated from the storage location corresponding to the node to be split to the storage location corresponding to the sibling node, and updating the parent node according to the first data of the node to be split and the first data of the sibling node after the migration; If the node resource utilization of all the brother nodes of the node to be split meets the migration condition, then writing the information that does not meet the migration condition into the balanced multi-way search tree creation script so that the balanced multi-way search tree creation script performs the node splitting task for the node to be split, including: If none of the queried brother nodes meet the migration condition, the information that does not meet the migration condition is written into the balanced multi-way search tree creation script, so as to call the balanced multi-way search tree creation script to use the node splitting algorithm to calculate the data to be migrated of the node to be split and the position of the new node, migrate the data to be migrated to a temporary storage space and migrate the data to the new node after creating the new node, and then update the parent node according to the first data of the node to be split and the first data of the new node.

21. A data storage device, characterized in that: include: A first calculation unit, used for calculating the order of the balanced multi-way search tree to be created according to the data block size of the device and the size of the unit data to be stored; A first writing unit, used for writing the order of the balanced multi-way search tree to be created into a balanced multi-way search tree creation script of the device, so as to call the balanced multi-way search tree creation script to create a balanced multi-way search tree for storing the data to be stored; A start control unit is used to start a node resource monitoring program for the balanced multi-way search tree; A first query unit is used to query the node resource utilization of the brother nodes of the node to be split when it is monitored that the balanced multi-way search tree meets the node splitting condition; A second writing unit is used for prohibiting the balanced multi-way search tree creation script from executing the node splitting task and writing the information of the brother node and the information of the node to be split into the balanced multi-way search tree creation script if the node resource utilization of the brother node meets the migration condition of receiving the data of the node to be split, so as to call the balanced multi-way search tree creation script to migrate the data of the node to be split to the brother node, and then update the parent node according to the data information of the brother node and the data information of the node to be split; If the node resource utilization of all the brother nodes of the node to be split meets the migration condition, the information that does not meet the migration condition is written into the balanced multi-way search tree creation script, so that the balanced multi-way search tree creation script executes the node splitting task for the node to be split.

22. A data storage device, characterized in that include: Memory for storing computer programs; A processor is used to execute the computer program, and when the computer program is executed by the processor, the steps of the data storage method according to any one of claims 1 to 20 are implemented.

23. A computer non-volatile readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the data storage method according to any one of claims 1 to 20 are implemented.

Citation Information

Patent Citations

  • B+ tree indexing method and device of real-time database

    CN102402602A

  • B+ tree operation device

    CN112989130A

  • Request allocation method and device, storage medium and electronic device

    CN116662019A

  • Data storage method, device and equipment and computer readable storage medium

    CN117312327A

  • Method of facilitating distributed data search in a federated cloud and system thereof

    US20190087445A1