A Method and System for Batch Sequential Deletion of B+ Trees

By splitting the B+ tree into multiple subtrees and using linked list management pointers, batch order deletion of the B+ tree is achieved, which solves the problem of inefficiency in the existing technology and improves the operation efficiency of frequent insertion and deletion.

CN114969033BActive Publication Date: 2025-07-18ZHUHAI GOTECH INTELLIGENT TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210501500.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-10
Publication Date
2025-07-18
Estimated Expiration
2042-05-10

AI Technical Summary

Technical Problem

The batch deletion operation of existing B+ trees is inefficient, especially in scenarios where frequent insertion and deletion are not effectively improved.

Method used

Divide the B+ tree into multiple subtrees, and manage the tree root node pointer and maximum value through a linked list. Use the p_tree_head pointer to mark the starting position of sequential deletion, and use the pointer to move backward to achieve batch deletion.

Benefits of technology

The efficiency of deletion of B+ trees is significantly improved, especially in scenarios where frequent insertion and deletion are required, and the efficiency of sorting and splicing data is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114969033B_ABST
    Figure CN114969033B_ABST
Patent Text Reader

Abstract

The present invention provides a method for batch sequential deletion of B+ trees, including the following steps: Divide a B+ tree into multiple ones; Define a linked list, and store the root node pointer and the maximum value of the tree into the linked list; Define a p_tree_head pointer for each B+ tree, and the p_tree_head pointer is used to mark the position of the first valid node after sequential deletion; When inserting and querying, first traverse the linked list to find the B+ tree to be inserted, and sequentially compare the maximum values of the trees in the linked list. If it is smaller than this value, query or insert into this B+ tree; For the deletion operation, find the p_tree_head pointer of the first node from the linked list, sequentially fetch the B+ trees backward, and then move the pointer backward; Until the key value corresponding to the p_tree_head pointer is the maximum value of this B+ tree; Delete the entire B+ tree and simultaneously delete the corresponding node in the linked list. On the premise of batch sequential deletion, the present invention can greatly improve the efficiency of the deletion operation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of software, and specifically provides a method and system for batch sequential deletion of B+ trees. Background Art

[0002] A B+ tree is a tree data structure, usually used in database and file index systems of operating systems. The B+ tree is characterized by being able to keep data stable and ordered, and its insertion and deletion have relatively stable logarithmic time complexity. The B+ tree is a balanced search tree, and all record nodes are stored in the leaf nodes of the same layer in the order of key values, and the leaf node pointers are connected.

[0003] In the prior art, the deletion operation of the B+ tree generally targets a single key value, and there is no effective method for batch deletion of nodes or multiple nodes. Some patents provide batch deletion operations, which are aimed at the low disk deletion efficiency. The B+ tree is read into the memory for operation and then the disk is batch modified. Its disadvantages are: first batch delete in the memory, and then batch write to the disk. Because the deleted nodes are not continuous, when operating the disk, it is still necessary to operate one by one, and the time complexity is still too high.

[0004] Scenario 1: Live broadcast scenario. The capture bit rate of the host client is extremely high, and multi-threaded upload to the server is enabled. The program segments uploaded by each thread are different, and each segment can be sorted according to time. Key: time slice; Value: pointer to program data. Due to multi-threaded upload to the server, the segments received by the server are messy. At this time, if the B+ tree is used for sorting and the key values are taken out in the order of time slices, a sequential program stream can be finally obtained.

[0005] Scenario 2: P2P download scenario. The client downloads data from multiple P2P nodes, and the obtained data is uncertain and repeated, so it needs to be sorted and output. At this time, if the B+ tree is used for sorting and then the key values are taken in order, a sequential program stream can be obtained.

[0006] Of course, there are many other specific application scenarios that require sequential deletion of key values. If deleted one node by one in the current way, the efficiency is too low. Summary of the Invention

[0007] In view of the deficiencies of the background art, the main purpose of the present invention is to provide a method for solving batch sequential deletion of key values in a B+ tree. The deletion is performed in the order of the sizes of the leaf nodes of the B+ tree, so as to improve the deletion efficiency, especially for scenarios with frequent insertion and deletion. And the method for deleting key values provided by the present invention is applicable to the case of deleting in the order of the sizes of the leaf nodes of the B+ tree.

[0008] The object of the present invention is achieved by the following technical solutions:

[0009] A method for batch sequential deletion of B+ trees, comprising the following steps:

[0010] S1. Divide a B+ tree into multiple ones;

[0011] S2. Define a linked list, and store the root node pointer and the maximum value of the tree into the linked list;

[0012] S3. Define a p_tree_head pointer for each B+ tree, and the p_tree_head pointer is used to mark the position of the first valid node after sequential deletion;

[0013] S4. When inserting and querying, first traverse the linked list to find the B+ tree to be inserted, and compare the maximum values of the trees in the linked list in sequence. If it is smaller than the maximum value of the trees in the linked list, query or insert into this B+ tree;

[0014] S5. During the deletion operation, find the p_tree_head pointer of the first node from the linked list, sequentially fetch the B+ trees backward, and then move the pointer backward until the key value corresponding to the p_tree_head pointer is the maximum value of this B+ tree;

[0015] S6. Delete the entire B+ tree, and at the same time delete the corresponding node in the linked list.

[0016] Preferably, in S1, the method of dividing a B+ tree into multiple ones is to divide according to the key size. In fact, it can be divided according to the actual data and needs, and generally it is divided according to a certain rule.

[0017] A system for batch sequential deletion of B+ trees, which includes:

[0018] A B+ tree splitting unit, which is used to divide a B+ tree into multiple ones;

[0019] A linked list generation unit, which is used to define a linked list and store the root node pointer and the maximum value of the tree into the linked list;

[0020] A pointer definition unit, which is used to define a p_tree_head pointer for each B+ tree, and the p_tree_head pointer is used to mark the position of the first valid node after sequential deletion;

[0021] An insertion and query unit, which is used to first traverse the linked list to find the B+ tree to be inserted, and compare the maximum values of the trees in the linked list in sequence. If it is smaller than the maximum value of the trees in the linked list, query or insert into this B+ tree;

[0022] A deletion unit, which is used to find the p_tree_head pointer of the first node in the linked list, sequentially retrieve the B+ tree backward, and then move the pointer backward; until the key value corresponding to the p_tree_head pointer is the maximum value of the B+ tree;

[0023] The deletion unit deletes the entire B+ tree and simultaneously deletes the corresponding node in the linked list.

[0024] Preferably, the B+ tree splitting unit is specifically used to divide one B+ tree into multiple according to the key size.

[0025] Compared with the prior art, the present invention has the following advantages:

[0026] The method for batch sequential deletion of B+ trees in the present invention; through the method of dividing trees, the deletion is completed by moving the pointer backward, and finally the entire tree is deleted to achieve batch deletion. On the premise of batch sequential deletion, the present invention can greatly improve the efficiency of deletion operations. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] Figure 1 It is a schematic diagram of the principle of an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] This embodiment provides a method for batch sequential deletion of B+ trees, including the following steps:

[0029] Divide one B+ tree into multiple: The specific splitting method can be to divide according to the key size, or according to the actual data and needs, and just divide according to a certain rule;

[0030] Define a linked list, and store the root node pointer and the maximum value of the tree in the linked list;

[0031] Define a p_tree_head pointer for each B+ tree, and the p_tree_head pointer is used to mark the position of the first valid node after sequential deletion;

[0032] When inserting and querying, first traverse the linked list to find the B+ tree to be inserted, and sequentially compare the maximum values of the trees in the linked list. If it is smaller than the maximum value of the tree in the linked list, query or insert into the B+ tree;

[0033] For the deletion operation, find the p_tree_head pointer of the first node in the linked list, sequentially retrieve the B+ tree backward, and then move the pointer backward; until the key value corresponding to the p_tree_head pointer is the maximum value of the B+ tree;

[0034] Delete the entire B+ tree and simultaneously delete the corresponding node in the linked list.

[0035] As Figure 1As shown below, this is an example of this embodiment:

[0036] Figure 1 In the figure, the linked list shows that a B+ tree is divided into 4, and the first and second B+ trees are displayed.

[0037] Query operation: If the key to be queried is 199, then through the p_root_node_head pointer, the pointers and maximum values of each B+ tree are obtained in turn. Comparing the maximum value 100 of the first root_node1, since 199 is greater than 100, continue to search downwards. Comparing root_node2, since 199 is less than 200, then query in the B+ tree root_node2.

[0038] Insertion operation: If the key to be inserted is 299, first query according to the above steps, and obtain that the B+ tree to be inserted is root_node3, then insert 299 into root_node3.

[0039] Deletion operation: p_tree_head traverses in the order of the leaf nodes of the B+ tree. Suppose p_tree_head points to 37, which means that all values less than 37 in root_node1 need to be deleted. According to the property that the leaf nodes of the B+ tree are doubly linked lists and each node is an array internally, traverse backwards. When the value of the next leaf node after traversing to 37 is 81, then it is necessary to wait until all the intermediate values are inserted and then continue to traverse backwards. This traversal process is actually a process of sequential deletion, but the actual deletion operation is not really executed. All the traversed leaf nodes are the nodes to be deleted.

[0040] Batch deletion operation: When p_tree_head points to 99, continue to traverse the leaf nodes of root_node1 to 100. Since the maximum value of root_node1 is 100, so p_tree_head points to the next B+ tree root_node2, which means that the entire B+ tree of root_node1 can be deleted, and then the entire tree of root_node1 is deleted. In this way, the efficiency of deleting the entire B+ tree is extremely high.

[0041] Figure 1 All the values in the figure are keys, and the values are not shown because the design of the value can be based on your own needs. For example, for video data, the value can be a pointer to video segment data.

[0042] Figure 1 The figure does not show how the key loops, that is, the key has a maximum value and cannot increase infinitely. According to the actual business, a maximum value of the key can be agreed upon, and when it exceeds this maximum value, another linked list can be started.

[0043] The principles and advantages of the embodiments of the present invention will be described below in combination with the two scenarios described in the background art:

[0044] Scenario 1: Live broadcast scenario. During a live broadcast, when the bitrate is extremely large, the efficiency of a single thread sending data to the server is low. If multiple threads are started to send data, then the video segments sent by each thread will be different. The client can agree on the splicing rules with the server. However, if a unique sequential key is set for each segment using a B+ tree, the video segments can be spliced through the B+ tree because the video data is continuous and sequential. After the splicing is completed, we read and delete the B+ tree sequentially. However, the deletion operation of the B+ tree is inefficient. If the above batch deletion operation method is adopted, the sorting efficiency of the B+ tree can be greatly improved.

[0045] Scenario 2: P2P download scenario. Regarding the problem of P2P multi-node download, here we discuss the situation where multiple nodes provide the same video source and data is pulled from multiple nodes simultaneously to splice into a video, rather than choosing to pull from a single P2P node. Then how can a B+ tree be used to splice the video? By slicing each video segment according to the same rule and marking it with a number, and inserting it into the B+ tree. After sorting through the B+ tree, the B+ tree is deleted sequentially after splicing to obtain a sequential video. The efficiency of single deletion is low. If the above batch deletion operation method is adopted, the efficiency can be greatly improved.

[0046] The above are only the preferred embodiments of the invention, and do not impose any form of limitation on the invention. Any simple modifications, equivalent changes, and decorations made to the above embodiments based on the technical essence of the invention still fall within the scope of the technical solution of the invention.

Claims

1. A method for batch sequential deletion of B+ trees, characterized in that, It includes the following steps: S1. Divide one B+ tree into multiple ones; S2. Define a linked list and store the root node pointer and the maximum value of the tree into the linked list; S3. Define a p_tree_head pointer for each B+ tree, and the p_tree_head pointer is used to mark the position of the first valid node after sequential deletion; S4. When inserting and querying, first traverse the linked list to find the B+ tree to be inserted, compare the maximum values of the trees in the linked list in sequence, and if it is smaller than the maximum value of the tree in the linked list, query or insert into this B+ tree; S5. During the batch deletion operation, find the p_tree_head pointer of the first node in the linked list, sequentially fetch the B+ trees backward, then move the pointer backward until the key value corresponding to the p_tree_head pointer is the maximum value of this B+ tree, delete the entire B+ tree, and at the same time delete the corresponding node in the linked list.

2. The method according to claim 1, wherein In the above S1, the method of dividing one B+ tree into multiple ones is to divide according to the key size.

3. A system for batch sequential deletion of B+ trees, characterized in that, It includes: A B+ tree splitting unit, which is used to divide one B+ tree into multiple ones; A linked list generating unit, which is used to define a linked list and store the root node pointer and the maximum value of the tree into the linked list; A pointer defining unit, which is used to define a p_tree_head pointer for each B+ tree, and the p_tree_head pointer is used to mark the position of the first valid node after sequential deletion; An insertion and querying unit, which is used to first traverse the linked list to find the B+ tree to be inserted, compare the maximum values of the trees in the linked list in sequence, and if it is smaller than the maximum value of the tree in the linked list, query or insert into this B+ tree; A batch deletion unit, which is used to find the p_tree_head pointer of the first node in the linked list, sequentially fetch the B+ trees backward, then move the pointer backward until the key value corresponding to the p_tree_head pointer is the maximum value of this B+ tree, delete the entire B+ tree, and at the same time delete the corresponding node in the linked list.

4. The system according to claim 3, characterized in that, The B+ tree splitting unit is specifically used to divide one B+ tree into multiple ones according to the key size.

Citation Information

Patent Citations

  • Method and system for processing data in treelike structures

    CN102867059A

  • Array tree data storage method, quick search method and readable storage medium

    CN111581215A