A Method and Device for Optimizing the Storage Space of B+ Trees

By converting the non-leaf node key value of the b+ tree into an hashkey of equal length and storing only the leaf node address, the metadata space consumption problem caused by the long key value in the b+ tree is solved, and more efficient storage and query performance is achieved.

CN115630063BActive Publication Date: 2025-05-30CHINA ELECTRONICS CLOUD DIGITAL INTELLIGENCE TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211233861.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-10-10
Publication Date
2025-05-30
Estimated Expiration
2042-10-10

AI Technical Summary

Technical Problem

In high-performance storage media, the b+ tree structure consumes a lot of metadata space due to the long key value, especially in scenarios where the key value needs to be sorted in dictionary order, the space occupied by the key value increases.

Method used

By converting the non-leaf node key value of the b+ tree into an hashkey of equal length by using the hash algorithm, and only the address of the leaf node is stored in the non-leaf node, without storing the real key value.

Benefits of technology

This method effectively reduces the storage space occupied by non-leaf nodes, reduces the level of the tree, improves the efficiency of key values, and improves the search performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115630063B_ABST
    Figure CN115630063B_ABST
Patent Text Reader

Abstract

The present application discloses a method and device for optimizing the storage space of a B+ tree, including: converting the key values of the non-leaf nodes of the B+ tree into equal-length hash keys through a hash algorithm, where the hash key includes a first comparison unit and a second comparison unit, and the values recorded by the first comparison unit and the second comparison unit are different. During the search and comparison, the first comparison unit and the second comparison unit are compared in a preset order to achieve internal sorting; for the non-leaf nodes of the B+ tree, the real key values are not stored, and only the address of the leaf node representing the lower bound of the range corresponding to the entry (rec) of the non-leaf node is stored. The method of this embodiment converts the key values of the non-leaf nodes of the B+ tree into equal-length hash keys through a hash algorithm, and the hash key includes a first comparison unit and a second comparison unit. For the non-leaf nodes of the B+ tree, only the address of the leaf node representing the lower bound of the range of the rec is stored, thereby optimizing the data structure of the B+ tree and reducing the metadata space consumption caused by the key values.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of storage technologies, and in particular, to a method and device for optimizing the storage space of a B+ tree. Background Art

[0002] In traditional storage technologies, B+ trees are often used as the organization form of disk data, which can effectively reduce the number of disk reads and writes and have stable read and write performance. For example, MySQL and Oracle are both storage engines based on B trees. The non-leaf nodes and leaf nodes of a B+ tree both store keys instead of actual data, thereby increasing the number of indexes and reducing the level of the tree.

[0003] In high-performance storage media (such as SCM), due to the high cost per unit capacity, it is necessary to consider minimizing space occupancy as much as possible. For scenarios with long keys, such as a B+ tree structure that needs to be sorted in lexicographical order, since the key length is a variable-length string, the key value occupies a relatively large amount of space. Summary of the Invention

[0004] Embodiments of this application provide a method and device for optimizing the storage space of a B+ tree, which are used to optimize the data structure of the B+ tree and reduce the metadata space consumption caused by key values.

[0005] Embodiments of this application provide a method for optimizing the storage space of a B+ tree, including:

[0006] Converting the key values of the non-leaf nodes of the B+ tree into equal-length hash keys through a hash algorithm, where the hash key includes a first comparison unit and a second comparison unit, and the values recorded by the first comparison unit and the second comparison unit are different. When performing a search comparison, the first comparison unit and the second comparison unit are compared in a preset order to achieve internal sorting;

[0007] For the non-leaf nodes of the B+ tree, the real key values are not stored, and only the address of the leaf node representing the lower bound of the range corresponding to the entry (rec) of the non-leaf node is stored.

[0008] Optionally, the first comparison unit and the second comparison unit have the same length, and the first comparison unit further includes equal-length first and second comparison subunits, where:

[0009] The first comparison subunit is used to record the string length;

[0010] The second comparison subunit is the first hash value obtained by using a first hash algorithm;

[0011] The second comparison unit is the second hash value obtained by using a second hash algorithm;

[0012] During the search and comparison, the first comparison unit is compared first, and then the second comparison unit is compared to achieve internal sorting.

[0013] Optionally, for a string with a length less than the length of the hash key, the length of the string is stored in the first length byte, and the string is stored in the second length byte, where the sum of the first length and the second length is the same as the length of the hash key;

[0014] During the search and comparison, the first comparison unit is compared with the third length byte first, and then the fourth length byte is compared to achieve internal sorting, where the third length byte and the fourth length byte are the same as the length of the second comparison unit.

[0015] Optionally, it further includes storing the real key in the leaf node. When searching for the corresponding level of the B+ tree, the corresponding leaf node is found through the address in the non-leaf node, and the key value of the 0th rec of the leaf node is obtained for comparison.

[0016] An embodiment of the present application also provides a terminal device, including a processor, which is configured to:

[0017] Convert the key value of the non-leaf node of the B+ tree into an equal-length hash key through a hash algorithm, where the hash key includes a first comparison unit and a second comparison unit, and the values recorded by the first comparison unit and the second comparison unit are different. During the search and comparison, the first comparison unit and the second comparison unit are compared in a preset order to achieve internal sorting;

[0018] For the non-leaf node of the B+ tree, the real key value is not stored, and only the address of the leaf node representing the lower bound of the range corresponding to the non-leaf node entry (rec) is stored.

[0019] Optionally, the first comparison unit and the second comparison unit have the same length, and the first comparison unit further includes an equal-length first comparison subunit and a second comparison subunit, where:

[0020] The first comparison subunit is used to record the string length;

[0021] The second comparison subunit is the first hash value obtained by using the first hash algorithm;

[0022] The second comparison unit is the second hash value obtained by using the second hash algorithm;

[0023] During the search and comparison, the processor is further configured to: compare the first comparison unit first, and then compare the second comparison unit to achieve internal sorting.

[0024] Optionally, the processor is further configured to:

[0025] For a string with a length less than the length of the hash key, store the length of the string in a first length byte and store the string in a second length byte, where the sum of the first length and the second length is the same as the length of the hash key;

[0026] During lookup comparison, first compare a third length byte with the first comparison unit, and then compare a fourth length byte to achieve internal sorting, where the third length byte and the fourth length byte are the same as the length of the second comparison unit.

[0027] Optionally, the processor is further configured to:

[0028] It further includes storing the real key in a leaf node. In the case of searching for a corresponding level of the B+ tree, find the corresponding leaf node through the address in the non-leaf node, and obtain the key value of the 0th rec of the leaf node for comparison.

[0029] An embodiment of the present application also proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing method for optimizing the storage space of the B+ tree are implemented.

[0030] The method of this embodiment converts the key value of the non-leaf node of the B+ tree into an equal-length hash key through a hash algorithm, and the hash key includes a first comparison unit and a second comparison unit. For the non-leaf node of the B+ tree, only store the address of the leaf node representing the lower bound of the range represented by the rec, thereby optimizing the data structure of the B+ tree and reducing the metadata space consumption caused by the key value.

[0031] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the specific embodiments of the present application are specifically given below. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] By reading the detailed description of the preferred embodiments below, various other advantages and benefits will become clear to those of ordinary skill in the art. The drawings are only for the purpose of showing the preferred embodiments and are not considered to be a limitation of the present application. And throughout the drawings, the same reference numerals are used to represent the same components. In the drawings:

[0033] Figure 1Hash key example for the b+ tree storage space optimization method of the embodiments of the present application;

[0034] Figure 2 Example of b+ tree storage space optimization for the embodiments of the present application;

[0035] Figure 3 Example of storing strings less than 15 bytes by the b+ tree storage space optimization method of the embodiments of the present application. Detailed implementation manners

[0036] The exemplary embodiments of the present disclosure will be described in more detail below with reference to the accompanying drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.

[0037] The embodiments of the present application provide a method for optimizing the storage space of a b+ tree, including:

[0038] Converting the key values of the non-leaf nodes of the b+ tree into equal-length hash keys through a hash algorithm. For string scenarios where sorting according to the input key is not required, the fixed length of the key values can be achieved through the hash algorithm, thereby increasing the order of the middle layer of the tree, reducing the number of levels, and reducing the storage space of non-leaf nodes. At the same time, the comparison efficiency of the key values can also be improved, and the search performance can be enhanced.

[0039] The hash key includes a first comparison unit and a second comparison unit, and the values recorded by the first comparison unit and the second comparison unit are different. When performing a search comparison, the first comparison unit and the second comparison unit are compared in a preset order to achieve internal sorting. In some specific examples, the first comparison unit and the second comparison unit have the same length, and the first comparison unit further includes equal-length first and second comparison subunits, where:

[0040] The first comparison subunit is used to record the string length;

[0041] The second comparison subunit is the first hash value obtained by using the first hash algorithm;

[0042] The second comparison unit is the second hash value obtained by using the second hash algorithm;

[0043] When performing a search comparison, the first comparison unit is compared first, and then the second comparison unit is compared to achieve internal sorting.

[0044] Such as Figure 1As shown, in this example, a hash key for a 16-byte double hash comparison mechanism can be designed. Among them, the first 8 bytes serve as the first comparison unit. The first 4 bytes of the first comparison unit serve as the first comparison subunit, which is used to record the string length. The last 4 bytes of the first comparison subunit serve as the second comparison subunit, and a 4-byte hash value is obtained using the dbj2 hash algorithm. The subsequent 8 bytes serve as the second comparison unit, and an 8-byte hash value is obtained through the murmur hash algorithm. During the search comparison, by first comparing the first 8 bytes and then the last 8 bytes, internal sorting is achieved.

[0045] For non-leaf nodes of the B+ tree, the real key values are not stored, and only the addresses of the leaf nodes corresponding to the lower bounds of the ranges represented by the corresponding entries (rec) of these non-leaf nodes are stored.

[0046] In some examples, it also includes storing the real key in the leaf nodes. When searching at the corresponding level of the B+ tree, the corresponding leaf node is found through the address in the non-leaf node, and the key value of the 0th rec of this leaf node is obtained for comparison.

[0047] In this way, the length of the keys in non-leaf nodes can be effectively reduced. Only the address of the corresponding leaf node, 8 bytes, needs to be recorded. As Figure 2 shown, calculated according to a 64-byte user key, for a 4096-page, the number of stored child nodes increases from 4096 / (64 + 8) = 56 to 4096 / (8 + 8) = 256. For a 3-layer tree, taking the number of kvs that 4096 can store as the order, the total number of indexable kvs increases from 56 3 = 175,616 to 256 3 = 16,777,216. The required nodes are 56 2 + 56 + 1 = 3193, 256 2 + 256 + 1 = 65,793. The proportion of non-leaf nodes decreases from 57 / 3193 = 0.017 (1.7%) to 257 / 65793 = 0.003 (0.3%), greatly reducing the levels of the same number of kvs and reducing the space occupied by non-leaf nodes.

[0048] In some embodiments, for a string with a length less than the length of the hash key, the length of the string is stored using the first length byte, and the string is stored using the second length byte, where the sum of the first length and the second length is the same as the length of the hash key;

[0049] When performing a search and comparison, first compare the first comparison unit of the third length byte, and then compare the fourth length byte to achieve internal sorting, where the third length byte, the fourth length byte have the same length as that of the second comparison unit.

[0050] As Figure 3 shown, in order to reduce the computational consumption of hashing, in this example, the design logic of the hash key is further optimized. For strings less than 15 bytes, no hash calculation is performed in this embodiment, and the key value is directly stored. For example, the first byte can be used to store the length, and the next 15 bytes store the string. When performing a search and comparison, the comparison principle of the first 8 bytes and the last 8 bytes is still followed.

[0051] In this embodiment, the variable-length key is converted into a 16-byte hash value for storage, which is applicable to kv storage with relatively long keys and does not require sorting in lexicographical order. Assuming a user key of 64 bytes is calculated, for a 4096-page, the number of stored child nodes increases from 4096 / (64 + 8) = 56 to 4096 / (16 + 8) = 170. For a three-layer tree, taking the number of kv that 4096 can store as the order, the total number of indexable kv increases from 56 3 = 175,616 to 170 3 = 4,913,000. The required nodes are 56 2 + 56 + 1 = 3193, 170 2 + 170 + 1 = 29071. The proportion of non-leaf nodes is reduced from 57 / 3193 = 0.017 (1.7%) to 171 / 29071 = 0.0058 (0.58%), effectively reducing the space of non-leaf nodes and the tree levels.

[0052] The method of this embodiment converts the key values of the non-leaf nodes of the B+ tree into equal-length hashkeys through a hash algorithm, and the hashkey includes a first comparison unit and a second comparison unit. For the non-leaf nodes of the B+ tree, only the address of the leaf node representing the lower bound of the range represented by the rec is stored, thereby optimizing the data structure of the b+ tree and reducing the metadata space consumption caused by the key values.

[0053] This application embodiment also proposes a terminal device, including a processor, which is configured to:

[0054] Convert the key values of the non-leaf nodes of the b+ tree into equal-length hashkeys through a hash algorithm, where the hashkey includes a first comparison unit and a second comparison unit, and the values recorded by the first comparison unit and the second comparison unit are different. When performing a search and comparison, compare the first comparison unit and the second comparison unit in a preset order to achieve internal sorting;

[0055] For non-leaf nodes of the B+ tree, the real key values are not stored, and only the addresses of the leaf nodes corresponding to the lower bounds of the ranges represented by the corresponding entries (rec) of the non-leaf nodes are stored.

[0056] In some embodiments, the first comparison unit and the second comparison unit have the same length, and the first comparison unit further includes a first comparison subunit and a second comparison subunit of equal length, where:

[0057] The first comparison subunit is used to record the string length;

[0058] The second comparison subunit is the first hash value obtained by using the first hash algorithm;

[0059] The second comparison unit is the second hash value obtained by using the second hash algorithm;

[0060] When performing a search comparison, the processor is further configured to: first compare the first comparison unit, and then compare the second comparison unit to achieve internal sorting.

[0061] In some embodiments, the processor is further configured to:

[0062] For a string with a length less than the length of the hash key, the length of the string is stored using a first length byte, and the string is stored using a second length byte, where the sum of the first length and the second length is the same as the length of the hash key;

[0063] When performing a search comparison, first compare the first comparison unit with a third length byte, and then compare with a fourth length byte to achieve internal sorting, where the third length byte and the fourth length byte are the same as the length of the second comparison unit.

[0064] In some embodiments, the processor is further configured to:

[0065] It further includes storing the real key in the leaf node. When searching for the corresponding level of the B+ tree, the corresponding leaf node is found through the address in the non-leaf node, and the key value of the 0th rec of the leaf node is obtained for comparison.

[0066] An embodiment of the present application also proposes a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the steps of the foregoing B+ tree storage space optimization method are implemented.

[0067] It should be noted that in this text, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements not only includes those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising such element.

[0068] The serial numbers of the embodiments of the present application above are for description only and do not represent the superiority or inferiority of the embodiments.

[0069] Through the description of the above embodiments, those skilled in the art can clearly understand that the above embodiment methods can be implemented by means of software plus a necessary general hardware platform. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation. Based on such an understanding, the technical solution of the present application, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions for causing a terminal (which may be a mobile phone, computer, server or network device, etc.) to execute the methods described in the various embodiments of the present application.

[0070] The embodiments of the present application have been described above in conjunction with the accompanying drawings. However, the present application is not limited to the above specific embodiments. The above specific embodiments are merely illustrative and not restrictive. Under the inspiration of the present application, those of ordinary skill in the art can also make many forms without departing from the purpose of the present application and the scope protected by the claims. All of these are within the protection scope of the present application.

Claims

1. A method for optimizing the storage space of a B+ tree, characterized in that, it includes: Converting the key values of the non-leaf nodes of the B+ tree into equal-length hash keys through a hash algorithm, where the hash key includes a first comparison unit and a second comparison unit, and the values recorded by the first comparison unit and the second comparison unit are different. During the search comparison, the first comparison unit and the second comparison unit are compared in a preset order to achieve internal sorting; For the non-leaf nodes of the B+ tree, the real key values are not stored, and only the addresses of the leaf nodes corresponding to the lower bounds of the ranges represented by the corresponding entries rec of the non-leaf nodes are stored; The first comparison unit and the second comparison unit have the same length, and the first comparison unit further includes equal-length first and second comparison subunits, where: The first comparison subunit is used to record the string length; The second comparison subunit is the first hash value obtained by using the first hash algorithm; The second comparison unit is the second hash value obtained by using the second hash algorithm; During the search comparison, the first comparison unit is compared first, and then the second comparison unit is compared to achieve internal sorting.

2. The method for optimizing the storage space of a B+ tree according to claim 1, characterized in that, For a string with a length less than the length of the hash key, the length of the string is stored using the first length byte, and the string is stored using the second length byte, where the sum of the first length and the second length is the same as the length of the hash key; During the search comparison, the first comparison unit of the third length byte is compared first, and then the fourth length byte is compared to achieve internal sorting, where the third length byte and the fourth length byte are the same as the length of the second comparison unit.

3. The method for optimizing the storage space of a B+ tree according to claim 1, characterized in that, It further includes storing the real key in the leaf nodes. When searching for the corresponding level of the B+ tree, the corresponding leaf node is found through the address in the non-leaf node, and the key value of the 0th entry of the leaf node is obtained for comparison.

4. A terminal device, characterized in that, it includes a processor configured to: Converting the key values of the non-leaf nodes of the B+ tree into equal-length hash keys through a hash algorithm, where the hash key includes a first comparison unit and a second comparison unit, and the values recorded by the first comparison unit and the second comparison unit are different. During the search comparison, the first comparison unit and the second comparison unit are compared in a preset order to achieve internal sorting; For the non-leaf nodes of the B+ tree, the real key values are not stored, and only the addresses of the leaf nodes corresponding to the lower bounds of the ranges represented by the corresponding entries rec of the non-leaf nodes are stored; The first comparison unit and the second comparison unit have the same length, and the first comparison unit further includes equal-length first and second comparison subunits, where: The first comparison subunit is used to record the string length; The second comparison subunit is the first hash value obtained by using the first hash algorithm; The second comparison unit is the second hash value obtained by using the second hash algorithm; When performing a search and comparison, the processor is further configured to: first compare the first comparison unit, and then compare the second comparison unit to achieve internal sorting.

5. The terminal device according to claim 4, wherein, the processor is further configured to: For a string with a length less than the length of the hash key, use the first length byte to store the length of the string, and use the second length byte to store the string, where the sum of the first length and the second length is the same as the length of the hash key; When performing a search and comparison, first compare the first comparison unit with the third length byte, and then compare with the fourth length byte to achieve internal sorting, where the third length byte and the fourth length byte are the same as the length of the second comparison unit.

6. The terminal device according to claim 4, wherein, the processor is further configured to: It further includes storing the real key in the leaf node. When searching the corresponding level of the B+ tree, find the corresponding leaf node through the address in the non-leaf node, and obtain the key value of the 0th entry of the leaf node for comparison.

7. A computer-readable storage medium, wherein, a computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the steps of the B+ tree storage space optimization method according to any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Data retrieval optimization method and device and computer equipment

    CN111382323A

  • Optimization method for effectively improving B+ tree retrieval efficiency on GPU

    CN111966678A