Tree-Based Data Structures
Patent Information
- Application Number
- JP2023564186
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2021-05-26
- Filing Date
- 2022-05-18
- Publication Date
- 2025-05-23
- Estimated Expiration
- 2042-05-18
AI Technical Summary
Existing tree-based data structures like B-trees face challenges in balancing the workload between writers and readers, where writing in sorted order benefits readers but burdens writers, while writing chronologically benefits writers but degrades reader performance.
Implement a tree structure with large and small blocks, where large blocks are sorted for efficient searches and small blocks contain recent updates, using optimizations like shortcut keys, back pointers, and order hints to maintain performance balance.
This approach enhances search performance while maintaining fast update capabilities, leveraging hardware acceleration for read-intensive operations, reducing latency and improving throughput.
Smart Images

Figure 00000000_0000_ABST
Abstract
Description
[Background technology]
[0001] background A tree structure, such as a B-tree, can be used as an index data structure to store key-value pairs (each pair comprises a value, which is what is stored, and a key, which is used to index that value). Each leaf node of the tree contains multiple items, which are key-value pairs. Going further up the tree from the leaf, each internal node of the tree contains an indication of the range of key values covered by each of its children (which may be leaves or other internal nodes). When a new item is added to the tree, the writer uses the tree structure to determine which leaf contains the range of keys that the new item's key falls into. If this leaf is not full, the new item can simply be added to the existing leaf. However, if the leaf is full (leaves and internal nodes usually have a maximum size in bytes), the leaf is split; that is, a new leaf is created. This also means updating the parent nodes of these two leaves to correctly reference the key ranges of the old and new leaves. If this update to the parent causes the parent to exceed its maximum size, the parent itself is split, the grandparent references are updated, and so on. If an item is deleted, this may also involve merging leaf or internal nodes.
[0002] There are several operations that a writer or reader may be required to perform in an application. A reader can read an individual item with a particular key, or can do a range scan to read from a range of keys. A writer can write a new entry, modify an existing entry, or delete an entry.
[0003] Index data structures are used in a wide variety of software systems, including data stores and databases, file systems, and operating systems. Index data structures store key-value pairs and support, among other operations, lookup, scanning, insertion, and deletion of key-value pairs. A B-tree is one such index data structure. A B-tree is an ordered index, which means that it also supports a scan operation. A scan operation returns all key-value pairs with keys within a specified range that are stored in the tree. Summary of the Invention
[0004] overview In tree-based data structures such as B-trees, it is desirable to balance the workload between writers and readers. For example, if new entries are simply added chronologically, and not sorted by key, this is very fast for writers. But readers must sort items on reads in order to do individual reads or range scans (at least range scans require sorting - individual reads are simply easier to implement both lookups and range scans the same way, using a sort). On the other hand, if a writer puts all new entries into sorted order every time it does a write, this makes reads very fast for readers, but puts a bigger burden on writers on writes. It is desirable to provide a compromise between these two approaches.
[0005] According to an aspect disclosed herein, a system is provided that includes a memory for storing a data structure, the data structure including a tree structure including a plurality of nodes each having a node ID. Some nodes are leaf nodes and other nodes are internal nodes, and each internal node is a parent to a respective set of one or more children in the tree structure. Each child is either a leaf node or another internal node, and each leaf node is a child but not a parent. Each of the leaf nodes includes a respective set of one or more items each having a key-value pair, and each of the internal nodes maps a node ID of each of the respective children to a range of keys encompassed by the respective child. The system further includes a writer implemented in software, hardware, or a combination thereof and arranged to write items to the leaf nodes, and a reader implemented in software, hardware, or a combination thereof and arranged to read items from the leaf nodes. Each of the leaf nodes includes a respective first block and a respective second block, the first block including a plurality of items of the respective leaf nodes sorted in key order in an address space of the memory. The writer is configured, when writing new items to the tree structure, to identify which leaf node to write to by following the mapping of keys to node IDs through the tree structure, and to write the new items to the second block of the identified leaf node in the order in which they were written, rather than sorted in key order. The reader, when reading one or more target items from the tree structure, is configured, when reading from the tree structure, to determine which leaf node to read from by following the mapping of keys to node IDs through the tree structure, and then to search for the determined leaf node for the one or more target items based on a) the order of the items already sorted in the first block, and b) the reader sorting the items in the second block by keys associated with the items in the first block.
[0006] In an embodiment, the first block may have a larger maximum size than the second block, and thus the first block may be referred to as a "large block" and the second block may be referred to as a "small block." In an embodiment, the writer may be implemented in software, while the reader may be implemented in custom hardware, such as a programmable gate array (PGA) or field programmable gate array (FPGA). However, more generally, the writer may be implemented in hardware (e.g., a hardware accelerator) and / or the reader may be implemented in software.
[0007] BRIEF DESCRIPTION OF THE DRAWINGS To aid in understanding the embodiments disclosed herein and to illustrate how such embodiments may be put into practice, reference is made, by way of example only, to the following drawings: [Brief description of the drawings]
[0008] [Figure 1] 1 illustrates a schematic of a range scan in the form of an inclusive scan (where the interval formed by the keys of the returned items is the largest subset that falls within the input interval) for an exemplary key interval [14,22]. [Diagram 2] 1 illustrates a schematic of a range scan in the form of a covering scan (where the interval formed by the keys of the returned items is the smallest superset that contains the input interval) for an exemplary key interval [14,22]. [Diagram 3] 1 illustrates a schematic diagram of a range scan in the form of a covering scan for an exemplary key section (14,22). [Figure 4] FIG. 1 is a schematic block diagram of a system according to an embodiment disclosed herein. [Diagram 5] 1 shows a schematic diagram of a data structure with a tree structure. [Figure 6] 1 illustrates a schematic diagram of a node with small and large blocks. [Figure 7]1 illustrates a schematic diagram of a method for inserting a new item into a small block. [Figure 8] 13 illustrates a schematic diagram of a method for including an order hint for sorting small blocks. [Figure 9] 1 illustrates a schematic diagram of a method for sorting small blocks based on order hints. [Figure 10] FIG. 1 is a schematic circuit diagram of a four-element small block sorter. [Figure 11] 1 illustrates a schematic diagram of a method for merging large and small blocks during an insert operation. [Figure 12] 1 illustrates a schematic diagram of a method for splitting leaf nodes upon insertion. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] Detailed Description of the Embodiments The present disclosure provides a node layout in which each node consists of a first and a second block. Preferably, the first block has a larger maximum size (in bits or bytes) than the second block, and thus the first and second blocks may be referred to as large and small blocks, respectively. The following is described with reference to large and small blocks as examples, but everywhere below this could in principle be replaced more generally with "first block" and "second block", respectively.
[0010] The large block is sorted, which allows for efficient node lookup. Sorted in this context means that the items are sorted by key such that the order of items by key is the same as the order in which the items appear in the memory's address space (virtual or physical address space, depending on the implementation). The small block, on the other hand, comprises the most recent insertions and deletions to the nodes and is time-ordered, which allows for fast updates. Alternatively, the entire node can be kept unsorted, which provides fast update performance but slower lookup performance, or the entire node can be kept sorted, which provides fast lookup performance but slower update performance.
[0011] Optionally, in large blocks, shortcut keys may be used, which allows the node search to look at only a small portion of the total nodes, resulting in improved performance.
[0012] As another, alternative or additional optimization, in embodiments, "back-pointers" may be included in the small block. These allow for establishing an order between items in the small and large blocks without comparing keys. This results in an efficient implementation (e.g., in hardware) since the entire key does not need to be compared.
[0013] As another optional optimization that can be used in conjunction with or independent of shortcut keys and / or back pointers, in an embodiment, "order hints" may be included in the small block. Order hints within the small block allow highly efficient hardware to establish a sorted order of items within the small block without comparing keys.
[0014] Yet another optional optimization is to provide synchronization of complex updates, where when a node is split or merged, the writer updates the tree by making a copy of the subtree (e.g., a thread-private copy) and then swapping it in with a single pointer update in the page table.
[0015] In an embodiment, the writer may be implemented in software running on one or more processors, while the reader may be implemented in specialized hardware, which may comprise dedicated (i.e. fixed function) circuitry such as an FPGA or PGA, or an ASIC (application specific integrated circuit).
[0016] Specialized hardware that implements index data structures is hardware that has a dedicated pipeline for performing data structure operations. This is in contrast to a general purpose CPU that implements an instruction set and executes a program that uses that instruction set to implement data structure operations. Specialized hardware has several advantages. It can have a higher throughput per watt, be more predictable and have lower latency, and be less expensive than the same functionality executed as software on a CPU. The main disadvantages of using specialized hardware are that it is more difficult to design, takes longer to build than software, and specialized hardware cannot be modified after it is built, so it is typically only used for very widely deployed functionality. However, it should be noted that the scope of this disclosure is not limited to implementing writes in software and reads in hardware. As another possibility, the writer may be implemented in some form of hardware such as an FPGA or ASIC, and / or the reader may be implemented in software. The embodiments disclosed herein provide an in-memory B-tree that benefits from the best of both worlds; that is, all operations can be performed in software on the CPU, and lookup and scan operations can be performed on the specialized hardware that we designed. In many workloads, lookups and scans are the dominant operations, meaning that most operations benefit from hardware offloading. Lookups and scans are simpler than update operations, meaning that the entire system can be built and deployed faster. In one exemplary implementation, the memory may be a DRAM-based host memory, but this is not limiting, and in another example, the B-tree may be stored in a storage NVM device, such as an SSD or a 3D-Xpoint drive.
[0017] Examples of the above mentioned techniques will now be described in more detail with reference to the figures.
[0018] The following is described in terms of B-trees, but more generally the ideas disclosed herein, such as different types of read and write operations, the use of big and small blocks, shortcut keys, back-pointers, ordering hints, and / or synchronizing complex updates (splits and merges), can be applied in any tree structure for storing key-value pairs.
[0019] Tree Structure Key-value stores such as B-trees store data as key-value pairs (also called items). Items may be of variable size, but depending on the implementation, the data structure may place a limit on the total allowable size of items (e.g., up to 512 bytes), or in some implementations may make them a fixed size to simplify the implementation. Each key is associated with a single respective value, and multiple key-value pairs with the same key are not allowed.
[0020] 1, 4, and 5 illustrate, by way of example, elements of a tree structure 103, such as a B-tree. The tree 103 is a form of memory-implemented data structure that can be used as a key-value store. The tree 103 comprises a number of nodes 102, some of which are leaf nodes 102L and others are internal nodes 102I. Each node 102 has a respective node ID, which may also be referred to herein as a logical ID (LID). Each leaf node 102L comprises a respective set of one or more items 101, each item 101 comprising a key-value pair (by way of illustration, FIG. 1 shows the nodes 102 undergoing a range scan, but this is not limiting and is merely one example of a type of read operation). The key-value pair serves as an index for the item, and comprises a key that is mapped to a value that is the content of the item. The numbers shown inside the items 101 in the figures are exemplary keys (in practice, the keys can reach much larger numbers, but this is merely for illustrative purposes). Each internal node 102I specifies a respective set of one or more children of the internal node. When the tree is constructed, at least one or some of the internal nodes will have two or more children each (although if the tree has only one leaf, e.g., when it is just created, the root node has only one child). Each internal node 102I also indicates what ranges of keys are encompassed by each of its respective children. The particular scheme shown in FIG. 5, in which keys are sandwiched between pointers, described below, is just one possible example for specifying this mapping.
[0021] Whatever means by which the mapping is implemented, the node IDs of the internal nodes 102I thus define the edges between parent and child nodes, thereby forming a tree structure. Each internal node 102I is the parent of at least one respective child. When the tree is constructed, at least one or some of the internal nodes each become the parent of multiple respective children (although, as mentioned above, if the tree has only one leaf, then the root has a single child). A child of a parent can be either a leaf node 102L or another internal node 102I. Leaf nodes 102L are only children, not parents (i.e., leaves are the bottom of the tree, or in other words, the terminus of the branches of the tree).
[0022] As used herein, a key is "contained" in a child means that the child is either a leaf node 102 that contains an item with that key (a key-value pair), or that the child is another internal node 102I that is connected, via a node ID-to-key mapping, ultimately to a leaf node 102I that contains an item with that key, one or more levels (or "generations") down in the hierarchy of the tree.
[0023] One of the internal nodes 102I at the top of the tree is the root node. When a writer or reader wants to write to or read from an item with a particular key, it starts by querying the mapping of node IDs to key ranges specified in the root node to find which child of the root contains the required key. If that child is itself another internal node 102I, the writer or reader queries the mapping specified in that node to find which child of the next generation contains the required key, and so on, until it finds a leaf node 102L that contains an item with the required key.
[0024] In an embodiment, the key comparison may follow the semantics of the C memcmp function, i.e., to preserve the semantics of integer comparisons, integer keys should be stored in big-endian format, but this is just one possible example.
[0025] In an embodiment, the B-tree may take the form of a B+ tree, whereby only leaf nodes 102L contain key-value pairs and only internal nodes 102I contain node ID to key range mappings.
[0026] Read / Write Operations A tree structure such as a B-tree may support read operation types of individual lookup and / or range scan. It supports write (update) operation types of insert, modify, and / or delete. Example semantics of such operations are described below.
[0027] Lookup: It takes a key as an argument and returns the value associated with that key if it exists in the tree. Otherwise, it returns a status of "not-found".
[0028] Scan: It takes two keys as input arguments and returns all key-value pairs from the tree with keys in the closed interval [low-key, high-key]. In an embodiment, two types of scans are supported: exhaustive scan and covering scan. Figures 1, 2 and 3 illustrate the difference. (l,u) means that the lower bound l and upper bound u are not included in the result even if they are in the tree. [l,u] means that they are included if they are in the tree. The combinations l is included but u is not included [l,u] and l is not included but u is included (l,u] are also possible.
[0029] In an inclusive scan (Figure 1), the interval formed by the keys of the returned items is the largest subset that falls within the input interval. An inclusive scan can be used when serving a database query such as "get all items with keys between 36 and 80". The result will not include an item with key 35, even if there is no item with key 36. In a covering scan (Figures 2 and 3), the interval formed by the keys of the returned items is the smallest superset that contains the input interval. In the case of Figure 2, an item with key 22 would be returned, but in Figure 3, it would not be returned. Covering scans are useful when querying storage metadata such as "get disk blocks that contain data in the offset range 550 to 800 from the beginning of the file". The result will contain the largest key before 550 and the smallest key after 800 (i.e., the keys exactly on either side of the specified range), i.e., all disk blocks that contain data in the requested range are returned.
[0030] Insert: This takes a new key-value pair to insert. If an item with the specified key does not exist in the tree, the key-value pair is inserted into the tree. Otherwise, this operation does not update the tree and returns the status "already-exists".
[0031] Modification: This takes an existing key-value pair to update. If an item with the specified key exists in the tree, the value is updated to the value of the input argument. Note that the value can be of a different size than before. Otherwise, the operation does not update the tree and returns a status of "not found".
[0032] Delete: This takes the key of the item to remove. If the item exists in the tree, it is removed. Otherwise, this operation returns a status of "not found".
[0033] System Architecture 4 shows an exemplary system according to an embodiment. The system comprises a writer 401, a reader 402, a system memory 403, and optionally a local memory 404 of the reader 402. The tree structure 103 is stored in the system memory 403. The writer 401 and the reader 402 are operatively coupled to the system memory 403. For example, the reader 402 may be operatively coupled to the writer 401 via a bus 405, such as a PCIe bus, while the writer 401 may be operatively coupled to the system memory 403 via a separate local or dedicated connection 406 that does not require communication via a bus. The writer 401 may be arranged to access the system memory 403 to write to the tree 103 via the local connection 406, but for the reader 402 to read from the tree, it may have to access the memory 403 via the bus 405. The reader 402 may be connected to its local memory 404 via its own local connection 407 and may thereby be arranged to cache a portion of the tree 103 in the reader's local memory 404 .
[0034] Memory 403 may take any suitable form and may comprise one or more memory media embodied in one or more memory units. For example, the memory media may comprise electronic silicon-based media such as SRAM (Static Random Access Memory) or DRAM (Dynamic RAM), EEPROM (Electrically Erasable Programmable ROM), flash memory; or magnetic media such as magnetic disks or tapes; or more exotic forms such as re-writable optical media or synthetic biological storage; or any combination of any of these and / or other types. The one or more memory units in which memory 403 is embodied may comprise an on-chip memory unit on the same chip as writer 401, or a separate unit on the same board as writer 401, or an external unit such as a solid-state drive (SSD) or hard drive, which may be implemented within the same housing or rack as writer 401, or elsewhere in a separate housing or rack, or any combination of these. In one particular embodiment, as shown, memory 403 comprises DRAM, which may be implemented on the same board or chip as writer 401. For example, in one particular example, the system uses a DRAM board that plugs into the board on which the CPU resides, but this is not intended to be limiting.
[0035] The reader's local memory 404, if used, may take any suitable form comprising any suitable medium in any one or more units. In one particular embodiment, as shown, the reader's local memory 404 comprises DRAM, which may be integrated into the reader 402 or implemented on the same chip or board as the reader 402.
[0036] In an embodiment, writer 401 is implemented in software stored in non-transitory computer-readable storage of the system and arranged to run on one or more processors, e.g., CPUs, of the system, while reader 402 is implemented in hardware, which may take the form of FPGA, PGA, or dedicated (fixed function) hardware circuitry, or a combination. Although the following may be described in terms of such an embodiment, it will be recognized that none of the techniques described below exclude the possibility that writer 401 could instead be implemented in hardware and / or reader 402 could be implemented in software.
[0037] The storage used to store the software of the writer 401 (and / or the reader in alternative embodiments where the reader is implemented in software) may take any suitable form, such as any of those mentioned above, or alternatively a read only memory (ROM), such as an electronic silicon-based ROM, or an optical medium, such as a CD ROM. The one or more processors may comprise, for example, a general purpose CPU (Central Processing Unit); or an application specific processor or accelerator processor, such as a GPU (Graphics Processing Unit), a DSP (Digital Signal Processor), a cryptographic processor, or an AI accelerator processor; or any combination of any of these and / or other types.
[0038] In one particular embodiment, as shown, writer 401 is implemented in software running on a CPU and reader 402 is implemented in an FPGA. It will be appreciated that the following embodiments may be described in terms of such examples, but this is not limiting.
[0039] B-trees store data in system memory 403. In an embodiment, this allows storing trees as large as system memory (typically hundreds of GB), at the cost of the reader 402 (e.g., FPGA) accessing system memory through PCIe 405. The reader (e.g., FPGA) 402 can keep a cache of the upper layers of the tree in its local memory (e.g., DRAM) 404 to speed up execution. In one example implementation, the reader 402 records the LID of the root node and the number of levels in the tree in its local register file.
[0040] B-trees may be designed to allow data to be stored on secondary storage as well, to allow for even larger trees to be stored. Emerging ultra-low latency Non-Volatile Memory Express (NVMe) devices make this an attractive option, even for low-latency scenarios.
[0041] In an embodiment, FPGA 402 performs only lookup and scan operations, and update operations are performed by CPU 401. Synchronization between the CPU and FPGA is preferably lightweight. Only when a B-tree node is split or merged, it may require communication over PCIe bus 405 (see Data Structures section for further details). Supporting only read operations to the FPGA allows read-dominant use cases to benefit from hardware offloading without having to implement all operations in hardware.
[0042] Data Structure In an embodiment, the B-tree 103 is a B+ tree, which means that key-value pairs 101 are stored only in the leaves 102L, and the internal nodes 102I store key indices and pointers to child nodes.
[0043] Although FIG. 5 shows an example of a two-level tree, the principles described herein can be applied to trees with any number of levels (ie, any number of parent and child levels).
[0044] The data structure comprising the tree 103 further comprises a page table 501. The page table 501 maps node IDs (also called logical IDs, LIDs) to the actual memory addresses where the respective nodes 102 are stored in the system memory 403. In this manner, a writer 401 or reader 402 can locate a node 102 in the memory 403 based on its node ID (LID) by looking up an address based on the node ID in the page table 501.
[0045] As mentioned above, each internal node 102I includes information mapping each of its child node IDs to an indication of the range of keys encompassed by the respective child. Figure 5 illustrates one example of how this may be implemented.
[0046] In this example, each internal node 102I comprises a set of node pointers 502 and key indices 503. Each node pointer 502 is a pointer to a child node specified by its node ID (logical ID, LID). Each key index 502 in this example specifies one of the keys at the boundary of the range encompassed by the child. In each internal node 102I, a pointer 502 with an index 503 of key l on its left side and key u on its right side points to a subtree that stores keys in the interval [l,u). In addition, each internal node 102I stores a pointer 502 to a subtree that has a key smaller than the smallest key in that internal node. This leftmost pointer is stored in the node's header 603 (see, for example, FIG. 6). In other words, each key index is sandwiched between a pair of pointers, and the key index specifies the boundary between the range encompassed by the children pointed to by the pointers on either side of the key index. In other words, an internal node stores a leftmost child lid. These essentially store X keys and X +l pointers. The leftmost is a +1 pointer. In the particular example shown in Figure 5, new items with keys in the interval [22,65) are inserted into the middle leaf, items with keys less than 22 are inserted into the leftmost leaf, and keys greater than or equal to 65 are inserted into the rightmost leaf.
[0047] It will be appreciated that the scheme for mapping children to ranges of keys shown in Figure 5 is only one example. Other examples would be, for example, to first store all the boundary keys in the node, and then all the LIDs follow them.
[0048] In some embodiments, each leaf 102L may also store a pointer to a leaf to its left and a pointer to a leaf to its right, to simplify range scan implementations.
[0049] A B-tree 103 grows from the bottom up (like any other B-tree). During insertion, if a new item 101 cannot fit into the leaf 102 it is mapped to, the leaf is split in two. The item in the old leaf is split in half and a pointer to the new leaf is inserted into the parent node. When the parent node becomes full, it is also split. Splitting can continue recursively up to the root of the tree. When the root is split, a new root is created, which increases the height of the tree. When a leaf becomes empty because all items have been removed from it, it is removed from the parent. Alternatively, another option is to merge leaves when they become less than half full, in order to maintain the spatial invariance of the B-tree.
[0050] Node Layout In an embodiment, the B-tree nodes 102 may be 8KB in size to align them with the DRAM page size. They may be allocated in pinned memory that cannot be moved by the operating system to allow the reader 402 (e.g., an FPGA) to access the nodes 102 using physical memory addresses. However, this is not required.
[0051] Preferably, the internal nodes 102I do not directly store the addresses of their child nodes. Instead, they store a logical identifier (LID) for each node. In one example, the LID may be 6 bytes in size, thereby limiting the maximum size of the tree to 2 61The number of bytes per page is limited. The mapping from a node's LID to the node's virtual and / or physical addresses is maintained in a page table 501. The page table 501 may be stored both in system memory 403 and in FPGA attached memory 404. When a new node's mapping is created or when a node's mapping is changed, the writer 401 (e.g., a CPU) updates the table 501 in system memory 403. If a cache is maintained on the reader side, the writer 401 also issues a command (e.g., over PCIe) to the reader 402 (e.g., an FPGA) to update its copy of the table in the reader's attached memory 404. Alternatively, the reader can poll the tree for updates.
[0052] Addressing nodes using their LIDs provides a level of indirection, such as allowing nodes to be stored on NVMe devices, and it also helps synchronize access to the tree, since in some cases update operations may perform copy-on-writes of nodes.
[0053] In an embodiment, the layout of the internal nodes 102I and the leaf nodes 102L are the same. See, for example, FIG. 6. The node header 603 may be stored, for example, in the first 32 bytes of the node 102. The header 603 may contain fields specifying the type of node 102 (internal or leaf), the number of used bytes in the node, a lock bit and a node version number for synchronizing access to the node, and / or several other fields described below. In the example shown in FIG. 5, the header 603 of the internal node 102I also stores the LID of its leftmost child. The header of the leaf node 102L may store the LIDs of its left and right siblings. Since the keys and values may be of variable size, they may be stored as blobs. Each blob has, for example, a 2-byte header that specifies its size. The LIDs may be stored without a header since their size is known. The values may be stored inline in the B-tree node 102L to improve sequential access performance. Values larger than 512 bytes may be stored outside the B-tree 103, and a pointer to the value is stored in the leaf 102L.
[0054] Regardless of the choice of specific implementation parameters such as the node size exemplified above, according to the embodiment disclosed herein, each leaf node 102L comprises two blocks: a large block 601 and a small block 602. To implement this, a pointer to the boundary between the blocks may be stored in the node's header 603. This pointer may be expressed, for example, as an offset relative to the start of the large block 601. The large block 601 stores items in sorted order, and the small block 602 stores a log of recent updates. When the small block 602 becomes larger than a threshold (e.g., set to 512 bytes), it is merged into the large block 601. By dividing the node 102L into large and small blocks, read operations benefit from accessing mostly sorted data without the overhead of sorting the node after every update. The entries in the small block 602 can be either newly inserted items or deletion markers. As an optional optimization, each entry in the small block 602 may further store a pointer (e.g., 2 bytes long) to an item in the large block 601. As a convenient label, this is sometimes called a "back pointer" (although the term "back" is not necessarily meant to limit it to a particular direction in physical address space or the like). In one example implementation, for a newly inserted item 101, the pointer points to the first item in the big block 601 with a key greater than the key of the new item added to the small block 602. Alternatively, for a deletion marker, the pointer points to the deleted item. The pointer can be expressed as an offset within the node 102. It can be used by the reader 402 to establish an order between the items 101 in the big and small blocks without comparing their keys, which is a useful optimization, especially when reading is implemented in hardware.
[0055] As another optional optimization, each item in the small block 602 may have a field (e.g., 1 byte long) that stores its index into the small block that was sorted at the time of the item's insertion; that is, this is the order of the items at the time of insertion. This is sometimes called an "order hint." These indexes are "regenerated" by the reader 402 as it scans the small block during a search to reconstruct the current sorted order. This order is stored in a temporary "indirection" array (e.g., FPGA registers) in the reader 402, which is used to access the items in sorted order. This makes the sort more efficient and does not introduce significant latency, especially when implemented in hardware.
[0056] Figure 7 shows an example of steps for inserting a new item 101 into a small block 602. Step a) shows the node 102 before insertion. In step b) the item 101 is copied into the small block 602. It is not yet visible to the reader 402. In step c) the size of the data shown in the node header 603 is changed to include the new item.
[0057] As another optional optimization, the large block 601 may be divided into roughly equal sized segments to optimize the search for keys within a node. Keys that are on the boundary of a segment are stored in a shortcut block 604, which may be stored immediately after the node header 603, along with a pointer to the boundary. The search for a key begins by scanning the shortcut block to identify segments that may contain the key. The search looks only at these segments and small blocks, which reduces the amount of data read from memory. For example, if the performance bottleneck is the PCIe bandwidth between the reader (e.g., FPGA) 402 and the system memory 403, this optimization can significantly improve performance. For keys of 32 bytes or less, the search reads roughly 1.5 KB of data from an 8 KB node, thereby providing a 5x performance improvement over a simpler implementation that scans the entire node. The shortcut key can be selected during the merging of the large block 601 and the small block 602.
[0058] Node Cache The FPGA 402 may have several gigabytes (e.g., 4 GB) of DRAM 404 attached directly to it. While this memory may not be large enough to store the entire B-tree 103 in all use cases, it may be large enough for the first few levels, so in an embodiment, it can be used as a cache to reduce B-tree latency and improve B-tree throughput. For example, the cache 404 may be allocated at, for example, 8 KB node granularity to reduce the amount of metadata, but data may be fetched in 256 byte chunks to reduce the amount of unnecessary data read over PCIe 405. The FPGA 402 may keep a 32-bit occupancy bitmap for each cache node 102 to keep track of the chunks that are in the cache 404. Page consistency between the cache 404 and the system memory 403 may be maintained by invalidating cached pages every time the page mapping is updated. Software on the CPU 401 may send a page table update command over PCIe 405 when the page mapping is updated, and the FPGA 402 may invalidate the pages by clearing the occupancy bitmap.
[0059] Page Table The reader 402 (e.g., FPGA) may maintain a local copy of the page table 501 stored in its local memory 404 (e.g., DRAM), which may be attached to the FPGA device 402, for example. In an embodiment, each entry in the table is 8 bytes in the FPGA memory but 16 bytes in the system DRAM, because in an embodiment, the copy in the system DRAM stores two addresses (virtual and physical) and the copy in the FPGA's DRAM memory stores one address (physical only).
[0060] The page table 501 entries store the memory addresses of the nodes 102 mapped to the respective LIDs. In general, these may be physical or virtual addresses depending on the implementation. In an embodiment, the writer employs virtual memory mapping, and the writer-side page table 501 stores the virtual addresses of the nodes mapped to the respective IDs. In an embodiment where a copy of the page table is kept in the FPGA 402, the entries of the FPGA page table copy may store the physical addresses of the nodes mapped to the respective LIDs. For example, with 8-byte entries and 4 GB of DRAM, the system supports a maximum of 4 TB of trees, which is large enough for the main memory trees of today's servers. Larger trees can be supported by having the reader 402 (e.g., FPGA) access the page table 501 in the system memory 403 or an NVMe device, or by increasing the node size of the B-tree.
[0061] The page table 501 is not absolutely required in all possible implementations. A less preferred alternative is to directly use memory addresses as node pointers for the internal nodes 102I.
[0062] Lookup The lookup starts at the root and scans the internal tree nodes 102I at each level to find a pointer to a node 102 at the next level. Each lookup ends at a leaf 102L. If the leaf contains the key, it returns the associated value. Otherwise, it returns a status indicating that the key is not stored in the tree.
[0063] When accessing an internal node 102I, the lookup may first fetch the first part of the node 102 (e.g., the first 512 bytes), including the node header 603 and the shortcut block 604. The lookup begins by searching the shortcut block 604 to find the last key below the target key. It follows the shortcut item's pointer to find the segment of the large block 601. The search fetches the large segment from memory and searches it to find the last key below the target key. Next (or in parallel), the lookup fetches the small block 602 from memory and searches it for the last key below the target key. If back pointers are used, the lookup follows the pointer stored in the larger of the items 101 found in the small block and large block without comparing keys. If the back pointer of the item in the small block points immediately after the item in the large block, it follows the small block's pointer, otherwise it follows the large block's pointer. If the target key is less than both the first key in the major block and the first key in the minor block, the lookup follows the leftmost pointer.
[0064] When the lookup reaches a leaf 102I, it searches it in a similar manner to an interior node, with the main difference being that it looks for an exact match in the small and large blocks, meaning that the items in the large and small blocks do not need to be ordered.
[0065] Small Block Search: A search of the small block 602 may begin by sorting the items 101 before inspecting them. Sorting allows the search of the small block to stop as soon as it encounters the first item with a key greater than the target key. In an embodiment, the sorting of the small block does not perform key comparisons. Instead, it establishes an ordering using an order hint field that is stored in each small block item. The order hint field stores the index of the items in the sorted small block at the time of the insertion of the item. These indexes are "regenerated" while scanning the small block to reconstruct the current sorted order. The established order is stored in a temporary indirection array (e.g., an FPGA register) without copying the items. The indirection array is used to access the items in the sorted order. The sorting does not introduce significant latency, especially in hardware, and it can be done in parallel while searching the shortcut block 604 and the large block 601. In software, the sorted order can be kept in a small indirection array.
[0066] FIG. 8 illustrates an exemplary insertion sequence into node 102. The numbers shown on each item 101 represent its key. Elements in the large block 601 have been omitted for clarity. The numbers above the items represent their order hints. In step a), the small block 602 in the node is initially empty. In step b), as an example, an item with key 90 is inserted, and since it is the first item in the small block, it is given an order hint of 0. Then in step c), an item with key 60 is inserted. Since it is the smallest item in the small block, it is given an order hint of 0. The order hints of the existing items are not changed. The order hint is determined from the sort order of the small block, which is established before inserting the new item. Then in step d), an item with key 30 is inserted. Again, since it is the smallest in the small block, its order hint is also 0. In step e), the last item has key 45, and its order hint is 1.
[0067] Figure 9 illustrates the sorting of small blocks 602. The figure shows the state of an indirect reference array 901. The indirect reference array stores the offsets of items within the small block, but in the figure, for illustration purposes, each offset is represented by its key.
[0068] To sort the small block 602, the items 101 are processed in the order in which they were stored, and their order hints are used to sort the items in the indirection array. The indirection array 901 stores the offsets of the items in the small block. When an item with order hint i is processed, its offset in the small block is inserted at position i in the indirection array. All items at position j≧i are shifted to the right. When the entire small block is processed, the indirection array contains the offsets of the items in sorted order. Figure 9 illustrates the indirection array sort of the small block shown in Figure 8. In one implementation, each step in the diagram is completed in a single cycle on the FPGA, in parallel with the key comparison. In the first step a), item 90 is inserted into the first slot. In the next step b), all items are shifted to the right and item 60 with order hint 0 is inserted into the first slot. In the next step d), all items are shifted to the right again to make room for item 30 with order hint 0. In the final step d), items 60 and 90 are shifted to the right and item 45, with ordering hint 1, is inserted into the slot with index 1. Final state e) shows the state of the indirection array after steps a) through d).
[0069] FIG. 10 shows an example implementation of an indirect array sorter. It shows an array that is 4 items wide. The sorter has registers that are 9 bits wide. The input to a register is either its output, the output of the register to its left, or the input offset value. The register input is selected by comparing the input index to the (constant) index of the register. If they are equal, the input offset is selected, if the input index is less than the register's index, the output of the register to its left is selected, if the input index is greater, the output of the register is written to itself. This example implementation does not cover the case of having a deleted item in a small block, but so that the input to a register can be adapted to also take the output of the register to its right.
[0070] Of course, the layout shown in Figure 10 is just one possible hardware implementation, and one of ordinary skill in the art will recognize, given the disclosure herein, that other designs may be made using conventional design tools. Software implementations on the read side 401 are also possible.
[0071] Implementing sorting on CPU 401 may follow a similar approach. The indirection array is small and fits comfortably in the L1 cache, so the operation is fast. The main difference is that it takes multiple cycles to move an item to the right in the array.
[0072] Range Scan A range scan operation traverses the tree 103 from the root as described above to find the start key of the target range. It then scans forward from the start key until it reaches the end key of the range or the end of the tree.
[0073] In an embodiment, the range scan may use sibling pointers to move between leaves 102L. In such an embodiment, each leaf 102L keeps pointers to its left and right siblings, allowing range scans to be done in both directions. However, the use of sibling pointers is not required. If sibling pointers are not used, the scan instead goes back up the tree to find the next leaf. In most cases, it only needs to look at the first parent and take the LID of the next sibling. However, if it gets to the end of the parent, it will go up to its parent to find its siblings. The scan may need to go all the way up to the root.
[0074] A range scan returns items 101 sorted by key. To keep items sorted, a scan of a leaf 102L may process large and small blocks 601 and 602 together. In a covering scan, the scan first finds the item with the largest key that is less than or equal to the start of the range in both the large and small blocks. A "standard" inclusive scan is similar, except that it starts with the smallest key that is greater than or equal to the start of the range. In either case, the scan then scans forward through the nodes in both the large and small blocks, looking for the end of the interval, and returns items that are in the interval. The scan stops when an item with a key greater than the right bound is found in both the large and small blocks. During the scan, if both the next smaller item and the next larger item are in the interval, the search can use the smaller item's back pointer to decide which one to return next. If the smaller item points to the larger item, the smaller item is returned and the search moves to the next smaller item, otherwise the larger item is returned and the search moves to the next larger item. To handle shortcut keys 604, the scan may note the end of the current large segment. In this case, at the beginning of the next segment, the key is taken from the shortcut block and the value is taken from the large block. In the middle of the segment, the key and value are in the large block.
[0075] Insert An insert operation begins with a traversal of the tree to find a node 120L into which to insert a new item 101. This traversal is similar to the traversal during lookup. The difference is that in an embodiment, it reads all items, regardless of version, to implement the semantics of the insert and modify operations. Before updating a node, writer 401 can lock it with a write lock, thereby ensuring that the state of the node is the one observed during the traversal (note that in an embodiment, memory 403 can support different types of locks, i.e., write locks and read locks, and the use of a write lock does not necessarily imply the use of a read lock). For example, a write lock can be stored in the node's header in a 32-bit lock word, which consists of a lock bit and a 31-bit node sequence number, which is incremented when the node is updated. Writer 401 can obtain a write lock by atomically setting the lock bit using an atomic compare-and-swap instruction on the lock word while keeping the sequence number the same as during the traversal. If the compare-and-swap is successful, this means that the node has not changed since the traversal, and the insertion can proceed. If the compare-and-swap is unsuccessful, the node has changed since the traversal, and the writer retries the operation by traversing the tree again.
[0076] Fast-path Insertion: In one common case, an insertion does not require splitting node 102L or merging large block 601 and small block 602. Insertion in this common case can be done in place (FIG. 7). Writer 401 atomically appends the new item by first copying small block 602 after it ends, and then adjusting the size of the small block to contain the new item 101. To do this, writer 401 may perform the following steps after obtaining a write lock: it copies the item to the small block, updates the size of the node, increments the node's sequence number, and unlocks the node after obtaining the lock. In an embodiment, the lock word in the header and the size of the small block may share the same 64-bit word, so writer 401 can increment the node version, update the size of the node, and release the lock in a single instruction. Because the new item is stored beyond the end of the small block currently specified in header 603, concurrent readers do not observe the new item while it is being copied. They observe the node with no items or with the items fully filled in.
[0077] In an embodiment, the reader 402 (e.g., an FPGA) does not cache the leaf node 102L. If it did, the writer 401 would have to invalidate the cache 404 by issuing a command over PCIe 405, which would introduce additional overhead to common case insertions.
[0078] Large and small merge: When a small block 602 becomes too large, the writer 401 merges the large block 601 and the small block 602 (see FIG. 11). It allocates a new memory buffer for the node 102L and sorts all items 101 in the node into a large block in the new buffer. In an embodiment, the new item is added to the merged large block (rather than a new small block). That is, the small block of the new node is empty after the merge. In principle, it could be done the other way (i.e., the new item is the first item added to the new small block), but storing the new item in the large block immediately results in slightly fewer large and small merges over time.
[0079] In embodiments that use shortcut keys 604, writer 401 selects shortcut keys as it sorts items (it selects them at the time because big blocks are immutable). For each item processed, writer 401 decides whether to put it in a shortcut or big block based, for example, on the number of bytes copied so far, the number of bytes remaining to copy, and / or the average size of the keys and values. It preferably maximizes the number of shortcut keys while keeping the big block segments of similar size.
[0080] Once all items are copied to the new buffer, writer 401 atomically replaces the mapping of the node's LID in page table 501 with the address of the new buffer. To update the LID mapping, in an embodiment, writer 401 updates the LID entry in both the CPU and FPGA copies of the page table. It acquires a software lock on the LID entry, issues a page table update command to the FPGA, and releases the software lock after the FPGA command is completed. Since the node's LID remains the same, there is no need to change the parent node. Finally, the writer puts the old memory buffer on a garbage collection list, unlocks it, and sets a "deleted" flag in the header to ensure that concurrent writers do not update the old buffer. An ongoing operation can still read from the old buffer, but if an operation needs to update the deleted node, it retries. When retrying, it uses the new node mapping.
[0081] The garbage collection used may be any suitable scheme, such as standard explicit garbage collection for concurrent systems, e.g. RCU, epoch-based memory manager, or hazard pointers. Figure 11 shows an example of merging of large and small blocks during an insert operation. Inserting a new item into node 102L would cause small block 602 to exceed its maximum size, so now the small block is merged into large block 601 together with the new item. Step a) shows the tree 103 before the insertion into the right-most node 102L. Step b) shows the allocation of a new buffer and merging it into it. The old pointer (dashed line) to the node is inserted. In step c) the page table is updated to point to the new buffer, and the old buffer is garbage collected (GC) once it is no longer needed by the ongoing operation.
[0082] The advantage of this approach is that complex updates such as merging can be done without the need to put a read lock on the old instance of the node at the original memory address (although some operations may still use a write lock that only prevents multiple writers and allows one writer and one or more readers to access the node). If a reader 402 attempts to read it while writer 401 is updating node 102L with a new buffer, but before the page table has been updated, reader 402 simply reads the old instance from the existing memory address currently specified in page table 501. Once the write to the new buffer (i.e., new memory location) is complete, writer 401 can then quickly update page table 501 to switch the respective page table entry for the node's LID to the new memory address. In an embodiment, this can be done in a single atomic operation. All subsequent reads then read the new instance of the node from the new address based on the updated page table entry. A similar approach can be used for other complex updates such as splitting a node or merging a node.
[0083] When a new memory buffer is allocated for a LID, its sequence number can be safely set to 0, since it is ensured that the buffer is unreachable by any operation before deleting and reusing it. To give some illustrative sizes, if large block 601 and small block 602 are merged when the small block is larger than 512 bytes and the size of the small block entry is at least 10 bytes (delete entry size is 10 bytes, insert entry size is at least 13 bytes), the maximum version of the buffer is 41, which means that the 31-bit sequence number never wraps around. In fact, in some implementations, the minimum size is indeed 6 bytes on average, but even with 1-byte entries the sequence number never overflows.
[0084] Node splitting If there is not enough space in the leaf 102L for the insertion, the writer 401 splits the leaf in two and inserts a pointer to the new item in the parent (Figure 12). If the parent is full, the writer 401 splits it as well. A split can propagate in this way all the way up to the root. An internal node 102I that is updated but not split is called the root of the split. The writer replaces the entire subtree under the root of the split by updating the mapping of the root of the split in the page table 501. Before updating the data, the writer can obtain write locks on all nodes 102 that will be updated by the operation (these are locks to prevent interference by other writers). If the locks fail due to a node version mismatch, the writer releases all the write locks it has obtained and resumes the operation. The writer also allocates all the memory buffers needed to complete the operation. It allocates two new nodes (with new LIDs and new buffers) for each node it splits, and a memory buffer for the root of the split. If the memory allocation fails due to the system being out of memory, the writer gives up and returns an appropriate status code to the caller. After acquiring the lock and allocating the memory, the writer processes the nodes from the leaf up and splits them into newly allocated buffers. Each of the two resulting nodes of the split ends up with roughly half the data of the original node. The writer merges the big and small blocks during the split. When it reaches the root of the split, the writer copies it to a new memory buffer. It modifies the new buffer to contain pointers to the new child nodes and the keys at the boundary between the nodes. To swap the two subtrees, the writer updates the mapping of the root of the split in the page table with the address of the new memory buffer. Swapping the subtrees in this way ensures that the reader observes either the old or the new subtree. The writer then puts the memory buffers of all the split nodes, their LIDs, and the old memory buffer of the root of the split into a garbage collection list.Finally, it unlocks all memory buffers and marks them as deleted.
[0085] Figure 12 shows an example of splitting a node on insertion. The dashed box represents the newly allocated node buffer. The filled box represents the new node with new LID and memory buffer. Step a) shows the tree 103 before the split. The right-most node is to be split. The root of the split is the root of the tree. In step b) two new nodes are allocated for the split and a new buffer is allocated for the root of the split. In step c) the mapping for the root of the split is updated and the node is put on the garbage collection (GC) list.
[0086] In an embodiment, the writer also updates the sibling leaf pointers used during the scan operation. It locks the sibling leaf and updates the pointers after it swaps in the new subtree. Even though the sibling pointers and the root of the split may not be updated atomically, the reader observes a consistent state of the tree.
[0087] If the root of the tree cannot accommodate the new item, the writer splits it to create a new root, thereby increasing the height of the tree. The writer performs the split as above, except that instead of only a memory buffer for the root of the split, it allocates a LID and memory buffer for the new root. It initializes the new root with its leftmost pointer set to the left half of the old root and a single item pointing to the right half of the old root. It may then send a command to the reader (e.g., an FPGA) to update the root and the height of the tree. In one example implementation, the reader 402 records the LID of the root node and the number of tree levels in its local register file, rather than in the cache 404. Data copies in the cache, e.g., older root nodes, are eventually invalidated when their LIDs are reused for another physical node, or are possibly evicted because the old data is no longer accessed after a while. Thus, the cache only needs to focus on maintaining the consistency of the cached node data.
[0088] delete Deleting an item 101 is similar to inserting a new item. The writer 401 inserts a delete entry into the small block 602. This entry points to the item it is deleting, so the deleted item is ignored during lookup operations. During the merge of the large block 601 and the small block 602, the space occupied by the deleted item is reclaimed. When a node 102 becomes empty, it is deleted. Node deletion can proceed from the leaf 102L up the tree 103, similar to node splitting. In an embodiment, the same techniques as above are used to ensure its atomicity.
[0089] Fixes Modification operations can be performed as a combination of insertions and deletions: the writer 401 atomically publishes the deletion entries of the old item and the new item by appending them to the small block 602 and updating the node's header 603.
[0090] conclusion It will be appreciated that the above embodiments are described by way of example only. More generally, according to one aspect disclosed herein, a system is provided comprising a memory for storing a data structure, the data structure comprising a tree structure comprising a plurality of nodes each having a node ID, some nodes being leaf nodes and other nodes being internal nodes, each internal node being a parent of a respective set of one or more children in the tree structure, each child being either a leaf node or another internal node, each leaf node being a child but not a parent, each leaf node comprising a respective set of one or more items each having a key-value pair, and each internal node mapping each node ID of each of its respective children to a range of keys encompassed by its respective child. The system further comprises a writer arranged to write items to the leaf nodes and a reader arranged to read items from the leaf nodes. Each of the leaf nodes comprises a respective first block and a respective second block, the first block comprising a plurality of items of the respective leaf node sorted in key order in the address space of the memory. The writer is configured, when writing new items to the tree structure, to identify which leaf node to write to by following the mapping of keys to node IDs through the tree structure, and to write the new items to the second block of the identified leaf node in the order in which they were written, rather than sorted in key order. The reader, when reading one or more target items from the tree structure, is configured, when reading from the tree structure, to determine which leaf node to read from by following the mapping of keys to node IDs through the tree structure, and then to search for the determined leaf node for the one or more target items based on a) the order of the items already sorted in the first block, and b) the reader sorting the items in the second block by keys associated with the items in the first block.
[0091] In an embodiment, the writer may be implemented in software stored in computer readable storage of the system and arranged to run on one or more processors of the system. In an embodiment, the reader may be implemented in a PGA, FPGA, or dedicated hardware circuitry. However, alternatively, the writer may be implemented in hardware and / or the reader may be implemented in software, or either may be implemented in a combination of software and hardware.
[0092] In an embodiment, the tree structure may take the form of a B-tree (eg, a B+ tree where items are not stored in intermediate nodes).
[0093] In an embodiment, in at least some of the leaf nodes and / or internal nodes, each first block may be divided into segments, and the node may further comprise a shortcut block comprising a plurality of shortcuts, each shortcut comprising i) an indication of a key at a boundary of a corresponding segment of the respective first block, and ii) an offset to the corresponding segment in the leaf node. In such an embodiment, the reader may be configured to use the shortcuts when performing said search of the first blocks to search only those segments that are likely to contain one or more items based on the key.
[0094] In an embodiment, the writer, when writing to at least one of the leaf nodes, may be configured to include in the write a back-pointer associated with each item in the respective second block, each back-pointer comprising an offset to an item in the first block having a next higher or lower key compared to the associated item in the second block. In such an embodiment, the reader, when performing said sorting of the second blocks, may be configured to use the back-pointers to sort the items of the second block by keys associated with the respective first blocks.
[0095] In an embodiment, the writer may be configured, when writing to at least one of the leaf nodes, to include, for each item in the second block, an order hint when writing, the order hint indicating an order of the item's key relative to the current key of the second block at the time of writing the item to the second block. In such an embodiment, the reader may be configured, when performing said sorting of the second block, to use the order hint to sort the items of the second block by the keys associated with their respective first blocks.
[0096] For example, a reader may be configured to use the ordering hints by processing the items in each second block in the order in which they were stored, and using the ordering hints to sort the items in the array, where the array stores and uses the offsets of the items in the second block. As each item with ordering hint i is processed, its offset in the second block is inserted at position i in the array, and all items at positions j>i are shifted to the right so that after the entire second block has been processed, the array contains the offsets of the items in sorted order.
[0097] In embodiments, the data structure may comprise a page table that maps node IDs to respective memory addresses, where the reader is configured to use the page table to determine a memory address of the determined node in memory when reading from the node of the determined node ID. In some such embodiments, the writer may be configured to write an updated version of the node to a new memory address when updating the node to write, split, or merge the node, and, once written, update the respective memory address in the page table to point to the new memory address.
[0098] In an embodiment, each second block may have a maximum size. The writer may be configured to merge the second block into the respective first block when writing a new item that causes the respective second block of the identified node to exceed the maximum size.
[0099] In an embodiment in which the data structure comprises a page table that maps node IDs to respective memory addresses, the writer may be configured, when updating a node to merge respective second blocks and first blocks, to write an updated version of the node including the merge to a new memory address, and, once the merge is complete, to update the respective memory addresses in the page table to point to the new memory addresses.
[0100] In an embodiment, each leaf node may further indicate the node IDs of one or more sibling leaf nodes that encompass a range of keys adjacent to the key of the respective leaf node. The reader may also be configured to perform a range scan to read items from a range of keys. In such an embodiment, the reader may be configured, when performing a range scan to read from a range of items that spans two or more leaf nodes, to determine one of the leaf nodes that encompasses one of the keys in the scanned range by following the mapping of keys to node IDs through the tree structure, and then to determine at least one other leaf node that encompasses at least one other key in the scanned range by using the ID of at least one of the sibling leaf nodes indicated in said one of the leaf nodes.
[0101] According to another aspect disclosed herein, a method is provided comprising operation of a memory, writer, and / or reader according to any embodiment disclosed herein. According to another aspect, a computer program may be provided embodied on one or more non-transitory computer-readable media, the computer program comprising code configured to write the writer and / or read the reader when executed on one or more processors. Other variations or applications of the disclosed techniques may be apparent to those of skill in the art given the disclosure herein. The scope of the present disclosure is not limited by the described embodiments, but only by the appended claims.
Claims
1. a memory for storing a data structure, the data structure comprising a tree structure with a number of nodes, each having a node ID, some nodes being leaf nodes and other nodes being internal nodes, each internal node being the parent of a respective set of one or more children in the tree structure, each child being either a leaf node or another internal node, each leaf node being a child but not a parent, each of the leaf nodes comprising a respective set of one or more items, each with a key-value pair, and each of the internal nodes mapping the node ID of each of its respective children to a range of keys encompassed by its respective child; a writer arranged to write items to said leaf nodes; a reader arranged to read items from said leaf nodes; A system comprising: each of the leaf nodes comprises a respective first block and a respective second block, the first block comprising a plurality of the items of the respective leaf node sorted in key order in an address space of the memory; the writer is configured, when writing a new item to the tree structure, to identify which leaf node to write to by walking through the tree structure through the mapping of keys to node IDs, and to write the new item to the second block of the identified leaf node in written order rather than sorted in key order; the reader is configured, when reading one or more target items from the tree structure, to determine which leaf node to read from by following the mapping of keys to node IDs through the tree structure, and then searching for the determined leaf node for the one or more target items based on a) the order of the items already sorted in the first block, and b) the reader sorting the items in the second block by keys associated with the items in the first block. system.
2. The system of claim 1 , wherein the writer is implemented in software stored in computer readable storage of the system and arranged to be executed on one or more processors of the system.
3. The system of claim 1 or 2, wherein the reader is implemented in a PGA, an FPGA, or a dedicated hardware circuitry.
4. The system of claim 1 or 2, wherein the tree structure takes the form of a B-tree.
5. 5. The system of claim 4, wherein the B-tree is a B+ tree in which items are not stored in intermediate nodes.
6. In at least some of the leaf nodes and / or internal nodes, the respective first blocks are divided into segments, and the nodes further comprise a shortcut block comprising a plurality of shortcuts, each shortcut comprising: i) an indication of the key at a boundary of a corresponding segment of the respective first block; and ii) an offset to the corresponding segment within the leaf node; 3. The system of claim 1 or 2, wherein the reader, when performing the search of the first block, is configured to use the shortcut to search only those segments that are likely to contain the one or more items based on a key.
7. the writer is configured, when writing to at least one of the leaf nodes, to include in the write a back-pointer associated with each item in the respective second block, each back-pointer comprising an offset to an item in the first block having a next higher or lower key compared to the associated item in the second block; 3. The system of claim 1 or 2, wherein the reader is configured to use the back-pointer when performing the sorting of the second block to sort the items of the second block by keys associated with the respective first blocks.
8. the writer is configured, when writing to at least one of the leaf nodes, to include, for each of the items in the second block, an order hint when writing that indicates an order of the key of the item relative to the current key of the second block at the time of writing the item to the second block; 3. The system of claim 1 or 2, wherein the reader is configured to use the order hints to sort the items in the second block by keys associated with the respective first blocks when performing the sorting of the second blocks.
9. The reader, - processing the items in each of said second blocks in the order in which they were stored; - using the order hint to sort the items in an array, the array storing offsets of the items in the second block; and configured to use the order hint by 9. The system of claim 8, wherein as each item with order hint i is processed, its offset in the second block is inserted at position i in the array, and all of the items at position j>i are shifted right so that after the entire second block has been processed, the array contains the offsets of the items in the sorted order.
10. the data structure comprising a page table that maps the node IDs to respective memory addresses; the reader is configured to use the page table to determine the memory address of the determined node in memory when reading from a node of a determined node ID; 3. The system of claim 1 or 2, wherein the writer is configured to write an updated version of the node to a new memory address when updating the node to write, split, or merge the node, and, once written, update the respective memory address in the page table to point to the new memory address.
11. each second block has a maximum size; 3. The system of claim 1, wherein if the writer writes a new item that causes the respective second block of the identified node to exceed the maximum size, the writer is configured to merge the second block into the respective first block.
12. the data structure comprising a page table that maps the node IDs to respective memory addresses; the reader is configured to use the page table to determine the memory address of the determined node in memory when reading from a node of a determined node ID; 12. The system of claim 11, wherein the writer is configured, when updating a node to merge the respective second and first blocks, to write an updated version of the node including the merge to a new memory address, and, once the merging is complete, to update the respective memory addresses in the page table to point to the new memory addresses.
13. each leaf node further indicates the node IDs of one or more sibling leaf nodes that encompass a range of keys adjacent to the key of the respective leaf node; The reader is configured to perform a range scan to read items from a range of keys, and when performing a range scan to read from a range of items that spans two or more leaf nodes: - determining one of the leaf nodes that contains one of the keys in the scanned range by following the mapping of keys to node IDs through the tree structure; - determining at least one other leaf node that contains at least one other key of the scanned range by using the ID of at least one of the sibling leaf nodes indicated in the one of the leaf nodes; The system according to claim 1 or 2, configured to:
14. storing a data structure in memory, the data structure comprising a tree structure with a number of nodes, each having a node ID, some nodes being leaf nodes and other nodes being internal nodes, each internal node being a parent to a respective set of one or more children in the tree structure, each child being either a leaf node or another internal node, each leaf node being a child but not a parent, each of the leaf nodes comprising a respective set of one or more items, each with a key-value pair, each of the internal nodes mapping the node ID of each of its respective children to a range of keys encompassed by its respective child; writing an item to the leaf node; Retrieving an item from the leaf node; A method comprising: each of the leaf nodes comprises a respective first block and a respective second block, the first block comprising a plurality of the items of the respective leaf node sorted in key order in an address space of the memory; wherein the writing comprises: when writing a new item to the tree structure, identifying which leaf node to write to by following the mapping of keys to node IDs through the tree structure; and writing the new item to the second block of the identified leaf node in written order rather than sorted in key order; the reading comprises: determining which leaf node to read from by following the mapping of keys to node IDs through the tree structure when reading one or more target items from the tree structure; and then searching the determined leaf node for the one or more target items based on a) the order of the items already sorted in the first block and b) a reader sorting the items in the second block by keys associated with the items in the first block.
15. 15. A computer program embodied on one or more non-transitory computer readable media, the computer program comprising code configured, when executed on one or more processors, to perform the writing and / or reading of claim 14.