Data query method and device, computer equipment and storage medium

By predicting and batch cache of inode nodes to memory buffers before querying, the problem of frequent foreign memory access in large-scale database queries is solved, and query efficiency and performance are improved.

CN120234341APending Publication Date: 2025-07-01SHENZHEN TENCENT COMP SYST CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202311871891.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2023-12-29
Publication Date
2025-07-01

AI Technical Summary

Technical Problem

In large-scale database queries, frequent external memory accesses affect query efficiency because the inode is stored in the external memory device, especially when the memory buffer is missed, the node needs to be read from the external memory, resulting in high dependency access delay.

Method used

The prefetch model predicts the node number corresponding to the target search key and batch caches these nodes to the memory buffer to avoid the need to read from the memory every time the access is accessed, and batch prefetching is used to improve query efficiency.

Benefits of technology

Reduces the number of reads to memory, improves the efficiency of data query, alleviates the problem of pointer pursuit, improves index access performance, and reduces I/O access latency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120234341A_ABST
    Figure CN120234341A_ABST
Patent Text Reader

Abstract

The invention discloses a data query method and device, computer equipment and a storage medium, and belongs to the technical field of computers. The method comprises the steps that an index and a target search key of to-be-queried data are obtained, the index comprises multiple layers, each layer comprises nodes, each node comprises a search key and has a number, the multiple layers comprise multiple prefetching layers, the multiple prefetching layers correspond to prefetching models, and the prefetching models are used for predicting the number corresponding to any search key; predicting a plurality of target numbers corresponding to the target search key through a prefetching model; caching a plurality of target nodes indicated by the plurality of target numbers from a memory to a memory buffer area; and querying the target search key in the index, and accessing the cached target node in the memory buffer area under the condition that the node needing to be accessed currently belongs to the target node. By caching the nodes in batches in advance, the problem that the memory needs to be read once when each node is accessed can be avoided, and the query efficiency is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and in particular to a data query method, apparatus, computer equipment, and storage medium. Background Art

[0002] With the development of computer technology, more and more databases are accelerating data queries by creating and maintaining indexes. Generally speaking, the index structure contains multiple layers of nodes, in which the search key and storage address of the data are recorded. By querying a search key in the index, the storage address of the corresponding data can be obtained, and the data can be accessed according to the storage address.

[0003] However, due to the existence of the hierarchical structure, when accessing the index, it is necessary to access the nodes layer by layer from the top layer to the bottom layer, and the access to the nodes of adjacent layers has a dependency relationship. When the data scale is large, the index nodes are usually stored in the external memory device. When accessing each layer of nodes, if the memory buffer cannot hit, the node needs to be read from the external memory device to the memory buffer, causing a large number of interdependent external memory accesses, which in turn affects the efficiency of data query. Summary of the invention

[0004] The embodiments of the present application provide a data query method, device, computer equipment and storage medium, which can improve the efficiency of data query. The technical solution is as follows:

[0005] In one aspect, a data query method is provided, the method comprising:

[0006] Obtaining an index and a target search key of the data to be queried, wherein the index includes multiple layers, each layer includes nodes, each node includes a search key, each node has a number, the multiple layers include multiple pre-fetch layers, the multiple pre-fetch layers correspond to pre-fetch models, the pre-fetch model is used to predict the number corresponding to any search key, and the number corresponding to the search key refers to the number of the node in the pre-fetch layer located on the access path when querying the search key;

[0007] Predicting multiple target numbers corresponding to the target search key through the pre-fetch model;

[0008] caching the multiple target nodes indicated by the multiple target numbers from the memory to the memory buffer;

[0009] The target search key is queried in the index, and when the node currently to be accessed belongs to the target node, the cached target node is accessed in the memory buffer.

[0010] In another aspect, a data query device is provided, the device comprising:

[0011] An acquisition module is used to acquire an index and a target search key of data to be queried, wherein the index includes multiple layers, each layer includes nodes, each node includes a search key, each node has a number, the multiple layers include multiple pre-fetch layers, the multiple pre-fetch layers correspond to pre-fetch models, the pre-fetch model is used to predict the number corresponding to any search key, and the number corresponding to the search key refers to the number of the node in the pre-fetch layer located on the access path when the search key is queried;

[0012] A prediction module, used for predicting a plurality of target numbers corresponding to the target search key through the pre-fetch model;

[0013] A cache module, used for caching the multiple target nodes indicated by the multiple target numbers from the memory to the memory buffer;

[0014] A query module is used to query the target search key in the index, and when the node currently required to be accessed belongs to the target node, access the cached target node in the memory buffer.

[0015] Optionally, the cache module is used to:

[0016] Obtaining a mapping table, wherein the mapping table stores a mapping relationship between the numbers of the nodes in the pre-fetch layer and the storage addresses of the nodes;

[0017] Querying the mapping table for multiple target storage addresses corresponding to the multiple target numbers;

[0018] The target nodes at the multiple target storage addresses are read in the memory, and the read multiple target nodes are cached in the memory buffer.

[0019] Optionally, each pre-fetch layer corresponds to a pre-fetch model, and the prediction module is used to:

[0020] For a target prefetch layer among the multiple prefetch layers, a target number corresponding to the target search key in the target prefetch layer is predicted through a prefetch model corresponding to the target prefetch layer. The target number corresponding to the target prefetch layer refers to the number of the node in the target prefetch layer located on the access path when querying the target search key. The target prefetch layer is any one of the multiple prefetch layers.

[0021] Optionally, the prediction module is used to:

[0022] Inputting the target search key into the pre-fetch model to obtain a triplet output by the pre-fetch model, wherein the triplet includes a reference number, an upper limit of a number error, and a lower limit of a number error;

[0023] The sum of the reference number and the upper limit of the number error is determined as the maximum number, and the difference between the reference number and the lower limit of the number error is determined as the minimum number;

[0024] A target number is determined among a plurality of numbers between the minimum number and the maximum number.

[0025] Optionally, the multiple search keys in each node are arranged in ascending order, and the multiple nodes in each layer are arranged in ascending order according to the search keys in the nodes; the device further comprises a creation module, which is used to:

[0026] Numbering the multiple nodes according to the arrangement order of the multiple nodes in the target pre-fetch layer;

[0027] Based on the minimum search key of each node in the target prefetch layer and the number of the node, a prefetch model corresponding to the target prefetch layer is created, and the prefetch model is used to linearly approximate the functional relationship between the minimum search keys of multiple nodes in the target prefetch layer and the numbers of the nodes.

[0028] Optionally, the prediction module is used to:

[0029] Determine at least one candidate number among a plurality of numbers between the minimum number and the maximum number, wherein the minimum search key in the node indicated by the candidate number is smaller than the target search key;

[0030] Among the at least one candidate number, the candidate number with the largest minimum search key in the indicated node is determined as the target number.

[0031] Optionally, the multiple search keys in each node are arranged in ascending order, and the multiple nodes in each layer are arranged in ascending order according to the search keys in the nodes; the device further comprises a creation module, which is used to:

[0032] Numbering the multiple nodes according to the arrangement order of the multiple nodes in the target pre-fetch layer;

[0033] Based on the maximum search key of each node in the target prefetch layer and the number of the node, a prefetch model corresponding to the target prefetch layer is created, and the prefetch model is used to linearly approximate the functional relationship between the maximum search keys of multiple nodes in the target prefetch layer and the numbers of the nodes.

[0034] Optionally, the prediction module is used to:

[0035] Determine at least one candidate number among a plurality of numbers between the minimum number and the maximum number, wherein the maximum search key in the node indicated by the candidate number is greater than the target search key;

[0036] Among the at least one candidate number, the candidate number with the smallest maximum search key in the indicated node is determined as the target number.

[0037] Optionally, the memory has a maximum number of requests, the maximum number of requests refers to the maximum number of access requests that the memory can process simultaneously, and the multiple target nodes are arranged in sequence according to the access order; the cache module is used to:

[0038] In the case that the number of the plurality of target nodes is greater than the maximum request number, determining the target nodes with the maximum request number in front according to the arrangement order of the plurality of target nodes, and caching the determined target nodes from the memory to the memory buffer;

[0039] The cache module is further used for:

[0040] When the maximum number of target nodes cached in the memory buffer have been accessed, the number of target nodes not greater than the maximum number of target nodes that have not been cached is determined, and the determined target nodes are cached from the memory to the memory buffer until the multiple target nodes have been cached in the memory buffer.

[0041] Optionally, the leaf nodes in the last layer of the index include all search keys, the multiple pre-fetch layers include the last layer of the index, the multiple nodes in each layer are arranged in ascending order according to the smallest search key in the node, and the target search key of the data to be queried is a search key between a lower bound search key and an upper bound search key of a target search key range;

[0042] The pass prediction module is used to:

[0043] Predicting, by means of the pre-fetch model, a plurality of target numbers corresponding to the lower bound search key and an upper bound leaf node number corresponding to the upper bound search key, the plurality of target numbers including a lower bound leaf node number, the lower bound leaf node number referring to the number of the leaf node where the lower bound search key is located in the last layer of the index, and the upper bound leaf node number referring to the number of the leaf node where the upper bound search key is located in the last layer of the index;

[0044] The cache module is used for:

[0045] The multiple target nodes indicated by the multiple target numbers and the multiple leaf nodes indicated by the multiple leaf node numbers between the lower boundary leaf node number and the upper boundary leaf node number are cached from the memory to the memory buffer.

[0046] Optionally, the pre-fetch model is created based on the minimum search key of each node in the target pre-fetch layer and the number of the node where the node is located; the device also includes an updating module, which is used to:

[0047] In response to a modification operation on a node in the prefetch layer, modify the node in the prefetch layer, and add a modification record corresponding to the modification operation to a log;

[0048] In the case where the update condition is currently met, based on the modification record in the log, determining the search key and number of the node in the updated prefetch layer, and updating the prefetch model based on the search key and number of the node in the updated prefetch layer;

[0049] Delete the modification record in the log.

[0050] Optionally, the multiple pre-fetch layers further correspond to a mapping table, and the mapping table stores a mapping relationship between the numbers of the nodes in the pre-fetch layers and the storage addresses of the nodes; the updating module is further used to:

[0051] In the case where the update condition is currently met, the mapping relationship between the serial number of the node in the pre-fetch layer and the storage address stored in the mapping table is updated based on the modification record in the log.

[0052] Optionally, the update condition includes any of the following:

[0053] The amount of data of the modification records stored in the log reaches a preset threshold;

[0054] The cache hit rate of the memory buffer is less than a preset hit rate, and the cache hit rate refers to the probability that a node to be accessed has been cached in the memory buffer.

[0055] On the other hand, a computer device is provided, comprising a processor and a memory, wherein the memory stores at least one computer program, and the at least one computer program is loaded and executed by the processor to implement the operations performed by the data query method described in the above aspects.

[0056] On the other hand, a computer-readable storage medium is provided, in which at least one computer program is stored. The at least one computer program is loaded and executed by a processor to implement the operations performed by the data query method described in the above aspects.

[0057] On the other hand, a computer program product is provided, including a computer program, wherein the computer program is loaded and executed by a processor to implement the operations performed by the data query method as described in the above aspects.

[0058] The method, apparatus, computer equipment and storage medium provided by the embodiments of the present application, the pre-fetch layer in the index corresponds to a pre-fetch model, when it is necessary to query the target search key of the data to be queried in the index, the target number of the node in the pre-fetch layer located on the access path when querying the target search key is first predicted by the pre-fetch model, and then multiple target nodes indicated by the target number are pre-cached in batches in the memory buffer. In the subsequent process of querying the target search key in the index, when the target node needs to be accessed, since the target node has been cached in the memory buffer in advance, there is no need to read the node from the memory, and the target node can be directly accessed in the memory buffer. The present application can avoid the problem of having to read the memory once for each node access by caching nodes in batches in advance, which is beneficial to improving the efficiency of data query. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0060] Figure 1 It is a schematic diagram of an implementation environment provided by an embodiment of the present application;

[0061] Figure 2 is a flow chart of a data query method provided by an embodiment of the present application;

[0062] Figure 3 is a flow chart of another data query method provided by an embodiment of the present application;

[0063] Figure 4 is a schematic diagram of a data query method provided in an embodiment of the present application;

[0064] Figure 5 is a flow chart of another data query method provided by an embodiment of the present application;

[0065] Figure 6 is a schematic diagram of another data query method provided in an embodiment of the present application;

[0066] Figure 7 is a flow chart of another data query method provided by an embodiment of the present application;

[0067] Figure 8 is a flow chart of a prefetcher updating method provided by an embodiment of the present application;

[0068] Fig. 9 is a comparison diagram of a data query method provided in an embodiment of the present application;

[0069] Fig.10 is a schematic diagram of another data query method provided in an embodiment of the present application;

[0070] Fig.11 It is a structural schematic diagram of a data query device provided in an embodiment of the present application;

[0071] Fig.12 is a structural schematic diagram of another data query device provided in an embodiment of the present application;

[0072] Fig.13 is a schematic diagram of the structure of a terminal provided in an embodiment of the present application;

[0073] Fig.14 It is a structural diagram of a server provided in an embodiment of the present application. DETAILED DESCRIPTION

[0074] In order to make the objectives, technical solutions and advantages of the embodiments of the present application clearer, the implementation methods of the present application will be further described in detail below in conjunction with the accompanying drawings.

[0075] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned description of the drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. It should be understood that the data used in this way can be interchangeable where appropriate, so that the embodiments of the present application described herein can be implemented in an order other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present application. On the contrary, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the attached claims.

[0076] Among them, at least one refers to one or more than one, for example, at least one node can be one node, two nodes, three nodes, or any other integer greater than or equal to one. Multiple refers to two or more than two, for example, multiple nodes can be two nodes, three nodes, or any other integer greater than or equal to two. Each refers to each of at least one, for example, each node refers to each node in the multiple nodes, and if the multiple nodes are 3 nodes, each node refers to each node in the 3 nodes.

[0077] It should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals (including but not limited to signals transmitted between user terminals and other devices, etc.) involved in this application are all fully authorized by users or relevant parties, and the collection, use and processing of relevant data need to comply with relevant laws, regulations and standards of relevant countries and regions.

[0078] Cloud technology refers to a hosting technology that unifies hardware, software, network and other resources within a wide area network or local area network to achieve data computing, storage, processing and sharing. Cloud technology is a general term for network technology, information technology, integration technology, management platform technology, application technology, etc. based on the cloud computing business model. It can form a resource pool that can be used on demand and is flexible and convenient. Cloud technology will become an important support. The background services of technical network systems require a large amount of computing and storage resources, such as video websites, picture websites and more portal websites. With the high development and application of the Internet industry, in the future, each item may have its own identification mark, and all need to be transmitted to the background system for logical processing. Data of different levels will be processed separately. All kinds of industry data require strong system backing support, which can only be achieved through cloud computing.

[0079] In short, a database can be seen as an electronic filing cabinet, that is, a place to store electronic files, where users can add, query, update, delete, etc. The so-called "database" is a collection of data that is stored together in a certain way, can be shared with multiple users, has as little redundancy as possible, and is independent of the application.

[0080] A database management system (DBMS) is a computer software system designed for managing databases. It generally has basic functions such as storage, retrieval, security, and backup. Database management systems can be classified according to the database model they support, such as relational, XML (Extensible Markup Language); or according to the type of computer they support, such as server clusters, mobile phones; or according to the query language used, such as SQL (Structured Query Language); or according to the performance focus, such as maximum scale, maximum operating speed; or other classification methods. Regardless of the classification method used, some DBMS can cross categories, for example, supporting multiple query languages ​​at the same time.

[0081] An embodiment of the present application provides a data query method, and the data queried in the embodiment of the present application may be data stored in a database based on the above-mentioned cloud technology.

[0082] For ease of understanding, the concepts involved in the embodiments of the present application are explained below:

[0083] A database index is a data structure used to speed up access to data operations. The B+ tree involved in the embodiment of the present application is also a widely used database index. Usually, when storing data, the search key of the data (also known as the Key, also known as the index value or keyword of the data) and the storage address of the data (or the data itself, also known as the Value, also known as the value of the data) are stored in the index structure. Although creating and maintaining the index structure requires additional space costs, the index can speed up the subsequent query of the data and avoid the high cost of directly scanning the data for querying. When it is necessary to query data, you can first query the index through the search key of the data to obtain the storage address of the data. In addition, data structures such as database indexes are also widely used in other fields, such as key-value storage, file systems, and search engines.

[0084] Sorted Index is a type of index among many index structures. Its most prominent feature is that in addition to supporting point queries, it also supports efficient range queries and prefix queries, etc. This is mainly achieved by explicitly organizing the index structure according to the order of the data. Point query is to query the existence of data based on the search key of the data, and return the storage address of the data if the data exists. Range query is to return the storage address of all data within a search key range given a search key range. Prefix query is to return the maximum value of all data less than the search key and its corresponding storage address given a search key. Other types of index structures, such as hash indexes, only support point queries and cannot support efficient range queries or prefix queries.

[0085] Learned Index is a new type of index structure built using machine learning methods and ideas. Learned index regards the index structure itself as taking the search key as input and the corresponding location information stored in the index structure as output. From this perspective, various machine learning models can be used to replace or accelerate the retrieval of the index model, and the functional relationship between input and output can be fitted to build the index. Through faster and simpler model calculation operations, the preliminary sequential search and binary search in the traditional index are replaced, thereby speeding up the overall query efficiency. When actually executing the query, the learned index predicts the location of the search key through the model. If there is an error in the model, additional search correction is required, and the output is finally obtained after searching through multiple layers of models.

[0086] B+-tree is a balanced search tree structure that can keep data stable and orderly. Its internal nodes are responsible for storing a set of search keys and node pointers for navigation, and leaf nodes are responsible for storing a set of ordered search keys and corresponding data (or data storage addresses). All leaf nodes are connected through linked lists to support efficient range queries. Given a search key to be queried, a binary search is performed in the nodes layer by layer from the root node to obtain the node pointer until the leaf node where the search key is located is found. This is a query process. In order to ensure the balance of the B+ tree, the insertion and deletion operations of the B+ tree may require splitting and merging nodes. The space complexity of the B+ tree is O(N) (N is the number of nodes), and the time complexity of the search operation, insertion operation, update operation, and deletion operation is O(logN). In addition, based on the basic structure of the B+ tree, some cache-friendly variant structures can also be derived.

[0087] Radix Tree is a multi-branch tree structure. Each node has multiple child nodes, each representing a different prefix. The characters on the path from the root node to a certain node are connected to represent the string corresponding to the node, which can also be understood as the search key corresponding to the node. The time complexity of its operation is O(K) (K is the length of the search key). At the same time, strings with a common prefix only need to be stored once, saving a lot of space. As an optimized structure of the radix tree, the Adaptive Radix Tree (ART) can adaptively adjust the data structure of the node according to the node size, has higher space utilization, and proposes delayed expansion and path compression optimization strategies for long strings.

[0088] Pointer-Chasing is a problem that exists in most pointer-intensive index structures (such as B+ trees, radix trees, skip lists, etc.). These indexes have similar hierarchical structures (such as tree structures, linked structures, etc.). When executing queries, it is necessary to dereference pointers layer by layer from the top layer (such as root nodes, linked list heads, etc.) to the bottom layer (such as leaf nodes, linked list tails, etc.), causing a large number of interdependent memory or external memory accesses, affecting the overall performance of the index.

[0089] Memory-Level Parallelism (MLP) is an architecture supported by most CPUs (Central Processing Units) that allows the CPU to issue multiple independent memory access requests in parallel. The memory access times of these requests overlap with each other, and ultimately multiple memory access results can be returned in one memory access cycle.

[0090] Prefetch can be divided into hardware prefetch and software prefetch according to the specific implementation. The specific process of hardware prefetch is transparent to the application. It is mainly performed by the hardware to predict the address that may be accessed according to the access mode, and prefetch the content of the corresponding address unit. Software prefetch refers to inserting prefetch instructions at appropriate locations in the program to ensure that the content of a certain address has been prefetched in advance when accessing it, which can reduce certain memory access overhead. The process of the prefetch node involved in the embodiment of the present application belongs to software prefetch.

[0091] Solid State Drive (SSD) is a high-performance non-volatile storage device. SSD uses flash memory chips and parallel read and write operations, and has the advantages of low latency, low power consumption, and high reliability. It is widely used in the new generation of data centers.

[0092] Index is one of the key technologies to improve database access performance. Indexes are mostly pointer-intensive data structures, such as B+ trees, radix trees, etc., which provide relatively stable time complexity and space complexity. However, due to the pointer-intensive structural characteristics of the index itself, there is a common pointer chasing problem, which affects the access performance of the index. Due to the existence of the pointer chasing problem, each query needs to start from the root node and dereference the pointer of the current node to the child node layer by layer until the leaf node. When the database contains a large amount of data, the memory capacity is limited, and the indexed nodes and data are stored in the external memory (such as a solid-state drive), so the nodes in the external memory need to be cached in the memory buffer. In the worst case, the nodes on the access path are not in the memory buffer, so each access to a node requires sending an I / O request to the external memory to read the node, resulting in a decrease in data query performance.

[0093] Based on this, the embodiment of the present application proposes a lightweight learning-based solution for pre-fetching nodes, which accelerates access with less time and space overhead without changing the basic structure and access process of the index. Before querying the search key, the nodes to be accessed are reasonably pre-fetched through an auxiliary pre-fetching model to improve the overall performance. For the update of the pre-fetching model, the embodiment of the present application proposes a log-based asynchronous update strategy, which can ensure that the accuracy of pre-fetching under dynamic workloads is relatively stable with less overhead. At the same time, the method provided in the embodiment of the present application has certain versatility and can be extended to different storage levels and different index structures. For detailed descriptions, please refer to the following embodiments.

[0094] Figure 1 is a schematic diagram of an implementation environment provided by an embodiment of the present application, see Figure 1 The implementation environment includes: a terminal 101 and a server 102, and the terminal 101 and the server 102 are connected via a wireless or wired network.

[0095] The server 102 is connected to a database, the terminal 101 is used to send a data query request to the server 102, the server 102 is used to obtain data corresponding to the data query request, and the database is used to store data and indexes. It should be noted that the embodiment of the present application does not limit the number of databases 102. For example, the database used to store data and the database used to store indexes are in the same database.

[0096] Among them, the terminal 101 sends a data query request to the server 102. After receiving the data query request, the server 102 determines the target search key of the data to be queried, and then queries the target search key in the index of the database, and obtains the data itself or the storage address of the data in the node where the target search key is located. If the storage address is obtained, the data is read according to the storage address, and then the server 102 returns the data to the terminal 101, thereby completing the data query.

[0097] In one possible implementation, the terminal 101 may be a smart phone, a tablet computer, a laptop computer, a desktop computer, a smart speaker, a smart watch, a car terminal, etc., but is not limited thereto. The server 102 may be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms.

[0098] Figure 2 is a flow chart of a data query method provided in an embodiment of the present application. The embodiment of the present application is executed by a computer device, for example, the computer device is the above-mentioned Figure 1 The server 102 in the illustrated implementation environment. Figure 2 , the method comprising:

[0099] 201. A computer device obtains an index and a target search key of data to be queried, wherein the index includes multiple layers, each layer includes nodes, each node includes a search key, each node has a number, the multiple layers include multiple pre-fetch layers, and the multiple pre-fetch layers correspond to pre-fetch models. The pre-fetch model is used to predict the number corresponding to any search key, and the number corresponding to the search key refers to the number of the node in the pre-fetch layer located on the access path when querying the search key.

[0100] The computer device obtains an index, which includes multiple layers, each layer includes at least one node, and each node includes at least one search key. In addition, each node also has its own number. Optionally, the node numbers are numbered in sequence based on the entire index, so the numbers of any two nodes in the entire index are different. Optionally, the node numbers are numbered in sequence based on a layer, so the numbers of nodes in two different layers may be the same or different.

[0101] Optionally, the node stores a search key, which is used to query data. Optionally, in addition to the search key, the internal node stores a pointer to the child node, and the leaf node stores the data indicated by the search key or the storage address of the data indicated by the search key.

[0102] The multiple layers of the index include multiple pre-fetch layers. Optionally, the multiple pre-fetch layers may be part of the multiple layers of the index, that is, part of the layers in the index are pre-fetch layers, and the nodes in these pre-fetch layers need to be pre-fetched, and the other part of the layers in the index are not pre-fetch layers, and the nodes in these non-pre-fetch layers do not need to be pre-fetched. Optionally, the multiple pre-fetch layers may be all of the multiple layers of the index, that is, all of the layers in the index are pre-fetch layers, and the nodes in each layer in the index need to be pre-fetched. Among them, the multiple pre-fetch layers in the index correspond to pre-fetch models, and the pre-fetch models can be used to predict which nodes in the pre-fetch layers need to be pre-fetched.

[0103] In one possible implementation, the index in the embodiment of the present application may be a structure such as a B+ tree or a radix tree that has a pointer chasing problem, which is not limited in the embodiment of the present application.

[0104] After receiving a data query request for the data to be queried, the computer device determines a target search key for the data, which is included in the index, so the target search key needs to be queried in the index to obtain the data or the storage address of the data. Before the computer device queries the index for the target search key, it first performs the following steps 202-203 to pre-fetch the nodes to be accessed.

[0105] 202. The computer device predicts multiple target numbers corresponding to the target search key through a pre-fetch model.

[0106] After determining the target search key of the data to be queried, the computer device predicts multiple target numbers corresponding to the target search key through the pre-fetch model. The multiple target nodes indicated by the multiple target numbers are the nodes in the multiple pre-fetch layers that need to be accessed when searching for the target search key in the index. That is, the subsequent query of the target search key needs to access the target nodes indicated by the multiple target numbers.

[0107] For each prefetch layer, when querying the target search key, it is necessary to access at least one target node in the prefetch layer, and there are multiple prefetch layers, so the number of target numbers that can be obtained by the prefetch model is also multiple.

[0108] 203. The computer device caches multiple target nodes indicated by multiple target numbers from the memory to the memory buffer.

[0109] After determining the multiple target numbers corresponding to the target search key, the computer device determines the multiple target nodes indicated by the multiple target numbers in the memory, and caches the multiple target nodes from the memory to the memory buffer. The nodes in the index are stored in the memory, for example, the memory is an external memory device, and the external memory device includes an SSD, etc. The embodiment of the present application does not limit the type of memory.

[0110] In the process of querying the search key in the index, it is necessary to access the nodes of each layer layer by layer starting from the root node until the leaf node, so it is necessary to cache all the nodes on the access path from the memory to the memory buffer, that is, to read all the nodes on the access path into the memory buffer so as to access these nodes in the memory buffer. In the embodiment of the present application, before querying the search key, multiple target nodes required to be accessed for querying the search key are predicted in advance, and the multiple target nodes are cached in batches from the memory to the memory buffer in advance, thereby realizing pre-fetching of the target nodes required to be accessed.

[0111] 204. The computer device queries the target search key in the index, and when the node currently to be accessed belongs to the target node, accesses the cached target node in the memory buffer.

[0112] After completing the pre-fetching of the target node, the computer device searches the index for the target search key, starting from the root node, and accesses the nodes of each layer layer by layer starting from the root node in the order pointed by the node pointer until the leaf node where the target search key is located is found. Among them, for the node currently to be accessed, if the node belongs to the target node, since the target node has been pre-fetched into the memory buffer, there is no need to perform a read operation on the memory to read the target node, and the cached target node can be directly accessed in the memory buffer.

[0113] In the related art, when querying a certain search key among multiple nodes of an index, every time a node needs to be accessed, the node needs to be first read from the memory into the buffer, and then the node can be accessed in the buffer, resulting in frequent read operations on the memory during the entire query process. In the embodiments of the present application, before querying a certain search key, the nodes that need to be accessed for querying the search key are predicted in advance, and multiple nodes that need to be accessed are read in batches into the buffer in advance. Then, during the subsequent query process, there is no need to perform read operations on the memory for these nodes, thereby improving the query efficiency.

[0114] In the method provided by the embodiments of the present application, a prefetch layer in the index corresponds to a prefetch model. When it is necessary to query a target search key of data to be queried in the index, the target numbers of the nodes in the prefetch layer that need to be accessed when querying the target search key are predicted first through the prefetch model, and then multiple target nodes indicated by the target numbers are cached in batches into the memory buffer in advance. During the process of querying the target search key in the index subsequently, when a target node needs to be accessed, since the target node has been cached in the memory buffer in advance, there is no need to read the node from the memory anymore, and the target node can be directly accessed in the memory buffer. By caching nodes in batches in advance, the present application can avoid the problem of performing a read operation on the memory for each accessed node, which is beneficial to improving the efficiency of data query.

[0115] The above Figure 2 embodiments are only a brief description of the data query method. For the detailed process of the data query method, reference can be made to the following Figure 3 embodiments. Figure 3 FIG. is a flowchart of another data query method provided by the embodiments of the present application. The embodiments of the present application are executed by a computer device. For example, the computer device is the server 102 in the above Figure 1 shown implementation environment. Referring to Figure 3 , the method includes:

[0116] 301. The computer device obtains an index and a target search key of the data to be queried. The index includes multiple layers, each layer includes nodes, each node includes a search key, each node has a number, and multiple prefetch layers are included in the multiple layers. The multiple prefetch layers correspond to a prefetch model, and the prefetch model is used to predict the number corresponding to any search key. The number corresponding to the search key refers to the number of the nodes on the access path in the prefetch layer when querying the search key.

[0117] The index includes the search keys of the stored data, as well as the data itself or the storage address of the data. Therefore, when it is necessary to query a certain piece of data, the data itself or the storage address of the data can be queried in the index through the target search key of the data. The index in the embodiments of the present application can be a B+ tree, a B-tree, or a radix tree, etc.

[0118] Taking this index as an example of a B+ tree, then this index has multiple levels, and each level includes at least one node. For example, the first level includes a root node, and the root node has no parent node. The last level includes multiple leaf nodes, and the leaf nodes have no child nodes. The nodes in each level other than the first and the last levels include both a parent node and a child node. The nodes in the B+ tree can be divided into internal nodes and leaf nodes. An internal node refers to a node that has at least one child node, and a leaf node refers to a node that has no child node. Among them, each node includes at least one search key, that is, multiple search keys can be stored in one node. Among them, for a certain node, the multiple search keys in this node are arranged in ascending order, and for a certain level, the multiple nodes in this level are arranged in ascending order of the search keys.

[0119] Among them, except for the leaf nodes, the search keys in each of the other internal nodes are used to separate the corresponding range of child nodes. And the multiple leaf nodes in the last level are connected in the form of a linked list in sequence. In addition, except for the leaf nodes, the other internal nodes in this index only store the search keys and the corresponding node pointers. Therefore, these nodes only play the role of indexing and do not have the function of data storage. Only the leaf nodes store the search keys and the corresponding data (or the storage address of the data). The leaf nodes play the role of data storage. Therefore, when querying a certain search key in this index, start from the root node and query layer by layer until the search key is found in the leaf nodes of the last level, then the data or the storage address of the data can be obtained to complete a query process.

[0120] Figure 4 It is a schematic diagram of a data query method provided by an embodiment of the present application. Figure 4 The index shown in Figure 4 is a B+ tree. As shown in

[0121] For the sake of easy explanation, taking the index as including three levels as an example, the first level in this index has 1 node, the second level in this index has 2 nodes, and the third level in this index has 5 nodes.

[0122] Among them, the pointer to the left of the search key 15 in the second layer points to the first node in the third layer. The first node includes the search keys 10 and 12, and the search keys in the first node are less than the search key 15. The pointer between the search keys 15 and 35 in the second layer points to the second node in the second layer. The second node includes the search keys 15 and 21, and the search keys in the second node are greater than or equal to the search key 15 and less than the search key 35. The pointer to the right of the search key 35 in the second layer points to the third node in the second layer. The third node includes the search keys 35 and 43, and the search keys in the third node are greater than or equal to the search key 35. The pointer to the left of the search key 63 in the second layer points to the fourth node in the third layer. The fourth node includes the search keys 54 and 59, and the search keys in the fourth node are less than the search key 63. The pointer to the right of the search key 63 in the second layer points to the fifth node in the third layer. The fifth node includes the search keys 63, 72, and 85, and the search keys in the fifth node are greater than or equal to the search key 63.

[0123] As Figure 4 shown, the third layer is the last layer of the index. Therefore, the nodes in the third layer have no child nodes, and multiple nodes in the third layer are connected in the form of a linked list in sequence.

[0124] 302. Each prefetch layer corresponds to a prefetch model. For the target prefetch layer among multiple prefetch layers, the computer device predicts a target number in the target prefetch layer corresponding to the target search key through the prefetch model corresponding to the target prefetch layer. The target number corresponding to the target prefetch layer refers to the number of the node in the target prefetch layer that needs to be accessed when querying the target search key. The target prefetch layer is any one of the multiple prefetch layers.

[0125] In the embodiment of the present application, the index includes multiple prefetch layers, each prefetch layer corresponds to a prefetch model, and the prefetch model in each prefetch layer is used to predict the number of the node in this prefetch layer that needs to be accessed when querying any search key, that is, each prefetch model is responsible for predicting the number of the node that needs to be accessed in its own prefetch layer.

[0126] Taking a certain target prefetch layer as an example, during the process of querying the target search key, a node in the target prefetch layer needs to be accessed. The computer device predicts a target number corresponding to the target search key through the prefetch model corresponding to the target prefetch layer. The target number is the number of the node in the target prefetch layer that needs to be accessed when querying the target search key. For each prefetch layer, the computer device performs the above operations, and then a target number corresponding to each prefetch layer can be obtained, thereby obtaining multiple target numbers.

[0127] In a possible implementation, the computer device predicts a target number corresponding to the target prefetch layer through the prefetch model corresponding to the target prefetch layer, including: inputting the target search key into the prefetch model to obtain a triple output by the prefetch model, where the triple includes a reference number, an upper bound of the number error, and a lower bound of the number error. The computer device determines the sum value of the reference number and the upper bound of the number error as the maximum number, and determines the difference value between the reference number and the lower bound of the number error as the minimum number; and determines a target number among multiple numbers between the minimum number and the maximum number.

[0128] Among them, the reference number is an estimated value predicted by the prefetch model, indicating that the node indicated by the reference number may need to be accessed. However, there is still a certain error in the reference number. Therefore, the prefetch model also outputs the upper bound of the number error and the lower bound of the number error. The upper bound of the number error represents the number range greater than the reference number among the numbers of the nodes that may need to be accessed, and the lower bound of the number error represents the number range less than the reference number among the numbers of the nodes that may need to be accessed. Therefore, the multiple numbers from the minimum number to the maximum number are all the ranges of the numbers of the nodes that may need to be accessed, and the numbers of the nodes that may need to be accessed belong to the interval between the minimum number and the maximum number. Therefore, the computer device determines the target number among multiple numbers between the minimum number and the maximum number, and this target number is the number of the node that needs to be accessed finally determined.

[0129] Optionally, the computer device performs a binary search among multiple numbers between the minimum number and the maximum number to determine the number of the node including the target search key. The specific search method is related to the creation method of the prefetch model. For the detailed process, refer to the following two creation methods and the corresponding search methods.

[0130] The first creation method: The multiple search keys in each node are arranged in ascending order, and the multiple nodes in each layer are arranged in ascending order of the search keys in the nodes. Then the creation process of the prefetch model corresponding to the target prefetch layer includes: numbering the multiple nodes according to the arrangement order of the multiple nodes in the target prefetch layer; creating the prefetch model corresponding to the target prefetch layer based on the minimum search key of each node in the target prefetch layer and the number of the node where it is located. The prefetch model is used to linearly approximate the functional relationship between the minimum search key of multiple nodes in the target prefetch layer and the number of the node where it is located. Among them, when creating the prefetch model corresponding to the target prefetch layer, the minimum search key of the node in the target prefetch layer is used as the input, and the number of the node where it is located is used as the output.

[0131] Taking the creation of the prefetch model corresponding to the target prefetch layer as an example, first number the multiple nodes according to the arrangement order of the multiple nodes in the target prefetch layer. For example, the numbering can start from 0 and increase incrementally, such as Figure 4As shown, the number of the first node in the second layer is 0, the number of the second node is 1, the number of the first node in the third layer is 0, the number of the second node is 1, the number of the third node is 2, the number of the fourth node is 3, and the number of the fifth node is 4.

[0132] Meanwhile, the computer device determines the minimum search key of each node and the number of the node where it is located. Based on the minimum search key and number of the node, a prefetch model is created, which is used to linearly approximate the functional relationship between the minimum search key and number of the node. Optionally, the prefetch model is a piecewise linear approximation model (PLA, Piecewise Linear Approximation), or the prefetch model is a machine learning model obtained by fitting the functional relationship between the minimum search key and number of the node, etc. The embodiments of the present application do not limit this.

[0133] Optionally, after numbering the nodes in the index, the computer device creates a mapping table, which includes the mapping relationship between the number of the node and the storage address of the node, and the mapping table also includes the minimum search key in the node indicated by the number.

[0134] In the first creation method, the method for determining a target number among multiple numbers from the smallest number to the largest number is as follows: among multiple numbers from the smallest number to the largest number, at least one candidate number is determined, and the minimum search key in the node indicated by the candidate number is less than the target search key; among the at least one candidate number, the candidate number with the largest minimum search key in the indicated node is determined as the target number.

[0135] Optionally, the computer device obtains the created mapping table, which includes the minimum search key in the node indicated by the number. Therefore, the computer device can, according to the mapping table, use the binary search method to determine at least one candidate number whose corresponding minimum search key is less than the target search key among consecutive multiple numbers, and then among the at least one candidate number, determine the candidate number with the largest corresponding minimum search key, so as to determine the candidate number as the target number. In the embodiments of the present application, in the case of creating a prefetch model based on the minimum search key of the node and the number of the node where it is located, the subsequent size relationship between the minimum search key in the node and the target search key is used to determine the target number.

[0136] The second creation method: The multiple search keys in each node are arranged in ascending order, and the multiple nodes in each layer are arranged in ascending order of the search keys in the nodes. The creation process of the prefetch model corresponding to the target prefetch layer includes: numbering the multiple nodes according to the arrangement order of the multiple nodes in the target prefetch layer; creating the prefetch model corresponding to the target prefetch layer based on the maximum search key of each node in the target prefetch layer and the number of the node where it is located. The prefetch model is used to linearly approximate the functional relationship between the maximum search key of multiple nodes in the target prefetch layer and the number of the node where it is located. Among them, when creating the prefetch model corresponding to the target prefetch layer, the maximum search key of the node in the target prefetch layer is used as the input, and the number of the node where it is located is used as the output.

[0137] Taking the creation of the prefetch model corresponding to the target prefetch layer as an example, first, number the multiple nodes according to the arrangement order of the multiple nodes in the target prefetch layer. For example, Figure 4 the numbers shown can start from 0 and increase incrementally. At the same time, the computer device determines the maximum search key of each node and the number of the node where it is located, and creates a prefetch model based on the maximum search key and number of the node. The prefetch model is used to linearly approximate the functional relationship between the maximum search key and number of the node. Optionally, the prefetch model is a piecewise linear approximation model or a machine learning model, etc., which is not limited in the embodiments of the present application.

[0138] Optionally, after numbering the nodes in the index, the computer device creates a mapping table, which includes the mapping relationship between the number of the node and the storage address of the node, and the mapping table also includes the maximum search key in the node indicated by the number.

[0139] In the second creation method, the method for determining a target number among the multiple numbers from the smallest number to the largest number is as follows: Among the multiple numbers from the smallest number to the largest number, determine at least one candidate number, and the maximum search key in the node indicated by the candidate number is greater than the target search key; among the at least one candidate number, determine the candidate number with the smallest maximum search key in the indicated node as the target number.

[0140] Optionally, the computer device obtains the created mapping table, which includes the maximum search key in the node indicated by the number. Therefore, the computer device can use the mapping table to determine at least one candidate number whose corresponding maximum search key is greater than the target search key among the consecutive multiple numbers by means of binary search, and then among the at least one candidate number, determine the candidate number with the smallest corresponding maximum search key, so as to determine the candidate number as the target number. In the embodiments of the present application, when creating a prefetch model based on the maximum search key of the node and the number of the node where it is located, the subsequent size relationship between the maximum search key in the node and the target search key is used to determine the target number.

[0141] It should be noted that in the above possible implementation methods, it is described that the mapping table stores the mapping relationship between the node numbers and the storage addresses of the nodes, as well as the boundary search keys of the nodes (i.e., the minimum search key or the maximum search key). Among them, storing the mapping relationship between the node numbers and the storage addresses of the nodes in the mapping table is to obtain the actual storage location of the target node indicated by the target number, so as to perform batch prefetching on multiple target nodes. Storing the boundary search keys of the nodes in the mapping table is to perform binary search within the error range to find the accurate node number, and this error range refers to the range between the above-mentioned lower bound of the number error and the lower bound of the number error, that is, multiple numbers from the minimum number to the maximum number.

[0142] It should be noted that in the embodiments of the present application, an example is given of determining a target number among multiple numbers from the minimum number to the maximum number. In some other embodiments, considering that if the above two methods are used to determine the target number, it is necessary to store the boundary search keys of the nodes in the mapping table for querying the target number, and this mapping table will also occupy additional storage space. In order to reduce the size of the mapping table to save storage space, the boundary search keys of the nodes may not be stored in the mapping table. Then, after determining the minimum number and the maximum number, multiple numbers from the minimum number to the maximum number can be determined as the target numbers, that is, the nodes indicated by these multiple numbers are all prefetched into the memory buffer. This not only saves storage space without storing the boundary search keys, but also improves the fault tolerance rate. In addition, considering that more nodes are prefetched in this way, it may cause some invalid prefetch requests, which may lead to buffer pollution and additional bandwidth consumption. To mitigate these side effects, the error parameter of the prefetch model can be set to reduce the error range output by the prefetch model and reduce the number of nodes to be prefetched.

[0143] 303. The computer device obtains the mapping table and queries the multiple target storage addresses corresponding to the multiple target numbers in the mapping table. The mapping table stores the mapping relationship between the numbers of the nodes in the prefetch layer and the storage addresses of the nodes.

[0144] After the computer device determines the multiple target numbers, it obtains the mapping table, which stores the mapping relationship between the numbers of the nodes in the prefetch layer and the storage addresses of the nodes. Therefore, the computer device can query the multiple target storage addresses corresponding to the multiple target numbers in this mapping table, and this target storage address is the storage address of the target node in the memory.

[0145] In a possible implementation method, as described in the two creation methods in step 302 above, the mapping table also stores the boundary search keys in the nodes indicated by the numbers.

[0146] In a possible implementation, each prefetch layer corresponds to its own mapping table. Taking the target prefetch layer as an example, the mapping table of the target prefetch layer stores the mapping relationship between the numbers of the nodes in the target prefetch layer and the storage addresses of the nodes. For the target number corresponding to the target prefetch layer, the computer device obtains the mapping table of the target prefetch layer and queries the target storage address corresponding to the target number in the mapping table of the target prefetch layer.

[0147] 304. The computer device reads the target nodes at multiple target storage addresses in the memory and caches the read multiple target nodes into the memory buffer.

[0148] After the computer device determines the target storage address of the target node, it reads multiple target nodes at the target storage address in the memory, so as to cache the target nodes into the memory buffer, so that these target nodes can be directly accessed in the memory buffer when querying the target search key in the index subsequently.

[0149] In a traditional index, due to the problem of pointer chasing, only when accessing the nodes of a certain layer, by performing a binary search on the search key in the node, can the position of the nodes to be accessed in the next layer be obtained. Therefore, for a single query request, each time a node is accessed. If the node is not in the memory buffer, an I / O (Input / Output) request needs to be synchronously sent to the memory and waited for completion, then the node is accessed, and then the next node to be accessed can be determined. Therefore, multiple nodes on the access path of a single query request cannot be read in batches by means of asynchronous I / O. In the embodiment of the present application, since multiple target nodes to be accessed are predicted through the prefetch model, multiple I / O requests can be sent to the memory by means of asynchronous I / O to read multiple target nodes in batches, thereby improving the utilization rate of the bandwidth of the memory (such as the external storage device SSD) and reducing the I / O access latency.

[0150] In a possible implementation, the computer device uses the io_uring asynchronous I / O framework to implement the asynchronous I / O access process. Its basic components include a submission queue for storing the data of the operations to be executed and a completion queue for storing the return results of the completed operations. The data of the operations to be executed in the submission queue are submission queue entries, and the return results of the completed operations in the completion queue are completion queue entries.

[0151] Among them, the process of the computer device using io_uring to implement asynchronous I / O access is as follows: The computer device initializes the submission queue, writes the I / O operations to be performed into the submission queue. One I / O operation corresponds to a target node to be prefetched. After writing the I / O operations corresponding to multiple target nodes into the submission queue, it submits multiple I / O requests in this queue to the kernel and waits for the kernel to finish processing.

[0152] Optionally, the io_uring framework also supports advanced features. For example, when initializing the queue, an additional kernel thread is started. This kernel thread is responsible for polling the queue, that is, periodically querying whether all the I / O operations in the queue have been completed. Therefore, there is no need to use the system call method to determine whether all the I / O operations have been completed, which can further reduce the I / O latency.

[0153] It should be noted that the above embodiments only take the process of using the io_uring asynchronous I / O framework to implement asynchronous I / O access as an example for illustration. In addition, the computer device can also use other asynchronous I / O methods to implement the asynchronous I / O access process, and the embodiments of this application do not limit this.

[0154] 305. The computer device queries the target search key in the index. When the currently required node to be accessed belongs to the target node, it accesses the cached target node in the memory buffer.

[0155] After the prefetch of the target node is completed, the computer device queries the target search key in the index, starts from the root node, and accesses the nodes in the order pointed to by the node pointers in turn until the leaf node where the target search key is located is found. Among them, for the currently required node to be accessed, if this node belongs to the target node, since this target node has been prefetched into the memory buffer, there is no need to perform a read operation on the memory to read this target node, and it can directly access the cached target node in this memory buffer.

[0156] In a possible implementation, when the currently required node to be accessed does not belong to the target node, if this node is not in the memory buffer, the storage address of this node is determined, and the node at this storage address in the memory is cached to the memory buffer, and then this node is accessed in the memory buffer.

[0157] Among them, in steps 302 - 304 above, only the target nodes required to be accessed in the prefetch layer are prefetched into the memory buffer. Therefore, the nodes required to be accessed in the non-prefetch layer are not prefetched into the memory buffer. In this case, this node may not be in the memory buffer, and it is necessary to read this node from the memory.

[0158] In addition, since the target nodes prefetched in the above steps 302-304 are the nodes that need to be accessed predicted by the prefetching model, and there are certain errors in the prefetching model (for example, after updating the nodes in the prefetching layer, the prefetching model has not been updated in time), it may cause that the nodes that need to be accessed are not predicted by the prefetching model. Therefore, there may be nodes that need to be accessed but not prefetched into the memory buffer. In this case, it is necessary to read the nodes from the memory.

[0159] 306. After the computer device queries the target search key in the indexed nodes, based on the node where the target search key is located, it obtains the data indicated by the target search key.

[0160] In a possible implementation, the indexed nodes store the search key and the data indicated by the search key. Therefore, after the computer device determines the node where the target search key is located, it can obtain the data in this node. In a possible implementation, the indexed nodes store the search key and the storage address of the data indicated by the search key. Therefore, after the computer device determines the node where the target search key is located, it obtains the storage address of the data in this node, and then reads the data at this storage address.

[0161] In a possible implementation, if the index is a B+ tree, after the computer device queries the target search key in the leaf nodes of the index, based on the leaf node where the target search key is located, it obtains the data indicated by the target search key.

[0162] Figure 4 is a schematic diagram of a data query method provided by an embodiment of the present application. As Figure 4 shown, taking the index as a B+ tree and querying the target search key 21 in this index as an example. The computer device predicts the numbers of the nodes on the access path when querying the target search key 21 through the prefetching model, and obtains the number 0 of the first layer, the number 0 of the second layer, and the number 1 of the third layer. The nodes indicated by the above three numbers are cached from the memory to the memory buffer. Then the query process starts: First, it is necessary to access the node

[54] indicated by the number 0 of the first layer. Since this node

[54] has been pre-cached in the memory buffer, it can directly access the node

[54] in the memory buffer. The target search key 21 is less than 54, so the nodes that need to be accessed in the next layer are the nodes [15, 35] indicated by the number 0 of the second layer. Since this node [15, 35] has been pre-cached in the memory buffer, it can directly access the node [15, 35] in the memory buffer. The target search key 21 is greater than 15 and less than 35, so the nodes that need to be accessed in the next layer are the nodes [15, 21] indicated by the number 1 of the third layer. Since this node [15, 21] has been pre-cached in the memory buffer, it can directly access the node [15, 21] in the memory buffer. By sequentially traversing the nodes [15, 21], the target search key 21 is found.

[0163] In the method provided by the embodiment of the present application, a prefetch layer in the index corresponds to a prefetch model. When it is necessary to query the target search key of the data to be queried in the index, first, the prefetch model is used to predict the target numbers of the nodes in the prefetch layer that need to be accessed when querying the target search key. Then, a plurality of target nodes indicated by the target numbers are pre-buffered in the memory buffer in batches. Subsequently, during the process of querying the target search key in the index, when it is necessary to access a target node, since the target node has been pre-buffered in the memory buffer in advance, there is no need to read the node from the memory again, and the target node can be directly accessed in the memory buffer. By pre-buffering nodes in batches in advance, the present application can avoid the problem that each time a node is accessed, a read operation needs to be performed on the memory, which is beneficial to improving the efficiency of data query.

[0164] Moreover, by introducing the prefetch model as an auxiliary tool to predict the nodes to be accessed, there is no need to change the original structure and access process of the index, and the invasiveness to the original index is small, which improves the applicability.

[0165] Moreover, after predicting the multiple nodes to be accessed, the multiple nodes to be accessed are prefetched in batches through the asynchronous I / O method, which makes greater use of the bandwidth of the memory and reduces the I / O access latency, thereby improving the overall access performance of the index and alleviating the problem of pointer chasing.

[0166] It should be noted that, in order to further save storage space, the computer device can adopt the following three methods to reduce the space overhead.

[0167] (1) Selective prefetch: The first case is to selectively prefetch specific nodes. For example, for a certain prefetch layer, only the nodes on the hot access path in the prefetch layer are prefetched, that is, only the nodes that are accessed more frequently are prefetched, which can reduce the size of the mapping table. The second case is to selectively prefetch the nodes in a specific layer. For example, considering that there are more nodes in the last layer of the index, the last layer is not set as the prefetch layer, that is, the nodes in the last layer are not prefetched, which can reduce the size of the prefetch model and the mapping table.

[0168] (2) Reduce the size of the prefetch model: By setting the prefetch model, the size of the prefetch model can be reduced at the cost of increasing the error range of the node numbers output by the prefetch model, so as to exchange a certain amount of computational overhead for a smaller space overhead.

[0169] (3) Reduce the size of the mapping table: As described in step 302 above, in the case where a plurality of numbers between the smallest number and the largest number are all determined as the target numbers, the boundary search keys of the nodes indicated by the numbers do not need to be stored in the mapping table, thereby reducing the size of the mapping table.

[0170] In some embodiments, the leaf nodes in the last layer of the index include all search keys. The multiple prefetch layers include the last layer of the index. The multiple nodes in each layer are arranged in ascending order of the search keys in the nodes. For example, the index is a B+ tree. The target search key of the data to be queried is a search key between the lower bound search key and the upper bound search key of the target search key range, that is, it is necessary to query the data corresponding to the search keys between the lower bound search key and the upper bound search key, rather than performing a point query on the data corresponding to a certain search key, but performing a range query on the data corresponding to a continuous multiple search keys. Then, the detailed process of data query can be seen in the following Figure 5 embodiments.

[0171] Figure 5 is a flowchart of another data query method provided by an embodiment of the present application. The embodiment of the present application is executed by a computer device. See Figure 5 and the method includes:

[0172] 501. The computer device obtains an index and the target search key of the data to be queried. The index includes multiple layers, each layer includes nodes, each node includes a search key, each node has a number, the multiple layers include multiple prefetch layers, and the multiple prefetch layers correspond to a prefetch model. The prefetch model is used to predict the number corresponding to any search key. The number corresponding to the search key refers to the number of the node on the access path in the prefetch layer when querying the search key.

[0173] This step 501 is the same as step 301 above, except that the target search key in this step 501 is a search key between the lower bound search key and the upper bound search key of the target search key range. The target search key range refers to the search key range to be queried. The lower bound search key is the smallest search key in the target search key range, and the upper bound search key is the largest search key in the target search key range. Therefore, the lower bound search key is less than the upper bound search key.

[0174] 502. The computer device predicts, through the prefetch model, multiple target numbers corresponding to the lower bound search key and the upper bound leaf node number corresponding to the upper bound search key. The multiple target numbers include the lower bound leaf node number, and the lower bound leaf node number refers to the number of the leaf node where the lower bound search key is located in the last layer of the index. The upper bound leaf node number refers to the number of the leaf node where the upper bound search key is located in the last layer of the index.

[0175] In the embodiment of the present application, the multiple prefetch layers include the last layer of the index. The last layer of the index includes leaf nodes, and the leaf nodes refer to the nodes without child nodes. Therefore, among the multiple target numbers predicted by the computer device through the prefetch model, the number of the leaf nodes in this last layer is included.

[0176] Among them, the process in step 502 where the computer device predicts multiple target numbers corresponding to the lower bound search key through the prefetch model is the same as the process in step 302 where the computer device predicts multiple target numbers corresponding to the target search key, and will not be elaborated here one by one.

[0177] Among them, the process in step 502 where the computer device predicts the upper bound leaf node number corresponding to the upper bound search key through the prefetch model is the same as the process in step 302 where the computer device predicts multiple target numbers corresponding to the target search key. The difference is that only the number of the leaf node on the last layer of the access path when querying the upper bound search key needs to be predicted. For example, the prefetch layer includes the last layer in the index, and each prefetch layer corresponds to a prefetch model. The computer device only needs to predict, through the prefetch model corresponding to the last layer, the number of a leaf node corresponding to the upper bound search key in the last layer, so as to obtain the upper bound leaf node number.

[0178] 503. The computer device caches multiple target nodes indicated by multiple target numbers, and multiple leaf nodes indicated by multiple leaf node numbers from the lower bound leaf node number to the upper bound leaf node number from the memory to the memory buffer.

[0179] The multiple target nodes are the nodes required to access the lower bound search key in the index, and the multiple leaf nodes are the nodes required to access the range between the lower bound search key and the upper bound search key after querying the lower bound search key in the index. Therefore, the computer device caches the multiple target nodes and the multiple leaf nodes from the memory to the memory buffer.

[0180] Among them, the process of caching multiple target nodes and the multiple leaf nodes from the memory to the memory buffer in step 503 is the same as the process of caching multiple target nodes to the memory buffer in steps 303 - 304 above, and will not be elaborated here one by one.

[0181] 504. The computer device queries the lower bound search key in the index. When the currently required accessed node belongs to the target node, it accesses the cached target node in the memory buffer.

[0182] The process of step 504 is the same as the process of step 305 above, and will not be elaborated here one by one.

[0183] 505. After the computer device queries the lower bound search key in the nodes of the index, it queries the search key between the lower bound search key and the upper bound search key in the nodes after the node where the lower bound search key is located. When the currently required accessed node belongs to the cached leaf node, it accesses the cached leaf node in the memory buffer.

[0184] The process of step 505 is the same as that of step 305 described above, and will not be elaborated here one by one.

[0185] Figure 6 It is a schematic diagram of another data query method provided by an embodiment of the present application. As Figure 6 shown, taking the index as a B+ tree and taking the query of search keys between search key 12 and search key 72 in this index as an example, search key 12 is the lower bound search key and search key 72 is the upper bound search key. The computer device predicts the numbers of the nodes on the access path when querying this search key 12 through a prefetch model, and obtains the number 0 of the first layer, the number 0 of the second layer, and the number 0 of the third layer. The number 0 of the third layer is the lower bound leaf node number. The computer device predicts the numbers of the nodes on the access path in the last layer when querying this search key 72 through a prefetch model, and obtains the number 4 of the third layer. The number 4 of the third layer is the upper bound leaf node number. The computer device caches the nodes indicated by the number 0 of the first layer, the number 0 of the second layer, and the numbers 0-4 of the third layer from the memory to the memory buffer.

[0186] Then start the query process: First, it is necessary to access the node

[54] indicated by the number 0 of the first layer. Since this node

[54] has been pre-cached in the memory buffer, it can be directly accessed in the memory buffer. The search key 12 is less than 54, so the node to be accessed in the next layer is the node [15, 35] indicated by the number 0 of the second layer. Since this node [15, 35] has been pre-cached in the memory buffer, it can be directly accessed in the memory buffer. The search key 12 is less than 15, so the node to be accessed in the next layer is the node [10, 12] indicated by the number 0 of the third layer. Since this node [10, 12] has been pre-cached in the memory buffer, it can be directly accessed in the memory buffer. By sequentially traversing the node [10, 12], the search key 12 is found, and then the nodes indicated by the numbers 1-4 of the third layer are sequentially accessed in the memory buffer until the search key 72 is found, thus completing the search for search keys between search key 12 and search key 72.

[0187] In the method provided by the embodiment of the present application, since the search keys in the index are arranged in sequence and the index structure is explicitly organized according to the order of the data, and the prefetch model for predicting the nodes to be accessed is layer-based and corresponds to each prefetch layer one by one, the embodiment of the present application can realize prefetching of the nodes to be accessed in the case of range query, further improving the versatility.

[0188] In practical applications, a memory has a maximum request quantity, which can refer to the maximum number of access requests that the memory can process simultaneously. Among them, one access request needs to be initiated for each accessed node. Then, when the number of multiple nodes to be prefetched is greater than the maximum request quantity, the multiple nodes need to be prefetched in batches.

[0189] For example, when the memory is an SSD, the maximum request quantity is called the queue depth of the SSD. The queue depth of the SSD is usually 16 or more. Taking a B+ tree with an index as an example, the number of layers of the B+ tree on the SSD is usually 4 layers. Therefore, for point queries, all nodes on the access path can be batch-cached to the memory buffer at one time using the method provided in the embodiments of the present application. However, for range queries, if the query range length is long, the number of multiple nodes to be prefetched is large. If all nodes are prefetched at one time, it will not only significantly increase the I / O latency but also contaminate the memory buffer. Therefore, the multiple nodes need to be prefetched in batches. Among them, for the process of prefetching multiple nodes in batches, reference can be made to the following Figure 7 embodiments.

[0190] Figure 7 is a flowchart of another data query method provided by the embodiments of the present application. The embodiments of the present application are executed by a computer device. Refer to Figure 7 which includes:

[0191] 701. The computer device obtains an index and a target search key of the data to be queried. The index includes multiple layers, each layer includes nodes, each node includes a search key, each node has a number, and multiple prefetch layers are included in the multiple layers. The prefetch layers correspond to a prefetch model, and the prefetch model is used to predict the number corresponding to any search key. The number corresponding to the search key refers to the number of nodes on the access path in the prefetch layer when querying the search key.

[0192] The process of step 701 is the same as that of step 301 above, and will not be elaborated here one by one.

[0193] 702. The computer device predicts multiple target numbers corresponding to the target search key through the prefetch model.

[0194] The process of step 702 is the same as that of step 302 above, and will not be elaborated here one by one.

[0195] 703. When the number of multiple target nodes is greater than the maximum request quantity, the computer device determines the first maximum request quantity of target nodes in the arranged order of the multiple target nodes, and caches the determined target nodes from the memory to the memory buffer.

[0196] The memory has a maximum number of requests, which refers to the maximum number of access requests that the memory can process simultaneously. One access request needs to be initiated for each node accessed, and multiple target nodes are arranged in sequence according to the access order. The multiple target nodes refer to the multiple target nodes corresponding to multiple target numbers. If the number of multiple target nodes is greater than the maximum number of requests, then according to the arrangement order of the multiple target nodes, the first maximum number of requests target nodes are determined, and the first maximum number of requests target nodes are first cached from the memory to the memory buffer, and the remaining target nodes are not processed temporarily.

[0197] For example, if the maximum number of requests is 16 and the number of the multiple target nodes is 40, then the first 16 target nodes are first cached from the memory to the memory buffer, and the 17th to 40th target nodes are not processed temporarily.

[0198] Among them, the process of caching the target node from the memory to the memory buffer is the same as the process of step 304 above, and will not be elaborated here one by one.

[0199] 704. The computer device queries the target search key in the index. When the currently required accessed node belongs to the target node, the cached target node is accessed in the memory buffer.

[0200] After the computer device completes a prefetch process, it can start to query the target search key in the index. The process of this step 704 is the same as the process of step 305 above, and will not be elaborated here one by one.

[0201] 705. When all the maximum number of requests target nodes cached in the memory buffer have been accessed, the computer device determines the first no more than the maximum number of requests target nodes that have not been cached, and caches the determined target nodes from the memory to the memory buffer until all the multiple target nodes have been cached in the memory buffer.

[0202] When all the maximum number of requests target nodes cached in the memory buffer have been accessed, that is, when all the target nodes of the first prefetch have been accessed, the query of the target search key is first paused, and the next prefetch process is started. The computer device determines the first no more than the maximum number of requests target nodes that have not been cached among the multiple target nodes, caches the determined target nodes from the memory to the memory buffer, and then continues to query the target search key from the current progress. When all the target nodes of the second prefetch have also been accessed, the query of the target search key is paused again, and the next prefetch process is continued until all the multiple target nodes have been cached in the memory buffer.

[0203] For example, the maximum number of requests is 16, the number of the multiple target nodes is 40. In the first prefetch, the 1st to 16th target nodes are prefetched; in the second prefetch, the 17th to 32nd target nodes are prefetched; and in the third prefetch, the 33rd to 40th target nodes are prefetched.

[0204] 706. After the computer device finds the target search key in the indexed nodes, based on the node where the target search key is located, it obtains the data indicated by the target search key.

[0205] The process of step 706 is the same as that of step 306 above, and will not be elaborated here one by one.

[0206] In the method provided by the embodiments of the present application, if the number of the multiple target nodes to be prefetched is large, the multiple target nodes are prefetched in batches, which can avoid increasing the I / O latency, reduce the pollution to the memory buffer, and thus improve the query efficiency.

[0207] In some embodiments, the prefetch model is created based on the minimum search key of each node in the target prefetch layer and the number of the node where it is located. The multiple prefetch layers also correspond to a mapping table, which stores the mapping relationship between the number of the nodes in the prefetch layer and the storage addresses of the nodes. Since the nodes in the prefetch layer can be modified, after the nodes in the prefetch layer are modified, in order to ensure the accuracy of the predicted node numbers and the accuracy of the storage addresses obtained according to the mapping table, it is necessary to update the prefetch model and the mapping table.

[0208] Among them, the index also corresponds to a log, which is used to store the modification records of the index. Subsequently, the prefetch model and the mapping table can be updated based on the modification records stored in the log, and the used log is deleted in time after the update. The prefetch model, the mapping table and the log constitute the basic components of the prefetcher. The prefetcher is responsible for the prefetch process in the embodiments of the present application. The update process of the prefetch model, the mapping table and the log can be understood as the update process of the entire prefetcher. For the detailed update process, please refer to the following Figure 8 embodiment.

[0209] Figure 8 is a flowchart of a method for updating a prefetcher provided by the embodiments of the present application. The embodiments of the present application are executed by a computer device. Please refer to Figure 8 The method includes:

[0210] 801. When the computer device responds to the modification operation on the nodes in the prefetch layer, it adds the modification record corresponding to the modification operation to the log.

[0211] In a possible implementation manner, the modification operation includes adding a search key to a node, deleting a search key from a node, node splitting, node merging, node deletion, and node addition, etc.

[0212] In a possible implementation, without touching a node, the search key in the node cannot be modified. Touching a node means reading the search key in the node or structurally modifying the node by splitting, merging, deleting, etc. Therefore, capture the modification operations on the node when touching the node, and add the corresponding modification records to the log.

[0213] In a possible implementation, the modification record includes the operation type of the modification operation, the storage address of the node where the change occurs, and the new search key in the node. Optionally, each prefetch layer corresponds to its own log, and the logs in each prefetch layer are globally shared.

[0214] In a possible implementation, the addition operation of adding the modification record to the log is an atomic operation of Fetch-and-Add (fetch and increment).

[0215] 802. When the computer device currently meets the update condition, based on the modification records in the log, determine the search key and number of the nodes in the updated prefetch layer. According to the search key and number of the nodes in the updated prefetch layer, update the prefetch model, and based on the modification records in the log, update the mapping relationship between the number and storage address of the nodes in the prefetch layer stored in the mapping table.

[0216] It should be noted that in the embodiments of the present application, after modifying the nodes in the prefetch layer, the prefetchers are not immediately updated. Instead, the modification records are first added to the log, and when the update condition is met, the prefetchers are updated.

[0217] The prefetchers include a prefetch model and a mapping table. The update of the prefetchers includes the update of the prefetch model and the update of the mapping table. When updating the prefetch model, since the prefetch model is created according to the search key and number of the nodes in the prefetch layer, it is necessary to first determine the search key and number of the nodes in the updated prefetch layer based on the modification records in the log, and then update the prefetch model according to the search key and number of the nodes in the updated prefetch layer. The process of updating the prefetch model is the same as the first creation method and the second creation method in step 302 above, and will not be elaborated here one by one. When updating the mapping table, since the mapping table includes the mapping relationship between the number and storage address of the nodes, it is only necessary to directly update the mapping relationship between the number and storage address of the nodes in the prefetch layer stored in the mapping table based on the modification records in the log.

[0218] In a possible implementation, the update condition includes any one of the following:

[0219] (1) The data volume of the modification records stored in the log reaches a preset threshold. Among them, whenever the prefetch model and the mapping table are updated according to the modification records in the log, the modification records in the log are deleted, that is, the log is cleared, and then the modification records are accumulated again. When the data volume of the accumulated modification records reaches the preset threshold, it indicates that enough modifications have been made to the index, and the performance of the current prefetching device has significantly declined. Therefore, it is necessary to update the prefetching device.

[0220] (2) The cache hit rate of the memory buffer is less than the preset hit rate. The cache hit rate refers to the probability that the node to be accessed has been cached in the memory buffer. If the cache hit rate of the memory buffer is less than the preset hit rate, it indicates that the accuracy of the nodes prefetched by the prefetching device is relatively low, indirectly indicating that enough modifications have been made to the index, and the performance of the current prefetching device has significantly declined. Therefore, it is necessary to update the prefetching device.

[0221] In a possible implementation, to avoid blocking the working threads responsible for processing query requests, two prefetching device copies are maintained in the computer device, including an active prefetching device copy and an inactive prefetching device copy, that is, one prefetching device copy is in an active state, and the other prefetching device copy is in an inactive state. The working threads predict the access nodes through the active prefetching device copy.

[0222] The prefetching device can be in a shared mode or a private mode. The shared mode means that the prefetching device is shared among multiple working threads, and the private mode means that the prefetching device is local and private to each working thread. Among them, the private mode can reduce the cache miss overhead when each working thread accesses the prefetching device after the prefetching device is updated, because the update of the prefetching device will cause the local caches of all working threads to become invalid due to obsolescence, and they can only gradually cache it locally by accessing the updated prefetching device. For the shared mode, when updating the two prefetching device copies, first merge the logs in the active prefetching device copy into the inactive prefetching device copy, update the mapping table in the inactive prefetching device copy and reconstruct the prefetch model, then switch the inactive prefetching device copy to the active state, switch the other unupdated active prefetching device copy to the inactive state, and finally update the other unupdated prefetching device copy, so as to realize the update of the two prefetching device copies. For the private mode, the working threads are not responsible for reconstructing the prefetching device, but a dedicated reconstruction thread is used to perform the above reconstruction operation to build the latest prefetching device copy. In this mode, it is only necessary to ensure that one of the two prefetching device copies of each working thread is the latest. That is, only need to copy the latest prefetching device copy built by the reconstruction thread to the inactive prefetching device copy of each working thread, then switch the inactive prefetching device copy to the active state, and switch the old version of the prefetching device copy to the inactive state, so as to ensure that all working threads have a latest prefetching device copy.

[0223] Optionally, for both the shared mode and the private mode, the update process of the prefetcher is executed by the reconstruction thread. In a split-memory scenario, the reconstruction thread may reside on a memory node with relatively weak computing power. In this case, each computing node is responsible for executing the update process and completing the update process through one-sided RDMA operations. Optionally, the process of switching the states of the two prefetcher copies is also executed by the reconstruction thread. Among them, the pointer of each worker thread points to the active prefetcher copy. Therefore, the switching process only needs to atomically switch the pointer of each worker thread from one prefetcher copy to the other.

[0224] Optionally, the worker thread can also periodically check whether the version of the inactive prefetcher copy is newer than the version of the active prefetcher copy. If so, a switch is required. Optionally, the check period can be determined by an operation counter, for example, determined by the number of index modifications recorded by the operation counter.

[0225] 803. The computer device deletes the modification records in the log.

[0226] After updating the prefetch model and the mapping table according to the modification records in the log, the modification records in the log are deleted, that is, the log is cleared, and then the modification records are accumulated again.

[0227] In the method provided by the embodiment of the present application, after modifying the nodes in the prefetch layer, the prefetcher is not immediately updated, but the modification records are first added to the log, and the prefetcher is updated only when the update condition is met. That is, the prefetcher is updated through an asynchronous update strategy based on the log, so as to avoid the additional overhead that may be generated by synchronous updates, enable the prefetcher to better handle dynamic workloads, and ensure the accuracy of prefetching and the performance improvement effect.

[0228] It should be noted that the embodiment of the present application takes the prefetcher including a prefetch model, a mapping table, and a log as an example to illustrate the update process of the prefetcher. In some other embodiments, the prefetcher does not include a mapping table, and the computer device uses other methods to determine the storage location of the target node indicated by the target number in the memory. Then, the update process of the prefetcher does not include the update process of the mapping table, that is, there is no need to execute the process of "updating the mapping relationship between the numbers and storage addresses of the nodes in the prefetch layer stored in the mapping table based on the modification records in the log" in step 802 above.

[0229] Fig. 9 It is a comparison diagram of a data query method provided by an embodiment of the present application, as Fig. 9As shown, take the query path of node r, node x, and node y as an example. In the related art, given a search key, the query starts from node r. First, node r needs to be cached from the memory to the memory buffer, and a binary search is performed to obtain the storage address of node x. Then, node x is cached from the memory to the memory buffer, and a binary search is performed to obtain the storage address of node y. Finally, node y is cached from the memory to the memory buffer, and a binary search is performed to determine whether the search key exists. In the embodiment of the present application, given a search key, the storage addresses of node r, node x, and node y to be accessed are predicted through a prefetching model, and then in an asynchronous I / O manner, node r, node x, and node y are batch-cached from the memory to the memory buffer. Since node r, node x, and node y on the access path have been prefetched to the buffer in advance, when querying this search key, only a binary search needs to be performed on node r, node x, and node y in the memory buffer in sequence.

[0230] Fig.10 is a schematic diagram of another data query method provided by the embodiment of the present application. As Fig.10 shown, taking the index including three layers as an example, the prefetching process is implemented by a prefetching device. The prefetching device includes a prefetching model, a mapping table, and a log. The prefetching model is used to predict the numbers of the nodes to be accessed. The mapping table is used to store the mapping relationship between the numbers of the nodes and the storage addresses of the nodes. The log is used to store modification records to update the prefetching device.

[0231] For a certain search key to be queried, the prefetching model in the prefetching device of each layer is responsible for calculating the numbers of the nodes to be accessed in this layer. Through the mapping relationship between the numbers of the nodes in the mapping table and the storage addresses of the nodes, the positions of the nodes to be accessed in the memory can be determined. Therefore, through the prefetching devices of multiple layers, the numbers of all the nodes to be accessed from the layer where the root node is located to the layer where the leaf node is located can be predicted, and then the nodes to be accessed are prefetched from the memory to the memory buffer in an asynchronous I / O manner, so as to ensure as much as possible that all the nodes to be accessed are located in the memory buffer when querying on the B+ tree.

[0232] In addition, when processing modification operations such as inserting nodes and deleting nodes, the modification records are added to the log, and the prefetching model can be asynchronously updated by using the modification records in the log subsequently.

[0233] An embodiment of the present application proposes a lightweight prefetcher for an ordered index of a database. The prefetcher has a clear hierarchy and logic, and does not require modifying the basic structure and access process of the original index. After encapsulating the prefetcher, it only needs to insert relevant interfaces at appropriate positions to be implemented. It can be extended to different index structures and different storage levels, and is convenient to be orthogonal to a variety of other optimization strategies, further exerting the advantages of prefetching, with small code intrusion and strong applicability.

[0234] Moreover, with the assistance of the prefetcher, the pointer chasing problem is better alleviated. By prefetching the nodes on the access path in advance, the speed of data query is accelerated, and the tail latency in the data query process under different workloads can be reduced with a small space overhead.

[0235] Moreover, using the asynchronous I / O method to batch prefetch multiple nodes can better exert the parallel access advantage of the memory and improve the utilization rate of the bandwidth.

[0236] Moreover, the asynchronous update strategy proposed for the prefetcher enables it to better adapt to dynamic workloads. Through operations such as delayed update and prefetcher switching, it can ensure that the accuracy rate of prefetching remains stable, thereby ensuring the performance improvement benefits brought by prefetching.

[0237] Fig.11 It is a schematic structural diagram of a data query device provided by an embodiment of the present application. Refer to Fig.11 This device includes:

[0238] An acquisition module 1101, configured to acquire an index and a target search key of data to be queried. The index includes multiple layers, each layer includes nodes, each node includes a search key, each node has a number, and multiple prefetch layers are included in the multiple layers. The prefetch layers correspond to prefetch models, and the prefetch models are used to predict the numbers corresponding to any search key. The number corresponding to the search key refers to the number of the node on the access path in the prefetch layer when querying the search key;

[0239] A prediction module 1102, configured to predict multiple target numbers corresponding to the target search key through the prefetch model;

[0240] A cache module 1103, configured to cache multiple target nodes indicated by the multiple target numbers from the memory to the memory buffer;

[0241] A query module 1104, configured to query the target search key in the index, and when the currently required accessed node belongs to the target node, access the cached target node in the memory buffer.

[0242] The data query device provided by the embodiment of the present application has a prefetch layer in the index corresponding to a prefetch model. When it is necessary to query the target search key of the data to be queried in the index, first, the prefetch model is used to predict the target numbers of the nodes on the access path in the prefetch layer when querying the target search key. Then, a plurality of target nodes indicated by the target numbers are batch-cached into the memory buffer in advance. Subsequently, during the process of querying the target search key in the index, when it is necessary to access a target node, since the target node has been cached into the memory buffer in advance, there is no need to read the node from the memory again, and the target node can be directly accessed in the memory buffer. By batch-caching nodes in advance, the present application can avoid the problem of reading the memory once for each accessed node, which is beneficial to improving the efficiency of data query.

[0243] Optionally, the cache module 1103 is configured to:

[0244] Obtain a mapping table, which stores the mapping relationship between the numbers of the nodes in the prefetch layer and the storage addresses of the nodes;

[0245] Query the multiple target storage addresses corresponding to the multiple target numbers in the mapping table;

[0246] Read the target nodes on the multiple target storage addresses from the memory, and cache the read multiple target nodes into the memory buffer.

[0247] Optionally, each prefetch layer corresponds to a prefetch model, and the prediction module 1102 is configured to:

[0248] For the target prefetch layer in the multiple prefetch layers, through the prefetch model corresponding to the target prefetch layer, predict a target number corresponding to the target search key in the target prefetch layer. The target number corresponding to the target prefetch layer refers to the number of the node on the access path in the target prefetch layer when querying the target search key, and the target prefetch layer is any one of the multiple prefetch layers.

[0249] Optionally, the prediction module 1102 is configured to:

[0250] Input the target search key into the prefetch model to obtain a triple output by the prefetch model. The triple includes a reference number, an upper bound of the number error, and a lower bound of the number error;

[0251] Determine the sum value of the reference number and the upper bound of the number error as the maximum number, and determine the difference value between the reference number and the lower bound of the number error as the minimum number;

[0252] Determine a target number among the multiple numbers from the minimum number to the maximum number.

[0253] Optionally, see Fig.12, the multiple search keys in each node are arranged in ascending order, and the multiple nodes in each layer are arranged in ascending order of the search keys in the nodes; the apparatus further includes a creation module 1105, configured to:

[0254] Number the multiple nodes according to the arrangement order of the multiple nodes in the target prefetch layer;

[0255] Create a prefetch model corresponding to the target prefetch layer based on the minimum search key of each node in the target prefetch layer and the number of the node where it is located, and the prefetch model is used to linearly approximate the functional relationship between the minimum search key of the multiple nodes in the target prefetch layer and the number of the node where it is located.

[0256] Optionally, refer to Fig.12 , a prediction module 1102, configured to:

[0257] Determine at least one candidate number among the multiple numbers from the smallest number to the largest number, and the minimum search key in the node indicated by the candidate number is less than the target search key;

[0258] Among the at least one candidate number, determine the candidate number with the largest minimum search key in the indicated node as the target number.

[0259] Optionally, refer to Fig.12 , the multiple search keys in each node are arranged in ascending order, and the multiple nodes in each layer are arranged in ascending order of the search keys in the nodes; the apparatus further includes a creation module 1105, configured to:

[0260] Number the multiple nodes according to the arrangement order of the multiple nodes in the target prefetch layer;

[0261] Create a prefetch model corresponding to the target prefetch layer based on the maximum search key of each node in the target prefetch layer and the number of the node where it is located, and the prefetch model is used to linearly approximate the functional relationship between the maximum search key of the multiple nodes in the target prefetch layer and the number of the node where it is located.

[0262] Optionally, refer to Fig.12 , a prediction module 1102, configured to:

[0263] Determine at least one candidate number among the multiple numbers from the smallest number to the largest number, and the maximum search key in the node indicated by the candidate number is greater than the target search key;

[0264] Among the at least one candidate number, determine the candidate number with the smallest maximum search key in the indicated node as the target number.

[0265] Optionally, the memory has a maximum number of requests, which refers to the maximum number of access requests that the memory can process simultaneously, and multiple target nodes are arranged in sequence according to the access order; a cache module 1103, configured to:

[0266] In the case where the number of multiple target nodes is greater than the maximum number of requests, determine the first maximum number of target nodes in the arranged order of the multiple target nodes, and cache the determined target nodes from the memory to the memory buffer;

[0267] The cache module 1103 is further configured to:

[0268] After all the maximum number of target nodes cached in the memory buffer have been accessed, clear the maximum number of target nodes from the memory buffer, determine the first no more than the maximum number of target nodes that have not been cached, and cache the determined target nodes from the memory to the memory buffer until all the multiple target nodes have been cached in the memory buffer.

[0269] Optionally, the leaf nodes in the last layer of the index include all search keys, the multiple prefetch layers include the last layer of the index, and the multiple nodes in each layer are arranged in ascending order of the smallest search key in the node. The target search key of the data to be queried is the search key between the lower bound search key and the upper bound search key of the target search key range;

[0270] A prediction module 1102, configured to:

[0271] Through a prefetch model, predict multiple target numbers corresponding to the lower bound search key and an upper bound leaf node number corresponding to the upper bound search key. The multiple target numbers include the lower bound leaf node number, where the lower bound leaf node number refers to the number of the leaf node where the lower bound search key is located in the last layer of the index, and the upper bound leaf node number refers to the number of the leaf node where the upper bound search key is located in the last layer of the index;

[0272] The cache module 1103, configured to:

[0273] Cache the multiple target nodes indicated by the multiple target numbers, and multiple leaf nodes indicated by the multiple leaf node numbers between the lower bound leaf node number and the upper bound leaf node number from the memory to the memory buffer.

[0274] Optionally, refer to Fig.12 , the prefetch model is created based on the smallest search key of each node in the target prefetch layer and the number of the node where it is located; the apparatus further includes an update module 1106, configured to:

[0275] In response to a modification operation on a node in the prefetch layer, modify the node in the prefetch layer, and add the modification record corresponding to the modification operation to the log;

[0276] When the current update condition is met, based on the modification records in the log, determine the search keys and numbers of the nodes in the updated prefetch layer, and update the prefetch model based on the search keys and numbers of the nodes in the updated prefetch layer;

[0277] Delete the modification records in the log.

[0278] Optionally, multiple prefetch layers also correspond to a mapping table, which stores the mapping relationship between the numbers of the nodes in the prefetch layer and the storage addresses of the nodes; the update module 1106 is further configured to:

[0279] When the current update condition is met, based on the modification records in the log, update the mapping relationship between the numbers of the nodes in the prefetch layer and the storage addresses stored in the mapping table.

[0280] Optionally, the update condition includes any one of the following:

[0281] The data volume of the modification records stored in the log reaches a preset threshold;

[0282] The cache hit rate of the memory buffer is less than the preset hit rate, and the cache hit rate refers to the probability that the nodes to be accessed have been cached in the memory buffer.

[0283] It should be noted that: for the data query device provided in the above embodiments, only the above division of each functional module is used for illustration. In practical applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of the computer device is divided into different functional modules to complete all or part of the functions described above. In addition, the data query device provided in the above embodiments and the embodiments of the data query method belong to the same concept, and the specific implementation process is detailed in the method embodiments, which will not be repeated here.

[0284] The embodiment of the present application also provides a computer device, which includes a processor and a memory. At least one computer program is stored in the memory, and the at least one computer program is loaded and executed by the processor to implement the operations performed in the data query method in the above embodiments.

[0285] Optionally, the computer device is provided as a terminal. Fig.13 The structural schematic diagram of the terminal 1300 provided by an exemplary embodiment of the present application is shown.

[0286] The terminal 1300 includes: a processor 1301 and a memory 1302.

[0287] The processor 1301 may include one or more processing cores, such as a quad-core processor, an octa-core processor, etc. The processor 1301 may be implemented in at least one hardware form of DSP (Digital Signal Processing), FPGA (Field Programmable Gate Array), or PLA (Programmable Logic Array). The processor 1301 may also include a main processor and a coprocessor. The main processor is a processor for processing data in the wake state, also known as the CPU (Central Processing Unit); the coprocessor is a low-power processor for processing data in the standby state. In some embodiments, the processor 1301 may be integrated with a GPU (Graphics Processing Unit), and the GPU is responsible for rendering and drawing the content to be displayed on the display screen. In some embodiments, the processor 1301 may further include an AI (Artificial Intelligence) processor, and the AI processor is used to process computational operations related to machine learning.

[0288] The memory 1302 may include one or more computer-readable storage media, and the computer-readable storage media may be non-transitory. The memory 1302 may further include high-speed random access memory and non-volatile memory, such as one or more disk storage devices and flash storage devices. In some embodiments, the non-transitory computer-readable storage media in the memory 1302 is used to store at least one computer program, and the at least one computer program is used to be possessed by the processor 1301 to implement the data query method provided in the method embodiments of the present application.

[0289] In some embodiments, the terminal 1300 may further optionally include: a peripheral device interface 1303 and at least one peripheral device. The processor 1301, the memory 1302, and the peripheral device interface 1303 may be connected through a bus or signal lines. Each peripheral device may be connected to the peripheral device interface 1303 through a bus, signal lines, or a circuit board. Optionally, the peripheral device includes at least one of a radio frequency circuit 1304, a display screen 1305, a camera assembly 1306, an audio circuit 1307, and a power supply 1308.

[0290] The peripheral device interface 1303 can be used to connect at least one I / O (Input / Output) related peripheral device to the processor 1301 and the memory 1302. In some embodiments, the processor 1301, the memory 1302, and the peripheral device interface 1303 are integrated on the same chip or circuit board; in some other embodiments, any one or two of the processor 1301, the memory 1302, and the peripheral device interface 1303 can be implemented on a separate chip or circuit board, and this embodiment does not limit this.

[0291] The radio frequency circuit 1304 is used to receive and transmit RF (Radio Frequency) signals, also known as electromagnetic signals. The radio frequency circuit 1304 communicates with the communication network and other communication devices through electromagnetic signals. The radio frequency circuit 1304 converts an electrical signal into an electromagnetic signal for transmission, or converts the received electromagnetic signal into an electrical signal. Optionally, the radio frequency circuit 1304 includes: an antenna system, an RF transceiver, one or more amplifiers, a tuner, an oscillator, a digital signal processor, a codec chipset, a subscriber identity module card, and so on. The radio frequency circuit 1304 can communicate with other devices through at least one wireless communication protocol. The wireless communication protocol includes but is not limited to: metropolitan area network, each generation of mobile communication network (2G, 3G, 4G, and 5G), wireless local area network, and / or WiFi (Wireless Fidelity) network. In some embodiments, the radio frequency circuit 1304 may further include a circuit related to NFC (Near Field Communication), and this application does not limit this.

[0292] The display screen 1305 is used to display the UI (User Interface). The UI may include graphics, text, icons, videos, and any combination thereof. When the display screen 1305 is a touch display screen, the display screen 1305 also has the ability to collect touch signals on or above the surface of the display screen 1305. The touch signals can be input as control signals to the processor 1301 for processing. At this time, the display screen 1305 can also be used to provide virtual buttons and / or a virtual keyboard, also known as soft buttons and / or a soft keyboard. In some embodiments, there may be one display screen 1305, which is provided on the front panel of the terminal 1300; in other embodiments, there may be at least two display screens 1305, which are respectively provided on different surfaces of the terminal 1300 or are in a foldable design; in other embodiments, the display screen 1305 may be a flexible display screen, which is provided on the curved surface or the folding surface of the terminal 1300. Even further, the display screen 1305 can also be set to an irregular non-rectangular shape, that is, a special-shaped screen. The display screen 1305 can be prepared using materials such as LCD (Liquid Crystal Display) and OLED (Organic Light-Emitting Diode).

[0293] The camera module 1306 is used to capture images or videos. Optionally, the camera module 1306 includes a front camera and a rear camera. The front camera is provided on the front panel of the terminal 1300, and the rear camera is provided on the back of the terminal 1300. In some embodiments, there are at least two rear cameras, which are any one of a main camera, a depth-of-field camera, a wide-angle camera, and a telephoto camera respectively, to implement functions such as background blurring by fusing the main camera and the depth-of-field camera, panoramic shooting by fusing the main camera and the wide-angle camera, and VR (Virtual Reality) shooting functions or other fused shooting functions. In some embodiments, the camera module 1306 may further include a flash. The flash can be a single-color temperature flash or a two-color temperature flash. A two-color temperature flash refers to a combination of a warm light flash and a cold light flash, which can be used for light compensation under different color temperatures.

[0294] The audio circuit 1307 may include a microphone and a speaker. The microphone is used to collect sound waves of the user and the environment, and convert the sound waves into electrical signals for input to the processor 1301 for processing, or input to the radio frequency circuit 1304 to implement voice communication. For the purpose of stereo collection or noise reduction, there may be multiple microphones, which are respectively arranged at different parts of the terminal 1300. The microphone may also be an array microphone or an omnidirectional collection microphone. The speaker is used to convert the electrical signal from the processor 1301 or the radio frequency circuit 1304 into sound waves. The speaker may be a traditional thin film speaker or a piezoelectric ceramic speaker. When the speaker is a piezoelectric ceramic speaker, it can not only convert the electrical signal into sound waves audible to humans, but also convert the electrical signal into sound waves inaudible to humans for uses such as ranging. In some embodiments, the audio circuit 1307 may further include a headphone jack.

[0295] The power supply 1308 is used to supply power to each component in the terminal 1300. The power supply 1308 may be alternating current, direct current, a disposable battery or a rechargeable battery. When the power supply 1308 includes a rechargeable battery, the rechargeable battery may support wired charging or wireless charging. The rechargeable battery may also be used to support fast charging technology.

[0296] Those skilled in the art can understand that Fig.13 the structure shown in

[0297] does not limit the terminal 1300, and may include more or fewer components than shown in the figure, or combine some components, or adopt different component arrangements. Fig.14 FIG. is a schematic structural diagram of a server provided by an embodiment of the present application. The server 1400 may vary greatly due to different configurations or performances, and may include one or more processors (Central Processing Units, CPUs) 1401 and one or more memories 1402. Among them, at least one computer program is stored in the memory 1402, and the at least one computer program is loaded and executed by the processor 1401 to implement the methods provided by the above-mentioned method embodiments. Of course, the server may also have components such as a wired or wireless network interface, a keyboard, and an input / output interface for input / output. The server may further include other components for implementing the functions of the device, which will not be elaborated here.

[0298] The embodiment of the present application further provides a computer-readable storage medium, in which at least one computer program is stored, and the at least one computer program is loaded and executed by a processor to implement the operations performed by the data query method in the above embodiment.

[0299] The embodiments of the present application also provide a computer program product, including a computer program, which is loaded and executed by a processor to implement the operations performed by the data query method in the above embodiments. In some embodiments, the computer program involved in the embodiments of the present application can be deployed to be executed on a computer device, or on multiple computer devices located at one place, or on multiple computer devices distributed at multiple places and interconnected through a communication network. The multiple computer devices distributed at multiple places and interconnected through a communication network can form a blockchain system.

[0300] Those of ordinary skill in the art can understand that all or part of the steps to implement the above embodiments can be completed by hardware, or can be completed by a program instructing relevant hardware. The program can be stored in a computer-readable storage medium, and the above-mentioned storage medium can be a read-only memory, a disk, or an optical disc, etc.

[0301] The above are only optional embodiments of the embodiments of the present application, and are not intended to limit the embodiments of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the embodiments of the present application shall be included in the protection scope of the present application.

Claims

1. A data query method, characterized in that, The method includes: Obtaining an index and a target search key of data to be queried. The index includes multiple layers, each layer includes nodes, each node includes a search key, each node has a number, the multiple layers include multiple prefetch layers, and corresponding prefetch models are provided for the multiple prefetch layers. The prefetch model is used to predict the number corresponding to any search key. The number corresponding to the search key refers to the number of the node on the access path in the prefetch layer when querying the search key. Predicting, through the prefetch model, multiple target numbers corresponding to the target search key. Caching multiple target nodes indicated by the multiple target numbers from the memory to the memory buffer. Querying the target search key in the index, and when the currently required accessed node belongs to the target node, accessing the cached target node in the memory buffer.

2. The method according to claim 1, wherein The caching of multiple target nodes indicated by the multiple target numbers from the memory to the memory buffer includes: Obtaining a mapping table, where the mapping table stores the mapping relationship between the numbers of nodes in the prefetch layer and the storage addresses of the nodes. Querying, in the mapping table, multiple target storage addresses corresponding to the multiple target numbers. Reading target nodes at the multiple target storage addresses in the memory, and caching the read multiple target nodes to the memory buffer.

3. The method according to claim 1, wherein Each prefetch layer corresponds to a prefetch model. The predicting, through the prefetch model, multiple target numbers corresponding to the target search key includes: For a target prefetch layer among the multiple prefetch layers, predicting, through the prefetch model corresponding to the target prefetch layer, a target number corresponding to the target search key in the target prefetch layer. The target number corresponding to the target prefetch layer refers to the number of the node on the access path in the target prefetch layer when querying the target search key, and the target prefetch layer is any one of the multiple prefetch layers.

4. The method according to claim 3, characterized in that, The predicting, through the prefetch model corresponding to the target prefetch layer, a target number corresponding to the target search key in the target prefetch layer includes: Inputting the target search key into the prefetch model to obtain a triple output by the prefetch model. The triple includes a reference number, an upper bound of number error, and a lower bound of number error. Determining the sum value of the reference number and the upper bound of number error as the maximum number, and determining the difference value between the reference number and the lower bound of number error as the minimum number. Determining a target number among multiple numbers from the minimum number to the maximum number.

5. The method according to claim 4, wherein The multiple search keys in each node are arranged in ascending order, and the multiple nodes in each layer are arranged in ascending order of the search keys in the nodes. The creation process of the prefetch model corresponding to the target prefetch layer includes: Numbering the multiple nodes according to the arrangement order of the multiple nodes in the target prefetch layer. Create a prefetch model corresponding to the target prefetch layer based on the minimum search key of each node in the target prefetch layer and the number of the node where it is located. The prefetch model is used to linearly approximate the functional relationship between the minimum search keys of multiple nodes in the target prefetch layer and the numbers of the nodes where they are located.

6. The method according to claim 5, characterized in that, Determining a target number among multiple numbers from the smallest number to the largest number includes: Determine at least one candidate number among multiple numbers from the smallest number to the largest number. The minimum search key of the node indicated by the candidate number is less than the target search key. Among the at least one candidate number, determine the candidate number with the largest minimum search key of the indicated node as the target number.

7. The method according to claim 4, characterized in that, The multiple search keys in each node are arranged in ascending order, and the multiple nodes in each layer are arranged in ascending order of the search keys in the nodes. The creation process of the prefetch model corresponding to the target prefetch layer includes: Number the multiple nodes in accordance with the arrangement order of the multiple nodes in the target prefetch layer. Create a prefetch model corresponding to the target prefetch layer based on the maximum search key of each node in the target prefetch layer and the number of the node where it is located. The prefetch model is used to linearly approximate the functional relationship between the maximum search keys of multiple nodes in the target prefetch layer and the numbers of the nodes where they are located.

8. The method according to claim 7, characterized in that, Determining a target number among multiple numbers from the smallest number to the largest number includes: Determine at least one candidate number among multiple numbers from the smallest number to the largest number. The maximum search key of the node indicated by the candidate number is greater than the target search key. Among the at least one candidate number, determine the candidate number with the smallest maximum search key of the indicated node as the target number.

9. The method according to claim 1, wherein The memory has a maximum request quantity, which refers to the maximum number of access requests that the memory can process simultaneously. The multiple target nodes are arranged in sequence according to the access order. Caching the multiple target nodes indicated by the multiple target numbers from the memory to the memory buffer includes: In the case where the number of the multiple target nodes is greater than the maximum request quantity, determine the first maximum request quantity of target nodes in accordance with the arrangement order of the multiple target nodes, and cache the determined target nodes from the memory to the memory buffer. The method further includes: When all the maximum request quantity of target nodes cached in the memory buffer have been accessed, determine the first non-cached target nodes not greater than the maximum request quantity, and cache the determined target nodes from the memory to the memory buffer until all the multiple target nodes have been cached to the memory buffer.

10. The method according to claim 1, characterized in that The leaf nodes in the last layer of the index include all search keys. The multiple prefetch layers include the last layer of the index. The multiple nodes in each layer are arranged in ascending order of the minimum search keys in the nodes. The target search key of the data to be queried is a search key between the lower bound search key and the upper bound search key of the target search key range. Predicting, by the prefetching model, a plurality of target numbers corresponding to the target search key includes: Predicting, by the prefetching model, a plurality of target numbers corresponding to the lower-bound search key and an upper-bound leaf node number corresponding to the upper-bound search key, where the plurality of target numbers include a lower-bound leaf node number, and the lower-bound leaf node number refers to the number of the leaf node where the lower-bound search key is located in the last layer of the index, and the upper-bound leaf node number refers to the number of the leaf node where the upper-bound search key is located in the last layer of the index; Caching, in the memory buffer, a plurality of target nodes indicated by the plurality of target numbers includes: Caching, in the memory buffer, a plurality of target nodes indicated by the plurality of target numbers, and a plurality of leaf nodes indicated by a plurality of leaf node numbers from the lower-bound leaf node number to the upper-bound leaf node number, from the memory.

11. The method according to claim 1, characterized in that, The prefetching model is created based on the minimum search key of each node in the target prefetching layer and the number of the node where it is located. The method further includes: In response to a modification operation on a node in the prefetching layer, modifying the node in the prefetching layer, and adding a modification record corresponding to the modification operation to the log; When the current update condition is satisfied, determining, based on the modification records in the log, the search key and number of the nodes in the updated prefetching layer, and updating the prefetching model based on the search key and number of the nodes in the updated prefetching layer; Deleting the modification records in the log.

12. The method according to claim 11, wherein The plurality of prefetching layers further correspond to a mapping table, and the mapping table stores a mapping relationship between the number of a node in the prefetching layer and the storage address of the node; Before deleting the modification records in the log, the method further includes: When the current update condition is satisfied, updating, based on the modification records in the log, the mapping relationship between the number of a node in the prefetching layer and the storage address stored in the mapping table.

13. The method according to claim 11, wherein The update condition includes any one of the following: The data volume of the modification records stored in the log reaches a preset threshold; The cache hit rate of the memory buffer is less than a preset hit rate, where the cache hit rate refers to the probability that the node to be accessed has been cached in the memory buffer.

14. A data query device, characterized in that, The apparatus includes: An acquisition module, configured to acquire an index and a target search key of data to be queried, where the index includes a plurality of layers, each layer includes nodes, each node includes a search key, each node has a number, the plurality of layers include a plurality of prefetching layers, the plurality of prefetching layers correspond to a prefetching model, and the prefetching model is used to predict the number corresponding to any search key, and the number corresponding to the search key refers to the number of the node on the access path in the prefetching layer when querying the search key; A prediction module, configured to predict, by the prefetching model, a plurality of target numbers corresponding to the target search key; A cache module, configured to cache, in the memory buffer, a plurality of target nodes indicated by the plurality of target numbers; A query module for querying the target search key in the index, and when the currently required node to be accessed belongs to the target node, accessing the cached target node in the memory buffer.

15. A computer device, characterized in that, The computer device includes a processor and a memory, and at least one computer program is stored in the memory. The at least one computer program is loaded and executed by the processor to implement the operations performed by the data query method according to any one of claims 1 to 13.

16. A computer-readable storage medium, characterized in that, At least one computer program is stored in the computer-readable storage medium. The at least one computer program is loaded and executed by a processor to implement the operations performed by the data query method according to any one of claims 1 to 13.

17. A computer program product, comprising a computer program, characterized in that, The computer program is loaded and executed by a processor to implement the operations performed by the data query method according to any one of claims 1 to 13.