Method for processing data search requests, processing device for data search requests, computer device, and computer program

The method enhances the efficiency of data search requests in distributed graph database systems by directly processing requests through determined processing models within the node device, eliminating the need for thread creation and thus speeding up search operations.

JP2025514920APending Publication Date: 2025-05-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2024560225
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2022-07-27
Filing Date
2023-05-23
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

In distributed data storage systems based on graph databases, processing parallel data search requests efficiently is challenging due to the need to create multiple threads, which increases search time.

Method used

A method where a first node device receives a data search request, determines a second processing model to process data at the starting vertex, and further determines a third processing model to process data of the first vertex or intermediate vertex, allowing direct processing without creating threads.

Benefits of technology

This approach enables efficient processing of data search requests by eliminating the need for thread creation, thereby increasing search speed and improving system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025514920000001_ABST
    Figure 2025514920000001_ABST
Patent Text Reader

Abstract

The present application discloses a method, an apparatus, a device, and a storage medium for processing a data search request, which belong to the field of databases. The method includes a step (502) of a first node device receiving a data search request transmitted from a second node device by a first processing model in the first node device, the data search request being for searching data related to a first vertex in a graph database, the data search request being assigned an identifier of a starting vertex, and the first node device storing data of the starting vertex; a step (504) of transmitting the data search request to a second processing model in the first node device by the first processing model, the second processing model being a processing model for processing data of the starting vertex in the first node device; a step (506) of determining a third processing model based on the data search request by the second processing model; and a step (508) of transmitting the data search request to the third processing model by the second processing model. According to the present application, it is possible to realize an increase in the speed of data search.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical field]

[0001] This application claims priority from a Chinese patent application filed on July 27, 2022, bearing application number 202210891780.8 and entitled "Method, device, apparatus, and storage medium for processing data search requests," the entire contents of which are incorporated herein by reference.

[0002] The present application relates to the field of databases, and in particular to a method, an apparatus, a device and a storage medium for processing a data search request. [Background technology]

[0003] Data based on a graph database is typically stored using a distributed data storage system, which shards and stores data on multiple node devices to support large-scale data storage.

[0004] Each node device that stores a data shard uses each core of a multi-core processor in the node device to jointly manage all data stored in the node device. When a node device receives a data search request for certain data stored therein, it creates a thread in response to the data search request, processes the data search request in the created thread, and obtains the search results.

[0005] When a distributed data storage system receives a parallel data search request, for example, a data search request for searching secondary neighbors of a start vertex in a graph database, the distributed data storage system searches data stored in different node devices in parallel. At this time, each node device that receives the data search request needs to create a thread to be used for each different data search request. If a large number of threads need to be created, the time required for data search increases. Summary of the Invention [Problem to be solved by the invention]

[0006] The present application provides a method, an apparatus, a device, and a storage medium for processing a data search request, which are configured as follows: [Means for solving the problem]

[0007] According to one aspect of the present application, there is provided a method for processing a data search request, the method being applied to a first node device, the first node device being one of at least two node devices in a distributed data storage system based on a graph database, the method comprising: a step of receiving, by the first node device, a data search request transmitted from a second node device by a first processing model in the first node device, the data search request being for searching data related to a first vertex in the graph database, the data search request being assigned an identifier of a starting vertex, and data of the starting vertex being stored in the first node device; a step of the first node device transmitting the data search request to a second processing model in the first node device by the first processing model, the second processing model being a processing model for processing data of the starting vertex in the first node device; a step of the first node device determining a third processing model based on the data search request by the second processing model, the third processing model being for processing data of the first vertex or an intermediate vertex, the intermediate vertex being a vertex between the starting vertex and the first vertex; The first node device transmits the data search request to the third processing model via the second processing model.

[0008] According to another aspect of the present application, there is provided an apparatus for processing a data search request, the apparatus being one of at least two node devices in a distributed data storage system based on a graph database, a receiving module that receives a data search request transmitted from a second node device by a first processing model in the device, the data search request being for searching data related to a first vertex in the graph database, the data search request being assigned an identifier of a starting vertex, and data of the starting vertex being stored in the device; and a sending module that sends the data search request by the first processing model to a second processing model in the device, the second processing model being a processing model for processing data of the starting vertex in the device; a determining module for determining a third processing model based on the data retrieval request by the second processing model, the third processing model being for processing data of the first vertex or an intermediate vertex, the intermediate vertex being a vertex between the starting vertex and the first vertex; The sending module sends the data retrieval request via the second processing model to the third processing model.

[0009] According to another aspect of the present application, there is provided a computing device comprising a processor and a memory, the memory storing at least one program, the at least one program, when loaded and executed by the processor, realizing the method for processing a data search request described in the above aspect.

[0010] According to another aspect of the present application, there is provided a computer-readable storage medium storing at least one program, which, when loaded and executed by a processor, realizes the method for processing a data retrieval request described in the above aspect.

[0011] According to another aspect of the present application, there is provided a computer program product or computer program comprising computer instructions, the computer instructions being stored in a computer readable storage medium, which a processor of a computing device reads from the computer readable storage medium and, when executed by the processor, causes the computing device to perform the method of processing a data retrieval request provided in various optional implementations of the above aspect. Effect of the Invention

[0012] The beneficial effects of the configuration provided herein include at least the following:

[0013] After a node device receives a data search request, the node device can directly process the data search request according to a processing model in the node device, thereby determining a processing model for processing the vertex data related to the data search request in the graph database, and further determining the processing result of the data search request. In the process of processing the data search request, there is no need to create a thread for processing the request, and the data search request is directly processed according to the processing model, thereby avoiding the creation of a corresponding thread when processing the data search request and accelerating the data search speed. [Brief description of the drawings]

[0014] [Figure 1] FIG. 2 is a schematic diagram of an attribute graph of a graph database provided in one exemplary embodiment of the present application. [Diagram 2] FIG. 2 is a schematic diagram of a B-tree based data structure provided in one exemplary embodiment of the present application; [Diagram 3] FIG. 2 is a schematic diagram of a computer system configuration provided in one exemplary embodiment of the present application. [Figure 4] FIG. 2 is a schematic diagram of a data retrieval request processing process provided in one exemplary embodiment of the present application; [Diagram 5]2 is a schematic flow diagram of a method for processing a data retrieval request provided in one exemplary embodiment of the present application; [Figure 6] FIG. 2 is a schematic flow diagram of another data retrieval request processing method provided in one exemplary embodiment of the present application. [Figure 7] FIG. 2 is a schematic diagram of a global consistent hash ring provided in one exemplary embodiment of the present application. [Figure 8] FIG. 2 is a schematic diagram of a local consistent hash ring provided in one exemplary embodiment of the present application; [Figure 9] FIG. 2 is a schematic diagram of a storage structure of vertices in a graph database provided in one exemplary embodiment of the present application. [Figure 10] FIG. 2 is a schematic diagram of the configuration of a data retrieval request processing device provided in one exemplary embodiment of the present application; [Figure 11] FIG. 2 is a schematic diagram of a configuration of a computer device provided in one exemplary embodiment of the present application. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS

[0015] First, the terminology used in this application will be introduced.

[0016] Graph database: a database that uses a graph structure to perform semantic search. A graph database represents and stores data by vertices, edges, and attributes. For example, FIG. 1 is an attribute graph of a graph database provided in one exemplary embodiment of the present application. As shown in FIG. 1, the attribute graph of a graph database is composed of vertices (circles), edges (arrows), and attributes. Each of the three vertices in FIG. 1 has a student label, and the attributes that the student has include name, and the values ​​of the attribute are user 1, user 2, and user 3, respectively. The edges between the vertices represent the relationship between the vertices, for example, older brother, younger brother, and have met. If the edge between two vertices is bidirectional, it means that the data of the two vertices has a bidirectional association relationship, and if the edge between two vertices is unidirectional, it means that the data of the two vertices has a unidirectional association relationship.

[0017] B-tree (Balance tree): A self-balancing tree-like data structure that can keep data in order. A data structure based on a B-tree can complete operations such as data search, sequential access, data insertion, and data deletion in logarithmic time, that is, as the amount of data increases, the time required for data processing increases logarithmically rather than exponentially, making data processing highly efficient.

[0018] For illustrative purposes, Fig. 2 is a schematic diagram of a data structure based on a B-tree provided in one exemplary embodiment of the present application. As shown in Fig. 2, this B-tree includes seven nodes, and each node is configured with an index of data. Here, node (6) is called the root node, nodes (4) and (11) are called intermediate nodes, and nodes (1, 2, 3), (5), (7, 8, 9), and (12, 13) are called leaf nodes.

[0019] B+ trees are an improvement of B trees, and differ from B trees in that their leaf nodes contain indexes for all data. Data structures such as B+ trees have the characteristics of good search performance and clear graph representation, and are therefore widely applied to graph databases. However, the process of parallel writing of data (writing data indexes of multiple data simultaneously) in data structures such as trees (B trees, B+ trees) is cumbersome because the data writing process involves splitting and merging of nodes. Therefore, in general, when performing parallel writing operations on B trees / B+ trees, different granularity locking processes are performed on the B trees / B+ trees to realize sequential single writing to the B trees / B+ trees. For example, the writing operation is realized by locking the entire tree. Currently, a finer granularity locking method can also be used to handle highly parallel writing cases.

[0020] Continuing with FIG2, when writing index (10) to the B-tree, node (7, 8, 9) is split and index (8) is merged into node (11) to form node (8, 11). When writing multiple data indexes in parallel, it is necessary to lock the B-tree and process the write requests sequentially in a sequential single-write manner. However, this leads to low efficiency in data processing.

[0021] Actor model: A type of mathematical model of concurrent computation. In response to a received message (task), the actor model can make local decisions about the task to be performed, and can also create more actor models to send more messages, and decide how to respond to the next message received. Messages are sent directly between actor models without any intermediate between them, and message sending and processing are asynchronous. Communication and interaction between modules is realized by sending messages between different actor models. The actor model adopts the philosophy that "everything is an actor". This is similar to the philosophy that "everything is an object" applied to some object-oriented programming languages, so its message sending is more in line with the original intent of object orientation. The actor model belongs to the concurrent component model. By defining the high-level stage of the concurrent programming paradigm in a component manner, users can avoid direct contact with basic concepts such as multithreading parallelism and thread pools.

[0022] Cloud storage: A new concept that has developed from the concept of cloud computing. A distributed cloud storage system (hereinafter abbreviated as storage system) refers to a storage system in which a large number of different types of storage devices (storage devices are also called storage nodes) in a network assemble and cooperate through application software and application interfaces by using functions such as cluster applications, grid technology, and distributed storage file systems to jointly provide data storage and business access functions to the outside. Illustratively, the distributed data storage system in the embodiment of the present application is a cloud storage system.

[0023] Currently, the storage method of the storage system is as follows: A logical volume is created, and when the logical volume is created, a physical storage space is allocated for each logical volume. The physical storage space may be composed of disks of a storage device or several storage devices. A client stores data in a logical volume, that is, stores the data in a file system. The file system divides the data into many parts, each part being an object, and the object includes not only data but also additional information such as a data identifier (ID: IDentity). The file system writes each object into the physical storage space of the logical volume respectively. Then, the file system stores storage location information of each object. In this way, when a client requests data access, the file system can allow the client to access the data based on the storage location information of each object. The process of the storage system allocating the physical storage space to the logical volume specifically includes dividing the physical storage space into stripes in advance according to an estimate of the capacity of the object to be stored in the logical volume (in many cases, this estimate has a large margin compared to the capacity of the object actually stored) and a group of redundant array of independent disks (RAID: Redundant Array of Independent Disks). A logical volume can be thought of as a stripe, which is how physical storage space is allocated to it.

[0024] Database: Simply put, it can be regarded as an electronic cabinet, that is, a place for storing electronic files, and users can add, search, update, delete, and perform other operations on the data in the files. A "database" is a collection of data that is stored together in a certain manner, can be shared by multiple users, has as little redundancy as possible, and is independent of applications. Illustratively, the database in the embodiment of this application is a graph database, and the data structure adopts a B-tree or B+tree structure.

[0025] A database management system (DBMS) is a computer software system designed to manage databases, typically with basic functions such as storage, pruning, security, and backup. Database management systems can be classified by the database model supported (e.g., relational, eXtensible Markup Language (XML), etc.), or by the type of computer supported (e.g., server cluster, cell phone), or by the query language used (e.g., Structured Query Language (SQL), XML Query, X Query), or by the emphasis of performance impulses (e.g., largest scale, fastest execution speed), or by another classification scheme. With any classification scheme, some DBMSs can cross categories, e.g., support multiple query languages ​​simultaneously.

[0026] 3 is a schematic diagram of a configuration of a computer system 300 provided in an exemplary embodiment of the present application. As shown in FIG. 3, the computer system 300 includes a first node device 301, a second node device 302, and a third node device 303.

[0027] The second node device 302 may be one server, a server cluster consisting of a few servers, or a virtual server in one cloud computing service center. Optionally, the second node device 302 is a node device for dispatching (also called "scheduling") data processing requests in the distributed data storage system. For example, the second node device 302 receives an external data search request, and dispatches the data search request based on the search target data of the data search request to a node device that stores the search target data or a node device that stores data related to the search target data, to perform processing. A connection may be established between the second node device 302 and the first node device 301 and the third node device 303 via a wired network or a wireless network.

[0028] The first node device 301 and the third node device 303 may be one server, a server cluster consisting of a few servers, or a virtual server in one cloud computing service center. The first node device 301 and the third node device 303 store the same, different, or partially identical data shards, thereby realizing distributed storage of data. When the same data is stored in different node devices, that is, when data copies are stored, the consistency problem of the data copies can be converted into a computation problem by using a conflict-free replicated data type (CRDT). In this way, the lock problem in the process of concurrently modifying data can be reduced. A connection may be established between the first node device 301 and the third node device 303 by a wired network or a wireless network.

[0029] It should be noted that the number of node devices in the computer system 300 described above is merely exemplary and does not limit the computer system 300 provided in the embodiment of the present application.

[0030] FIG. 4 is a schematic diagram of a processing process of a data search request provided in one exemplary embodiment of the present application. As shown in FIG. 4, a second node device is a node device in a distributed data storage system based on a graph database. After receiving a data search request 401, the second node device sends the data search request 401 to a first processing model 4021 of a first node device 402. Here, the data search request is for searching data related to a first vertex in a graph database, the data search request is assigned an identifier of a starting vertex, and the first node device 402 stores data of the starting vertex. The first vertex is also called a target vertex. The target vertex is a vertex for processing the data search request. In other words, the target vertex is a vertex for providing a processing result of the data search request.

[0031] The first node device 402 transmits the data search request 401 to a second processing model 4022 in the first node device 402 by a first processing model 4021. The second processing model 4022 is a processing model for processing data of the starting vertex in the first node device 402. The first node device 402 determines a third processing model based on the data search request 401 by the second processing model 4022. The third processing model is for processing data of the first vertex or an intermediate vertex, where the intermediate vertex is a vertex between the starting vertex and the first vertex.

[0032] Optionally, if the first vertex exists among the vertices connected to the first node device, the first node device 402 can directly determine the first vertex through the second processing model 4022 and determine a third processing model for processing the data of the first vertex. If the first vertex does not exist among the vertices connected to the first node device, the first node device 402 can determine an intermediate vertex connected to the first vertex through the second processing model 4022 and determine a third processing model for processing the data of the intermediate vertex. After determining the third processing model, the first node device 402 sends the data search request to the third processing model through the second processing model 4022. If the third processing model is used to process the data of the first vertex, the third processing model processes the data search request and sends the processing result of the data search request to the second node device. If the third processing model is used to process the data of the intermediate vertex, the third processing model continues to determine a fourth processing model and continue to send the data search request through the above method until the processing result of the data search request is obtained.

[0033] Optionally, the third processing model belongs to the first node device 402 or the third node device 403. If the third processing model belongs to the first node device 402, the first node device 402 directly sends the data search request to the third processing model through the second processing model 4022. If the third processing model belongs to the third node device 403, the first node device 402 sends the data search request to a processing model in the third node device 403 for dispatching the data search request through the second processing model 4022. Then, the processing model sends the data search request to the third processing model.

[0034] After determining a target processing model (e.g., a third processing model) for processing the data of the first vertex through the above method, the target processing model processes the data search request and transmits the processing result of the data search request to a processing model for dispatching the data search request in the node device where it is located. Then, the processing model for dispatching the data search request transmits the processing result of the data search request to the second node device. If there are multiple processing models that transmit the processing result of the data search request to the second node device, the second node device finally obtains the processing result of the data search request by performing a consolidation process of the processing results of the data search request transmitted from all the processing models. Optionally, the above processing model is an actor model. Here, the above "multiple" may be understood to mean "at least two" and more than two.

[0035] It should be noted that the methods provided in the embodiments of the present application are applicable to scenarios related to distributed data storage systems, scenarios related to big data processing, scenarios related to processing training data related to machine learning models, scenarios related to processing data in applications (including applications in fields such as social, short video, instant messaging, games, e-commerce, finance, etc.), and scenarios related to data exploration (e.g., product exploration scenarios).

[0036] Exemplarily, when the method provided in the embodiment of the present application is applied to a data processing scenario of a social application, the vertices in the graph database represent data related to a certain user account, the edges between the vertices represent that there is an association relationship between the user accounts, and the attributes of the edges are for reflecting the attributes of the association relationship, such as a friend relationship, a blacklist relationship, a block relationship, etc. The node device is a back-end server of the social application, and the processing model is an actor model in the back-end server. Exemplarily, the start node is a node corresponding to a first user account, and the target node is a node corresponding to a second user account among the friend accounts of the first user account, which has been in a friendship relationship with the first user account for the longest period of time. According to the method provided in the embodiment of the present application, it is possible to realize a search for a second user account that satisfies the condition that the friend relationship has been established for the longest period of time among the friend accounts of the first user account, and to search for data related to the second user account.

[0037] Exemplarily, when the method provided in the embodiment of the present application is applied to a data processing scenario of a short video application, the vertices in the graph database represent data related to a certain user account, the edges between the vertices represent that there is an association relationship between the user accounts, and the attributes of the edges are for reflecting the attributes of the association relationship, such as a follow relationship, a like relationship, a favorite relationship, etc. The node device is a back-end server of the short video application, and the processing model is an actor model in the back-end server. Exemplarily, the start node is a node corresponding to a third user account, and the target node is a node corresponding to a fourth user account that similarly follows the third user account among the user accounts followed by the third user account. According to the method provided in the embodiment of the present application, it is possible to realize a search for a fourth user account that satisfies the condition of following the third user account among the user accounts followed by the third user account, and data related to the fourth user account can be searched for.

[0038] Exemplarily, when the method provided in the embodiment of the present application is applied to a data processing scenario of a game-based application, the vertices in the graph database represent data related to a certain user account, the edges between the vertices represent interaction behaviors between the user accounts, and the attributes of the edges are for reflecting the attributes of the interaction behaviors, for example, reflecting that there is a trading behavior of a virtual item between the user accounts. The node device is a back-end server of the game-based application, and the processing model is an actor model in the back-end server. Exemplarily, the start node is a node corresponding to the fifth user account, and the target node is a node corresponding to the sixth user account having the highest price of the virtual item to be traded among the user accounts having the trading behavior of a virtual item with the fifth user account. According to the method provided in the embodiment of the present application, it is possible to realize a search for the sixth user account that satisfies the condition that the price of the virtual item to be traded is the highest among the user accounts having the trading behavior of a virtual item with the fifth user account, and data related to the sixth user account can be searched for.

[0039] Optionally, the method provided in the embodiments of the present application is also applicable to various scenarios, such as cloud technology, artificial intelligence, Internet of Vehicles, smart transportation, driving assistance, etc., and processes relevant data related to the above scenarios. When applied to a scenario related to the Internet of Vehicles, the above node device may be an in-vehicle device, and the in-vehicle device may be an in-vehicle terminal.

[0040] After receiving the data search request, the node device can directly process the data search request according to a processing model in the node device. This processing includes determining a processing model of the data of the vertex (in the graph database for the data search request) or determining a processing result of the data search request. In the process of processing the data search request, there is no need to create a thread for processing the request, and the data search request is directly processed according to the processing model. This can avoid creating a corresponding thread when processing the data search request, and can speed up the data search.

[0041] Fig. 5 is a schematic flow diagram of a data search request processing method provided in one exemplary embodiment of the present application. The method is applied to a first node device in a system as shown in Fig. 3, and the first node device is one of a plurality of node devices in a distributed data storage system based on a graph database. As shown in Fig. 5, the method includes the following steps:

[0042] In step 502, the first node device receives a data search request transmitted from the second node device by a first processing model in the first node device.

[0043] The second node device is a node device for dispatching a data search request in the distributed data storage system, and can also be used to dispatch data processing requests, such as a data write request, a data delete request, and a data modification request.

[0044] The data search request is for searching data related to a first vertex in a graph database, and the data search request is accompanied by an identifier of a starting vertex. Optionally, the data search request is used to search the first vertex based on the starting vertex, search data of the first vertex based on the starting vertex, search a path (edge ​​in the graph database) from the starting vertex to the first vertex, search edges or edge attributes between the starting vertex and the first vertex, and search a subgraph including the starting vertex and the first vertex in the graph database. Optionally, the data search request may be accompanied by an identifier of the first vertex, for example, the data search request is used to search a path from the starting vertex to the first vertex.

[0045] Exemplarily, the data retrieval request is used to retrieve data of secondary neighborhoods of a starting vertex, i.e., data of vertices connected to an intermediate vertex, where the intermediate vertex is connected to the starting vertex and the intermediate vertex is connected to a first vertex, where there are one or more intermediate vertices and one or more first vertices.

[0046] The first node device stores data of the starting vertex. After receiving the data search request, the second node device determines the node device in which the data of the starting vertex is stored, and transmits the data search request to the processing model in the determined node device.

[0047] In addition, the node device includes at least two types of processing models: a processing model for task dispatching (e.g., a processing model for dispatching data search requests and data processing requests) and a processing model for data shard management (one processing model can manage one or more data shards, for example, a first node device may include a processing model for managing the data shard of vertex 1 and a processing model for managing the data shards of vertex 2 and vertex 3).

[0048] The first processing model is a model for dispatching a data search request (data processing request) in the first node device. That is, the first processing model belongs to a processing model for task dispatch. After receiving a data search request, the first processing model transmits the data search request to a processing model for processing data of a starting vertex in the first node device. Optionally, the first processing model does not have data of a vertex that needs to be processed, and a processing model with the same function as the first processing model exists in each of the node devices that store the data of the vertex in the distributed data storage system. Optionally, the second node device stores an address of a processing model for dispatching a data search request of a different node device, and the second node device transmits the data search request to an address corresponding to the first processing model after determining the first node device.

[0049] Optionally, multiple starting vertices are included, in which case the second node device decomposes the data search request for each starting vertex, determines a node device corresponding to each data search request, and sends the corresponding data search request to the corresponding processing model in the determined node device.

[0050] In step 504, the first node device transmits a data search request to a second processing model in the first node device via the first processing model.

[0051] The second processing model is a processing model for processing data of the starting vertex in the first node device. That is, the second processing model belongs to a processing model for data shard management. Optionally, each data shard of the multiple vertices stored in the first node device has a corresponding processing model. After receiving a data search request, the first processing model determines the second processing model by determining a processing model for processing data of the starting vertex in the first node device based on an identifier of the starting vertex.

[0052] Optionally, the first node device may determine a second processing model based on a local consistent hash ring according to the first processing model, and then send the data search request to the second processing model. The local consistent hash ring is for reflecting the corresponding relationship between the data shards of each vertex stored in the first node device and the processing model in the first node device. Optionally, when there are multiple starting vertices, the first node device also decomposes the data search request according to the first processing model before sending it.

[0053] Optionally, metadata including an address of each processing model in the first node device is stored locally in the first processing model, and the first node device transmits the data search request by the first processing model to a second processing model determined based on the metadata.

[0054] In step 506, the first node device determines a third processing model based on the data retrieval request according to the second processing model.

[0055] The third processing model is for processing data of the first vertex or an intermediate vertex, and the intermediate vertex is a vertex between the starting vertex and the first vertex. That is, the third processing model belongs to a processing model for data shard management. For example, if the starting vertex and the first vertex are connected and the search condition of the data search request is satisfied, the third processing model is used to process data of the first vertex. If there is an intermediate vertex between the starting vertex and the first vertex, the third processing model is used to process data of the intermediate vertex. Optionally, the intermediate vertex and the starting vertex are connected (there is an edge connecting the intermediate vertex and the starting vertex in the graph database).

[0056] Optionally, when the second processing model processes the data search request and obtains partial search results for the data search request, for example, by searching a path from the starting vertex to the first vertex, the first node device transmits the search results determined by the second processing model to the first processing model, and then the first node device transmits the search results to the second node device by the first processing model, and the second node device performs a process of compiling the search results for the received data search request.

[0057] The third processing model is a processing model in the first node device or a processing model in the third node device.

[0058] In step 508, the first node device transmits a data search request to the third processing model via the second processing model.

[0059] When the start vertex and the first vertex are connected and the search condition of the data search request is satisfied, that is, when the third processing model is used to process the data of the first vertex, the third processing model processes the data search request and transmits the processing result of the data search request to the second node device. When an intermediate vertex exists between the start vertex and the first vertex, the third processing model continues to determine the fourth processing model by the above method and continues to transmit the data search request to the fourth processing model until the processing result of the data search request is finally obtained, and then the fourth processing model continues to process the data search request by the above method. The fourth processing model is a model for dispatching a data search request (data processing request) in the intermediate vertex. That is, the fourth processing model belongs to a processing model for task dispatch.

[0060] Optionally, metadata including an address of a processing model for processing data of a vertex connected to the starting vertex is stored locally in the second processing model, and the first node device transmits the data search request to the third processing model based on the metadata via the second processing model.

[0061] As described above, in the method provided in this embodiment, after a node device receives a data search request, the node device can directly process the data search request according to a processing model in the node device, thereby determining a processing model for processing the vertex data related to the data search request in the graph database, and further determining the processing result of the data search request. In the process of processing the data search request, there is no need to create a thread for processing the request, and the data search request is directly processed according to the processing model, thereby avoiding the creation of a corresponding thread when processing the data search request and accelerating the data search speed.

[0062] Fig. 6 is a schematic flow diagram of another data search request processing method provided in one exemplary embodiment of the present application. This method is applied to a first node device in the system as shown in Fig. 3, and the first node device is one of multiple node devices in a distributed data storage system based on a graph database. As shown in Fig. 6, the method includes the following steps:

[0063] In step 602, the first node device receives a data search request transmitted from the second node device by a first processing model in the first node device.

[0064] The second node device is a node device in the distributed data storage system for dispatching a data processing request, the data search request being for searching data related to a first vertex in a graph database, and the data search request is accompanied by an identifier of a starting vertex.

[0065] After receiving the data search request, the second node device determines a node device in which data of the starting vertex is stored, and transmits the data search request to a processing model in the node device. The first node device stores data of the starting vertex. Optionally, each node device of the distributed data storage system stores a data shard of at least one vertex in the graph database, and a correspondence relationship between the data shard of the vertex stored in the node device and the node device is established by a global consistent hash ring. For example, the node device 1 stores a data shard of the vertex 1 in the graph database, the node device 2 stores data shards of the vertex 2 and the vertex 3 in the graph database, and the node device 3 stores a data shard of the vertex 4 in the graph database. The correspondence relationship between the node devices and the vertices in the graph database can be established by a global consistent hash ring.

[0066] For example, a fixed value can be obtained by performing a modulus operation on the identifier of the node device. Also, a fixed value can be obtained similarly by performing a modulus operation on the identifier of each vertex in the graph database. For example, all are modulus operated by 2^32. A ring can be constructed based on the result value of the modulus operation of the identifier of the node device or the identifier of the vertex by 2^32, and this ring can be called a global consistent hash ring. In terms of a clock, the circle of the clock can be understood as a circle consisting of 60 points, but in the case of a global consistent hash ring, it can be imagined as a circle consisting of 2^32 points. The value obtained by the above modulus operation is, in other words, the position of a different vertex or node device in the global consistent hash ring. Then, based on the position of the vertex in the global consistent hash ring, a node device is searched for in the clockwise direction in the global consistent hash ring, and the first node device found is determined as the node device for storing the data shard of the vertex. Note that 2^32 used in the modulus operation is merely exemplary, and other values ​​may be used. "^" represents a power; for example, "2^32" is 2 to the power of 32.

[0067] After receiving the data lookup request, the second node device can determine a second location in the global consistent hash ring based on the identifier of the starting vertex, and then determine the first node device corresponding to the data lookup request based on the second location and the location of each node device of the distributed data storage system in the global consistent hash ring. Optionally, the second node device does not store data of the vertices in the graph database.

[0068] Illustratively, FIG. 7 is a schematic diagram of a global consistent hash ring provided in one exemplary embodiment of the present application. As shown in FIG. 7, after receiving a data search request, the second node device calculates a hash value based on an identifier of a starting vertex, and determines a second position in the global consistent hash ring based on the hash value. For example, the second position is the position of vertex 701. Then, the second node device searches for the first node device in the global consistent hash ring in a clockwise direction, that is, node device 703, and determines node device 703 as the first node device. The position of the node device in the global consistent hash ring is determined by calculating a hash value of the identifier of the node device. Alternatively, the second position is the position of vertex 702. In this case, the second node device searches for the first node device in the global consistent hash ring in a clockwise direction, that is, node device 704. At this time, the second node device determines node device 704 as the first node device.

[0069] The first processing model is a model in the first node device for dispatching a data retrieval request. After receiving the data retrieval request, the first processing model sends the data retrieval request to a processing model in the first node device for processing the data of the starting vertex.

[0070] In step 604, the first node device determines a second processing model within the first node device according to the first processing model.

[0071] The second processing model is a processing model for processing data of the starting vertex in the first node device. Optionally, each data shard of the multiple vertices stored in the first node device has a corresponding processing model. After receiving the data search request, the first processing model determines the second processing model by determining a processing model for processing data of the starting vertex in the first node device based on an identifier of the starting vertex.

[0072] Optionally, the first node device stores data shards for each of the multiple vertices in the graph database, and a correspondence between each data shard and a processing model in the first node device is established by a local consistent hash ring. The first node device can determine a first position in the local consistent hash ring based on an identifier of a starting vertex by the first processing model. Then, the first node device can determine a second processing model corresponding to a data search request based on the first position and the position of each processing model in the first node device in the local consistent hash ring by the first processing model.

[0073] Illustratively, FIG. 8 is a schematic diagram of a local consistent hash ring provided in one exemplary embodiment of the present application. As shown in FIG. 8, after the first processing model receives a data search request, the first node device calculates a hash value based on an identifier of a starting vertex by the first processing model, and determines a first position in the local consistent hash ring based on the hash value. For example, the first position is the position of vertex 801. Then, the first node device searches for a first processing model, i.e., processing model 803, in the local consistent hash ring in a clockwise direction, and determines the processing model 803 as the second processing model. The position of the processing model in the local consistent hash ring is determined by calculating a hash value of an identifier of the processing model. Alternatively, the first position is the position of vertex 802. In this case, the first node device searches for a first processing model, i.e., processing model 804, in the local consistent hash ring in a clockwise direction by the first processing model. At this time, the first node device determines the processing model 804 as the second processing model by the first processing model.

[0074] In step 606, the first node device transmits a data search request to a second processing model in the first node device via the first processing model.

[0075] Optionally, the first node device sends a data search request to a message queue of a second processing model by a first processing model, where the message queue of the second processing model is for storing tasks to be processed of the second processing model. Optionally, the second processing model sequentially processes the tasks to be processed in the message queue until the processing of all the tasks to be processed is completed.

[0076] In step 608, the first node device determines, using the second processing model, a third processing model for processing the data of the first vertex based on the storage structure of the starting vertex in the graph database and the search condition in the data search request.

[0077] The storage structure of the start vertex in the graph database includes attributes of the start vertex, identifiers of vertices having a successive relationship with the start vertex, and attributes of the successive relationship. Optionally, the vertices having a successive relationship with the start vertex can be divided into those connected to the start vertex and having edges pointing to the start vertex, and those connected to the start vertex and having edges pointing to vertices connected to the start vertex. The attributes of the connection relationship include attributes of edges in different directions.

[0078] Optionally, each processing model stores resources related to its corresponding vertex, i.e., stores the memory structures of the vertices it is responsible for processing. For example, processing model 1 is used to process data of vertex 1 and vertex 4, processing model 2 is used to process data of vertex 2, and processing model 3 is used to process data of vertex 3. Processing model 1 stores the memory structure of vertex 1 and the memory structure of vertex 4, processing model 2 stores the memory structure of vertex 2, and processing model 3 stores the memory structure of vertex 3.

[0079] Illustratively, Fig. 9 is a schematic diagram of a storage structure of vertices in a graph database provided in one exemplary embodiment of the present application. As shown in Fig. 9(a), in the graph database, vertex 1 is connected to each of vertex 2, vertex 3, and vertex 4. Here, the edge between vertex 1 and vertex 2 points to vertex 1, the edge between vertex 1 and vertex 3 points to vertex 1 and vertex 3, respectively, and the edge between vertex 1 and vertex 4 points to vertex 4. Vertex 2 is connected to vertex 3, and the edge between vertex 2 and vertex 3 points to vertex 2. Vertex 2 is connected to vertex 4, and the edge between vertex 2 and vertex 4 points to vertex 4. Vertex 3 is connected to vertex 4, and the edge between vertex 3 and vertex 4 points to vertex 3.

[0080] As shown in FIG. 9B, processing model 1 is used to process data of vertex 1 and vertex 4, processing model 2 is used to process data of vertex 2, and processing model 3 is used to process data of vertex 3. Processing model 1 stores the memory structures of vertex 1 and vertex 4. Here, Vp1 represents the attribute of vertex 1 (V1), V1_in includes the identifier of a vertex whose edge points to vertex 1 among the vertices connected to vertex 1, and the attribute of the edge between the vertex and vertex 1, for example, Ep3 represents the attribute of the edge between vertex 1 and vertex 3 that points to vertex 1, and Ep2 represents the attribute of the edge between vertex 1 and vertex 2 that points to vertex 1. V1_out includes the identifier of a vertex whose edge points to the connected vertex among the vertices connected to vertex 1, and the attribute of the edge between the vertex and vertex 1. Vp4 represents an attribute of vertex 4 (V4), V4_in includes an identifier of a vertex whose edge points to vertex 4 among the vertices connected to vertex 4, and V4_out includes an identifier of a vertex whose edge points to the connected vertex among the vertices connected to vertex 4, and an attribute of the edge between the vertex and vertex 4. The processing model 2 stores a memory structure of vertex 2. Here, Vp2 represents an attribute of vertex 2 (V2), V2_in includes an identifier of a vertex whose edge points to vertex 2 among the vertices connected to vertex 2, and V2_out includes an identifier of a vertex whose edge points to the connected vertex among the vertices connected to vertex 2, and an attribute of the edge between the vertex and vertex 2. The processing model 3 stores a memory structure of vertex 3. Here, Vp3 represents the attributes of vertex 3 (V3), V3_in includes the identifier of a vertex connected to vertex 3 whose edge points to vertex 3, and V3_out includes the identifier of a vertex connected to vertex 3 whose edge points to the connected vertex and the attribute of the edge between the vertex and vertex 3.

[0081] The first node device searches for a vertex that satisfies the search condition from among the vertices that have a consecutive relationship with the starting vertex based on one or more of the attribute of the starting vertex, the vertices that have a consecutive relationship with the starting vertex, and the attribute of the consecutive relationship, using the second processing model. In response to the presence of a vertex that satisfies the search condition among the vertices that have a consecutive relationship with the starting vertex, the first node device determines the vertex that satisfies the search condition as a first vertex using the second processing model. Thereafter, the first node device can determine a third processing model for processing data of the first vertex using the second processing model. Optionally, the second processing model stores information of a processing model for processing data of a vertex connected thereto, and after determining the first vertex connected thereto, the third processing model can be determined based on the information.

[0082] 9, for example, the data search request is used to search for data of a vertex whose edge points to vertex 1 and whose attribute satisfies a search condition among the first-order neighborhoods of vertex 1. The first node device can determine vertex 2 as the first vertex based on the memory structure of vertex 1 through the second processing model, and then determine a third processing model for processing the data of vertex 2.

[0083] In step 610, the first node device determines, using the second processing model, a third processing model for processing the data of the intermediate vertex based on the storage structure of the starting vertex in the graph database and the search condition in the data search request.

[0084] The intermediate vertex is a vertex between the starting vertex and the first vertex. When the first node device searches the storage structure of the starting vertex by the method in step 608 using the second processing model, in response to the absence of a vertex that satisfies the search condition among the vertices that have a continuous relationship with the starting vertex, the first node device determines the vertex that has a continuous relationship with the starting vertex as an intermediate vertex by the second processing model, for example, determines all vertices that have a continuous relationship with the starting vertex as intermediate vertices. Thereafter, the first node device determines a third processing model for processing data of the intermediate vertex by the second processing model. Optionally, the second processing model stores information of a processing model for processing data of a vertex connected thereto, and after determining the intermediate vertex connected thereto, the third processing model can be determined based on the information.

[0085] 9, for example, the data search request is used to search the second-order neighborhood of vertex 1. If vertex 1 and vertex 4 are not directly connected, the first node device only searches for connected vertices 2 and 3 based on the memory structure of vertex 1 through the second processing model. In this case, vertex 2 and vertex 3 are determined as intermediate vertices, and both the processing model for processing data of vertex 2 and the processing model for processing data of vertex 3 are determined as the third processing model.

[0086] In step 612, the first node device transmits a data search request to the third processing model via the second processing model.

[0087] If the third processing model belongs to the first node device, the first node device transmits a data search request to the third processing model in the first node device according to the second processing model.

[0088] When the third processing model belongs to the third node device, the first node device transmits a data search request to a fourth processing model in the third node device through the second processing model, and the fourth processing model is a processing model in the third node device for dispatching the data search request, where the third node device transmits the data search request to the third processing model through the fourth processing model, so that the third processing model can receive the data search request.

[0089] Optionally, when transmitting a data search request across node devices, for example, when a first node device transmits a data search request to a fourth processing model in a third node device by a second processing model, the data search request is transmitted to the fourth processing model based on a remote procedure call (RPC). The first node device transmits the data search request to a message queue of the third processing model (fourth processing model) by the second processing model.

[0090] When the start vertex and the first vertex are connected and the search condition of the data search request is satisfied, i.e., when the third processing model is used to process the data of the first vertex, the third processing model processes the data search request and transmits the processing result of the data search request to the second node device. When an intermediate vertex exists between the start vertex and the first vertex, i.e., when the third processing model is used to process the data of the intermediate vertex, the third processing model successively determines the fourth processing model according to the above method and successively transmits the data search request to the fourth processing model until the processing result of the data search request is finally obtained, and then the fourth processing model successively processes the data search request according to the above method.

[0091] Exemplarily, when the third processing model determines the processing result of the data search request, i.e., when the third processing model is used to process the data of the first vertex, the node device where the third processing model is located transmits the processing result of the data search request to a processing model for dispatching the data search request in the node device where the third processing model is located by the third processing model, and then the node device where the third processing model is located further transmits the processing result of the data search request to the second node device by the processing model for dispatching the data search request.

[0092] Optionally, there are a plurality of processing models that transmit the processing results of the data search request to the second node device. For example, to search for a path from the starting vertex to the first vertex, it is necessary to search for a plurality of edges and feed them back to the second node device. In this case, the second node device finally obtains the processing results of the data search request by performing a process of consolidating the processing results of the data search request transmitted from all the processing models.

[0093] Optionally, multiple threads of each processor core of the first node device are bound one-to-one to multiple processing models in the first node device, thereby realizing independent execution of the processing models for each thread, and avoiding resource robbery between the processing models. All of the processing models in the distributed data storage system are actor models, for example, the above first processing model, second processing model, third processing model, and fourth processing model are actor models. Optionally, the first node device includes multiple hard disks for storing data, and the hard disks have a binding relationship with the processor core of the second node device. The above binding relationship allows each actor model to have its own storage and calculation resources, and realizes unlocked processing in the data processing process.

[0094] It should be noted that the method provided in the embodiment of the present application can be used for processing data processing requests in addition to processing data search requests. When processing a data processing request, the second node device determines a node device in which data of the vertex is stored based on an identifier of the vertex attached to the data processing request, and transmits the data processing request to a processing model in the node device for dispatching a data processing request (data search request). The processing model then transmits the data processing request to a processing model in the node device for processing data of the vertex, and the processing model in charge of processing the data of the vertex processes the data processing request. In this scenario, data of the vertices in the graph database is sharded and stored in the node device, and different processing models in the node device are responsible for data processing of the data of the different vertices. When processing data of different vertices in parallel, it is possible to realize data processing of different vertices in parallel by different processing models without locking the database, thereby increasing the efficiency of data processing. When a data structure based on a B-tree or a B+tree is used, locks of different granularities on the trees can be avoided, thereby increasing the efficiency of data processing.

[0095] In response to the problem of traditional data modification synchronization lock, the method provided in the embodiment of the present application adjusts data shards based on the characteristics of current multi-core processors, manages each data shard by an independent actor model, and does not bind the actor model to a specific central processing unit (CPU) core. In this way, the logical shards in each conventional node device are converted into real physical shards, and data can be managed with finer granularity and more suitable for highly parallel modification scenarios. In addition, such a fine-grained management method can be more flexible in terms of scalability, and is particularly suitable for the expansion needs on the cloud.

[0096] As described above, in the method provided in this embodiment, after a node device receives a data search request, the node device can directly process the data search request according to a processing model in the node device, determine a processing model for processing the vertex data related to the data search request in the graph database, and further realize the determination of the processing result of the data search request. In the process of processing the data search request, there is no need to create a thread for processing the request, and the data search request is directly processed according to the processing model, thereby avoiding the creation of a corresponding thread when processing the data search request and accelerating the data search speed.

[0097] In addition, in the method provided in this embodiment, a simple method for determining the third processing model is provided by determining the third processing model based on the storage structure of the starting vertex in the graph database.

[0098] In addition, in the method provided in this embodiment, when a first vertex is determined based on a memory structure, a third processing model for processing the data of the first vertex is determined to process a data retrieval request, so that the data retrieval request can be quickly assigned to the corresponding processing model for processing.

[0099] In addition, in the method provided in this embodiment, when an intermediate vertex is determined based on a memory structure, a third processing model for processing data of the intermediate vertex is determined to process a data retrieval request, thereby realizing that the third processing model continues to search for the first vertex based on the data retrieval request.

[0100] In addition, in the method provided in this embodiment, a local consistent hash ring is used to maintain the vertices that each processing model in each node device is responsible for, so that the number of processing models in each node device can be dynamically increased or decreased without affecting the relationship between the processing models and their corresponding vertices.

[0101] In addition, in the method provided in this embodiment, a data search request is sent to a message queue of a processing model, and when a message is present in the message queue of the processing model, the request is processed immediately, thereby improving the efficiency of request processing.

[0102] In addition, the method provided in this embodiment provides a scheme for dispatching a data retrieval request by sending the data retrieval request to a processing model for dispatching the request in a different node device.

[0103] In addition, the method provided in this embodiment provides a scheme for transmitting a data search request by transmitting the data search request to a processing model in a different node device.

[0104] Furthermore, in the method provided in this embodiment, when transmitting across node devices, the RPC method is used for transmission, thereby providing a method for transmitting a data search request across node devices.

[0105] In addition, in the method provided in this embodiment, in the data search process, the method of directly pushing the search results after determining the search results realizes a higher performance search than the conventional method of sequentially pulling the search results.

[0106] Furthermore, the method provided in this embodiment provides a method for efficiently determining data search results by consolidating the processing results of data search requests.

[0107] In addition, in the method provided in this embodiment, the processing model can be bound to the threads of the processor core, and the processing model can be expanded or contracted at the true processor core level by increasing or decreasing the number of processor cores.

[0108] In addition, the method provided in this embodiment provides a model for processing data retrieval requests by processing the data retrieval requests through an actor model.

[0109] In addition, the order of steps of the method provided in the embodiments of the present application may be adjusted as appropriate, and the number of steps may be increased or decreased according to circumstances. Any method that a person skilled in the art can easily think of modifying within the technical scope set forth in the present application should be included in the scope of protection of the present application, and therefore further description will be omitted.

[0110] In one specific example, still referring to FIG. 4, the data search request is for searching the second-order neighbors of vertex m (id=$id). The data search request can be expressed as Match(m)-[*2]->(n) where id(m)=$id return n. When a user sends a data search request to a search layer (i.e., a second node device) of a distributed data storage system, the search layer sends the data search request to a node device where the data of vertex m is stored based on a global data shard arrangement (global consistent hash ring). When the node device receives the data search request, the node device sends the data search request to a message queue of an actor model responsible for processing the data of vertex m based on a local data shard rule (local consistent hash ring) by an actor model for dispatching. After receiving a message (data search request), the actor model analyzes the message and executes an instruction according to the message instruction. For example, in this case, the actor model obtains all the first-order neighbors of vertex m, and then calls the vertex dispatch (DispatchVertex) function to assign a message to the actor model related to the first-order neighbors, and sends the data related to the second-order neighbors searched by the actor model to the message queue of the actor model in charge of search assignment in the node device, and feeds back the partial search result to the search layer. At the same time, after receiving the message, other actor models immediately perform a search locally based on the information in the message after receiving the message, and return the search result to the actor model in charge of search assignment in the node device where it is located, and feed back the partial search result to the search layer. When sending a message across node devices, the message is transmitted to the actor model in charge of search assignment in the node device, and the actor model in charge of search assignment performs the above-mentioned assignment operation, and finally returns the data search result to the search layer. The above search process can be likened to Map(Actors[DestList()],GetDestList()), whose meaning is as follows:The method for getting the target list (GetDestList()) is applied to the related actor models simultaneously, and then all the returned data are aggregated to finally obtain the processing result of the data retrieval request.

[0111] 10 is a schematic diagram of a configuration of a data search request processing device provided in one exemplary embodiment of the present application. The device is one of a plurality of node devices in a distributed data storage system based on a graph database. As shown in FIG. 10, the device: a receiving module 1001 that receives a data search request transmitted from a second node device by a first processing model in the device, the data search request being for searching data related to a first vertex in the graph database, the data search request being assigned an identifier of a starting vertex, and data of the starting vertex being stored in the device; and a sending module 1002 for sending the data search request by the first processing model to a second processing model in the device, the second processing model being a processing model for processing data of the starting vertex in the device; a determining module 1003 for determining a third processing model based on the data retrieval request by the second processing model, the third processing model being for processing data of the first vertex or an intermediate vertex, the intermediate vertex being a vertex between the starting vertex and the first vertex; The sending module 1002 sends the data retrieval request via the second processing model to the third processing model.

[0112] In one optional design, the determination module 1003: The third processing model is determined based on a storage structure of the starting vertex in the graph database and a search condition in the data search request using the second processing model.

[0113] In one optional design, the storage structure of the starting vertex in the graph database includes an attribute of the starting vertex, an identifier of a vertex that has a successive relationship with the starting vertex, and an attribute of the successive relationship, and the determining module 1003: using the second processing model, searching for a vertex that satisfies the search condition from among the vertices that have the successive relationship with the starting vertex based on one or more of an attribute of the starting vertex, the vertices that have the successive relationship with the starting vertex, and an attribute of the successive relationship; in response to a vertex that satisfies the search condition being present among the vertices that have the successive relationship with the starting vertex, determining, by the second processing model, the vertex that satisfies the search condition as the first vertex; The third processing model for processing the data of the first vertex is determined by the second processing model.

[0114] In one optional design, the storage structure of the starting vertex in the graph database includes an attribute of the starting vertex, an identifier of a vertex that has a successive relationship with the starting vertex, and an attribute of the successive relationship, and the determining module 1003: using the second processing model, searching for a vertex that satisfies the search condition from among the vertices that have the successive relationship with the starting vertex based on one or more of an attribute of the starting vertex, the vertices that have the successive relationship with the starting vertex, and an attribute of the successive relationship; in response to a fact that no vertex that satisfies the search condition is present among the vertices that have the successive relationship with the starting vertex, determining, by the second processing model, the vertex that has the successive relationship with the starting vertex as the intermediate vertex; The third processing model for processing the data of the intermediate vertices is determined by the second processing model.

[0115] In one optional design, the device stores a data shard for each of a plurality of vertices in the graph database, and a correspondence between each of the data shards and a processing model in the device is established by a local consistent hash ring. determining, by the first processing model, a first location in the local consistent hash ring based on an identifier of the starting vertex; The first processing model determines the second processing model corresponding to the data retrieval request based on the first location and a location in the local consistent hash ring of each processing model in the device.

[0116] In one optional design, the transmitting module 1002 includes: sending, by the first processing model, the data retrieval request to a message queue of the second processing model; The message queue of the second processing model is for storing tasks to be processed of the second processing model.

[0117] In one optional design, the third processing model belongs to the device, and the sending module 1002 is The data retrieval request is sent by the second processing model to the third processing model within the device.

[0118] In one optional design, the third processing model belongs to a third node device, and the sending module 1002 further comprises: transmitting the data search request to a fourth processing model in the third node device by the second processing model, the fourth processing model being a processing model in the third node device for dispatching the data search request; The third node device is for transmitting the data search request to the third processing model by the fourth processing model.

[0119] In one optional design, the transmitting module 1002 includes: The second processing model sends the data retrieval request to the fourth processing model based on an RPC.

[0120] In one optional design, when the third processing model determines a processing result of the data search request, a node device in which the third processing model is located transmits, by the third processing model, the processing result of the data search request to a processing model for dispatching the data search request in the node device in which the third processing model is located; The node device in which the third processing model is located further transmits a processing result of the data search request to the second node device by a processing model for dispatching the data search request within the node device in which the third processing model is located.

[0121] In one optional design, the third processing model determines a result of processing the data retrieval request; When the third processing model belongs to the first node device, the first node device transmits a processing result of the data search request to the first processing model by the third processing model, the first processing model being a processing model for dispatching the data search request in the first node device, and the first node device further transmits a processing result of the data search request to the second node device by the first processing model; When the third processing model belongs to a third node equipment, the third node equipment transmits a processing result of the data search request to a fourth processing model via the third processing model, the fourth processing model being a processing model for dispatching the data search request within the third node equipment, and the third node equipment further transmits the processing result of the data search request to the second node equipment via the fourth processing model.

[0122] In one optional design, there are a plurality of processing models for transmitting processing results of the data search request to the second node device; The second node device performs a process of compiling the processing results of the data search requests transmitted from the at least two processing models.

[0123] In one optional design, multiple threads of each processor core of the device are bound one-to-one to multiple processing models within the device.

[0124] In one optional design, the first processing model, the second processing model, and the third processing model are actor models.

[0125] According to an embodiment of the present application, a computer device is provided having a processor and a memory, and at least one program is stored in the memory, which, when loaded and executed by the processor, realizes the method of processing a data search request provided in each of the embodiments of the above methods.

[0126] Optionally, the computer device is a server. Illustratively, Figure 11 is a schematic diagram of a configuration of a computer device provided in one exemplary embodiment of the present application.

[0127] The computing device 1100 includes a central processing unit (CPU) 1101, a system memory 1104 including a random access memory (RAM) 1102 and a read-only memory (ROM) 1103, and a system bus 1105 connecting the system memory 1104 and the central processing device 1101. The computing device 1100 further includes a basic input / output system (I / O system) 1106 that supports information transfer between components within the computing device, and a mass storage device 1107 for storing an operating system 1113, applications 1114, and other program modules 1115.

[0128] The basic input / output system 1106 includes a display 1108 for displaying information and an input device 1109, such as a mouse or keyboard, for a user to input information, where both the display 1108 and the input device 1109 are connected to the central processing unit 1101 via an input / output controller 1110 connected to the system bus 1105. The basic input / output system 1106 may further include an input / output controller 1110 for accepting and processing input from a number of other devices, such as a keyboard, a mouse, or an electronic touch pen. Similarly, the input / output controller 1110 may also provide output to a display, a printer, or other type of output device.

[0129] The mass storage device 1107 is connected to the central processing unit 1101 via a mass storage controller (not shown) that is connected to the system bus 1105. The mass storage device 1107 and its associated computer readable storage media provide non-volatile storage for the computing device 1100. That is, the mass storage device 1107 may include a computer readable storage medium (not shown), such as a hard disk or a Compact Disc Read-Only Memory (CD-ROM) drive.

[0130] Without loss of generality, the computer readable storage media may include computer storage media and communication media. Computer storage media includes volatile and non-volatile, removable and non-removable media implemented by any method or technology for storing information such as computer readable instructions, data structures, program modules, or other data. Computer storage media includes RAM, ROM, erasable programmable read only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other solid state storage, CD-ROM, digital versatile disc (DVD) or other optical storage, tape cartridge, magnetic tape, magnetic disk storage or other magnetic storage. Of course, those skilled in the art will appreciate that the computer storage media is not limited thereto. The above system memory 1104 and mass storage device 1107 may be collectively referred to as memory.

[0131] The memory stores one or more programs, the one or more programs configured to be executed by one or more central processing units 1101, the one or more programs including instructions for implementing the above method embodiments, and the central processing unit 1101 executes the one or more programs to implement the methods provided in each of the above method embodiments.

[0132] According to various embodiments of the present application, the computer device 1100 may be implemented by connecting to a remote computer device on a network, such as the Internet, via a network. That is, the computer device 1100 may be connected to a network 1112 via a network interface unit 1111 connected to the system bus 1105, or in other words, the computer device 1100 may be connected to another type of network or remote computer device system (not shown) using the network interface unit 1111.

[0133] The memory further includes one or more programs stored in the memory, the one or more programs including those for performing steps executed by a computer device in the methods provided in the embodiments of the present application.

[0134] In an embodiment of the present application, a computer-readable storage medium storing at least one program is further provided, which, when loaded and executed by a processor of a computer device, realizes the data retrieval request processing method provided in each of the above method embodiments.

[0135] The present application further provides a computer program product or computer program comprising computer instructions, the computer instructions being stored in a computer readable storage medium, a processor of a computing device reading the computer instructions from the computer readable storage medium, and, when executed by the processor, causing the computing device to perform the method for processing a data retrieval request provided in each of the above method embodiments.

Claims

1. A method for processing a data search request, executed by a first node device, the first node device being one of at least two node devices in a distributed data storage system based on a graph database, the method for processing the data search request comprising: a step of receiving, by the first node device, a data search request transmitted from a second node device by a first processing model in the first node device, the data search request being for searching data related to a first vertex in the graph database, the data search request being assigned an identifier of a starting vertex, and data of the starting vertex being stored in the first node device; a step of the first node device transmitting the data search request to a second processing model in the first node device by the first processing model, the second processing model being a processing model for processing data of the starting vertex in the first node device; a step of the first node device determining a third processing model based on the data search request by the second processing model, the third processing model being for processing data of the first vertex or an intermediate vertex, the intermediate vertex being a vertex between the starting vertex and the first vertex; The first node device transmits the data search request to the third processing model via the second processing model. How to handle data retrieval requests.

2. The step of determining a third processing model based on the data search request by the first node device using the second processing model includes: The first node device determines the third processing model based on a storage structure of the starting vertex in the graph database and a search condition in the data search request by using the second processing model.

2. The method of claim 1, wherein the data retrieval request is processed by a processor.

3. a storage structure for the starting vertex in the graph database includes an attribute of the starting vertex, an identifier of a vertex having a successive relationship with the starting vertex, and an attribute of the successive relationship; The step of the first node device determining the third processing model based on a storage structure of the starting vertex in the graph database and a search condition in the data search request using the second processing model, a step in which the first node device searches for a vertex that satisfies the search condition from among the vertices that have the successive relationship with the starting vertex, based on one or at least two of an attribute of the starting vertex, the vertices that have the successive relationship with the starting vertex, and an attribute of the successive relationship, using the second processing model; in response to a vertex that satisfies the search condition being present among the vertices that have the successive relationship with the starting vertex, the first node device determines, by the second processing model, the vertex that satisfies the search condition as the first vertex; and determining, by the first node device, the third processing model for processing the data of the first vertex according to the second processing model.

3. A method for processing a data retrieval request according to claim 2.

4. a storage structure for the starting vertex in the graph database includes an attribute of the starting vertex, an identifier of a vertex having a successive relationship with the starting vertex, and an attribute of the successive relationship; The step of the first node device determining the third processing model based on a storage structure of the starting vertex in the graph database and a search condition in the data search request using the second processing model, a step in which the first node device searches for a vertex that satisfies the search condition from among the vertices that have the successive relationship with the starting vertex, based on one or at least two of an attribute of the starting vertex, the vertices that have the successive relationship with the starting vertex, and an attribute of the successive relationship, using the second processing model; in response to a fact that no vertex that satisfies the search condition is present among the vertices that have the successive relationship with the starting vertex, the first node device determines, by the second processing model, the vertex that has the successive relationship with the starting vertex as the intermediate vertex; and determining, by the first node device, the third processing model for processing the data of the intermediate vertex according to the second processing model.

3. A method for processing a data retrieval request according to claim 2.

5. The first node device stores a data shard for each of at least two vertices in the graph database, and a correspondence between each of the data shards and a processing model in the first node device is established by a local consistent hash ring; The method for processing the data search request includes: determining, by the first node device, a first position in the local consistent hash ring based on an identifier of the starting vertex according to the first processing model; The method further includes a step of determining, by the first node device, the second processing model corresponding to the data retrieval request based on the first location and a location in the local consistent hash ring of each processing model in the first node device, using the first processing model.

2. The method of claim 1, wherein the data retrieval request is processed by a processor.

6. The step of transmitting the data search request to a second processing model in the first node device by the first processing model, The first node device transmits the data retrieval request to a message queue of the second processing model by the first processing model; the message queue of the second processing model is for storing tasks to be processed of the second processing model; 2. The method of claim 1, wherein the data retrieval request is processed by a processor.

7. the third processing model belongs to the first node device; The step of transmitting the data search request to the third processing model by the first node device through the second processing model includes: The first node device transmits the data search request to the third processing model in the first node device by the second processing model.

2. The method of claim 1, wherein the data retrieval request is processed by a processor.

8. the third processing model belongs to a third node device; The step of transmitting the data search request to the third processing model by the first node device through the second processing model includes: a step of the first node device transmitting the data search request to a fourth processing model in the third node device by the second processing model, the fourth processing model being a processing model for dispatching the data search request in the third node device; the third node device is for transmitting the data search request to the third processing model by the fourth processing model; 2. The method of claim 1, wherein the data retrieval request is processed by a processor.

9. The step of transmitting the data search request by the first node device to a fourth processing model in the third node device by the second processing model, the first node device transmitting the data retrieval request to the fourth processing model based on a remote procedure call via the second processing model; 9. A method for processing a data retrieval request according to claim 8.

10. the third processing model determines a result of processing the data retrieval request; When the third processing model belongs to the first node device, the first node device transmits a processing result of the data search request to the first processing model by the third processing model, the first processing model being a processing model for dispatching the data search request in the first node device, and the first node device further transmits a processing result of the data search request to the second node device by the first processing model; When the third processing model belongs to a third node device, the third node device transmits a processing result of the data search request to a fourth processing model by the third processing model, the fourth processing model being a processing model for dispatching the data search request within the third node device, and the third node device further transmits the processing result of the data search request to the second node device by the fourth processing model.

2. The method of claim 1, wherein the data retrieval request is processed by a processor.

11. There are at least two processing models for transmitting a processing result of the data search request to the second node device, the second node device performs a process of consolidating the processing results of the data search requests transmitted from the at least two processing models; A method for processing a data retrieval request according to claim 10.

12. At least two threads of each processor core of the first node device are bound one-to-one to at least two processing models within the first node device; 2. The method of claim 1, wherein the data retrieval request is processed by a processor.

13. the first process model, the second process model, and the third process model are actor models; 2. The method of claim 1, wherein the data retrieval request is processed by a processor.

14. A data search request processing device, the data search request processing device being one of at least two node devices in a distributed data storage system based on a graph database; a receiving module that receives a data search request transmitted from a second node device by a first processing model in the data search request processing device, the data search request being for searching data related to a first vertex in the graph database, the data search request being assigned an identifier of a starting vertex, and data of the starting vertex being stored in the data search request processing device; a sending module that sends the data retrieval request by the first processing model to a second processing model in a processing device of the data retrieval request, the second processing model being a processing model for processing data of the starting vertex in the processing device of the data retrieval request; a determining module for determining a third processing model based on the data retrieval request by the second processing model, the third processing model being for processing data of the first vertex or an intermediate vertex, the intermediate vertex being a vertex between the starting vertex and the first vertex; the transmission module transmits the data retrieval request to the third processing model via the second processing model; The processor of data retrieval requests.

15. A computer device comprising a processor and a memory, wherein the memory stores at least one program, and the at least one program, when loaded and executed by the processor, realizes a method for processing a data search request according to any one of claims 1 to 13.

16. A computer program causing a computer to execute the method for processing a data search request according to any one of claims 1 to 13.

Citation Information

Patent Citations

  • Duplication elimination in depth based searches for distributed systems

    US20220179859A1

  • Network graph generation method and decision-making assistance system

    WO2014083655A1