Incremental query method and device for graph data

By using the incremental query method in the graph query, the incremental graph data is obtained and N-hop expansion is performed to form an expanded sub-graph, which solves the efficiency problem of processing dynamic changing graph data in the graph query and achieves efficient graph query performance.

CN120179875APending Publication Date: 2025-06-20ALIPAY (HANGZHOU) INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510237889.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-20

AI Technical Summary

Technical Problem

During the execution of graph query, how to efficiently process massive and dynamically changing graph data to avoid wasting computing resources and slow query speed caused by full calculations.

Method used

The incremental query method is adopted to obtain the incremental edges and nodes in the incremental graph data, and N-hop neighbor expansion is performed to form an expanded subgraph, and the target graph query is performed in the expanded subgraph to avoid full graph calculation.

Benefits of technology

Effectively prevent computing resources from wasting on duplicate node and path searches, improve the execution performance of graph queries, and meet business scenarios with high real-time requirements.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120179875A_ABST
    Figure CN120179875A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides an incremental query method for graph data, which comprises the following steps of: determining a target graph query which relates to an N-hop query; obtaining incremental graph data, wherein the incremental graph data comprises a plurality of incremental edges and incremental nodes connected with the incremental edges; and performing N-hop neighbor expansion by taking the incremental node as a starting point to obtain an expanded sub-graph. And in the expanded sub-graph, executing matching for the target graph query to obtain a query result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] One or more embodiments of this specification relate to the technical field of computer data query, and in particular to an incremental query method and device for graph data. Background Art

[0002] With the development of big data and artificial intelligence, graph-structured data (hereinafter simply referred to as "graph data") is used to record and process business data in more and more scenarios due to its efficient expression ability for complex relationships. For example, in social platforms, graph data is used to depict the social relationships between users; in payment platforms, graph data is used to construct complex payment relationship graphs. In order to fully utilize the potential of graph data, the industry has designed graph data query languages, which can execute graph query statements with the help of graph databases to achieve querying and analyzing of graph data.

[0003] In practical applications, graph query faces a severe challenge: how to efficiently process massive and dynamically changing graph data during the execution of graph query. On the one hand, the graph data constructed based on specific business scenarios is not static, but continuously updated with the continuous generation of business data. On the other hand, large-scale graph data is difficult to be completely processed in one graph query, and usually needs to be reasonably divided and dynamically queried in batches. During the execution of graph query, the dynamic changes of graph data make it difficult to guarantee the real-time performance and accuracy of graph query results.

[0004] Even more troublesome is that when the graph data changes, existing graph databases usually need to re-execute graph query statements on the updated graph data. This traditional full-scale calculation and processing method not only causes waste of computing resources, but also slows down the execution speed of graph query, and it is difficult to meet business scenarios with high requirements for real-time performance.

[0005] Therefore, in the process of executing queries on dynamic graph data using graph query languages, how to improve the execution efficiency of graph query is a technical problem that needs to be solved currently. Summary of the Invention

[0006] One or more embodiments of this specification describe an incremental query method and device for graph data. By expanding subgraphs based on incremental graph data and executing graph queries in the subgraphs, it is possible to avoid full-scale graph calculations caused by incremental graph data, effectively prevent waste of computing resources in repeated node and path searches, thereby improving the execution performance of graph query and solving the above technical problems.

[0007] According to a first aspect, an incremental query method for graph data is provided, including:

[0008] Determine a target graph query, which involves N-hop queries.

[0009] Obtain incremental graph data, which includes a number of incremental edges and the incremental nodes connected by the incremental edges.

[0010] Starting from the incremental nodes, perform N-hop neighbor expansion to obtain an extended subgraph.

[0011] In the extended subgraph, perform matching for the target graph query to obtain a query result.

[0012] According to one implementation, the target graph query does not include a constraint condition that limits the starting node to a uniquely specified node.

[0013] According to one implementation, the value of N is greater than or equal to 2 and less than a preset first threshold.

[0014] According to one implementation, the incremental graph data is the latest batch of graph data streamed into the graph database.

[0015] According to one implementation, the graph database storing the graph data obtains the incremental data corresponding to the time window at every preset time window; the incremental graph data is the incremental data of the most recent time window.

[0016] According to one implementation, the performing N-hop neighbor expansion to obtain an extended subgraph includes:

[0017] Execute N rounds of message propagation. A single round of message propagation includes propagating a first message from a first node to its neighbor nodes; wherein, in the first round of message propagation, the incremental node is used as the first node, and in non-first-round message propagation, the node that received the first message in the previous round is used as the first node.

[0018] Include the nodes and edges involved in the N rounds of message propagation in the extended subgraph.

[0019] In a scenario of the above implementation, the first message carries the sender id; the non-first-round message propagation includes:

[0020] Determine the target neighbor nodes of any first node. The target neighbor nodes are the neighbor nodes of the first node and the node id is different from the sender id carried in the first message received by the first node.

[0021] Send a first message from the first node to the target neighbor nodes, where the node id of the first node is used as the sender id.

[0022] According to one implementation, the performing matching for the target graph query to obtain a query result includes:

[0023] Perform N-hop traversal in the extended subgraph according to the query conditions in the target graph query.

[0024] Determine the query result according to the result of the N-hop traversal.

[0025] In a scenario of the above implementation, the N-hop traversal includes adding a first mark in the single step of the traversal in response to traversing an incremental edge.

[0026] The determination of the query result includes screening out the paths containing the first mark from the traversal paths that meet the query conditions as the result paths.

[0027] In an example of the above scenario, after adding the first mark, the method further includes:

[0028] Keep recording the first mark in each subsequent single step of the traversal.

[0029] According to a second aspect, there is provided an incremental query device for graph data, the device includes:

[0030] A determination module configured to determine a target graph query, which involves an N-hop query.

[0031] An acquisition module configured to acquire incremental graph data, which includes a number of incremental edges and incremental nodes connected by the incremental edges.

[0032] An expansion module configured to perform N-hop neighbor expansion starting from the incremental nodes to obtain an expanded subgraph.

[0033] An execution module configured to perform matching for the target graph query in the expanded subgraph to obtain a query result.

[0034] According to a third aspect, there is provided a computer program product including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the steps of the method described in the first aspect are implemented.

[0035] According to a fourth aspect, there is provided a computing device including a memory and a processor, characterized in that an executable code is stored in the memory, and when the processor executes the executable code, the method described in the first aspect is implemented.

[0036] In summary, by using the above methods and devices disclosed in the embodiments of this specification, after obtaining incremental graph data, an expanded subgraph that may generate an incremental query result can be determined according to the incremental edges and incremental nodes. When performing a graph query, only perform graph matching that meets the query conditions in the expanded subgraph, so as to implement incremental query of graph data, avoid full-scale graph calculation caused by incremental graph data, and thus improve the execution performance of graph queries. Description of the Drawings

[0037] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following will briefly introduce the accompanying drawings required for the description of the embodiments. Obviously, the accompanying drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other accompanying drawings can be obtained based on these drawings.

[0038] Figure 1 An exemplary graph data disclosed in this specification;

[0039] Figure 2 An execution schematic diagram of graph query in a typical streaming computing scenario disclosed in this specification;

[0040] Figure 3 A method architecture diagram for performing incremental query on graph data disclosed in this specification;

[0041] Figure 4 A flowchart of an incremental query method for graph data provided according to an embodiment of this specification;

[0042] Figure 5 A schematic diagram of streaming to obtain incremental graph data provided according to an embodiment of this specification;

[0043] Figure 6 An example of point-edge determination of incremental graph data provided according to an embodiment of this specification;

[0044] Figure 7 A schematic diagram of extended subgraph extraction provided according to an embodiment of this specification;

[0045] Figure 8 A schematic diagram of an incremental query device for graph data provided according to an embodiment of this specification. Detailed implementation manners

[0046] The following will describe the solutions provided in the embodiments of this specification in conjunction with the accompanying drawings.

[0047] Graph Query Language (GQL) can be used to standardize the writing of graph query statements for performing retrievals on graph data. Currently, common graph query languages include: Cypher, Gremlin, SPARQL, etc. The syntax and functions of each language are different, but they can all process graph data and complete graph query tasks. In order to unify graph query languages, the ISO / IEC standardization organization has developed the GQL standard (ISO / IEC 39075:2024) as the standard graph query language specification.

[0048] In one or more embodiments of this specification, taking a graph query written based on GQL as an example, the problem of duplicate calculation that may occur when facing dynamic graph data during the execution of a graph query will be elaborated, and a technical solution to solve this problem will be elaborated. It should be noted that although in some embodiments of this specification, the GQL graph query language specification is used for description, it does not represent a limitation on the application scenarios and technical tools of the embodiments of the present invention. The technical concept embodied in each embodiment of this specification can be applied to other graph databases that support other graph query languages.

[0049] A graph database is a new type of database implemented based on graph theory. It mainly includes a data storage module for storing graph data and a graph query engine for performing graph queries on graph data. The atomic objects for operations in a graph database are the basic elements of a graph in graph theory: nodes and connecting edges.

[0050] Figure 1 An exemplary graph data is disclosed. For the sake of simplicity in description, in this example, the labels and attributes on the nodes and connecting edges are omitted, and each node in the graph is numbered with a letter. In the following description, the connecting edge between nodes is represented as <source node, destination node>, and the direction of the connecting edge points from the source node to the destination node.

[0051] Referring to the accompanying drawings, the graph data consists of nodes (shown as circles) and connecting edges (shown as arrow lines). On the graph data, graph query statements can be executed to perform operations such as data mining, relationship analysis, and path finding. A graph query statement is a query statement written in accordance with a specific graph query language specification. In one example, the graph query statement can be written using the GQL standard graph query language, and the example is as follows:

[0052] MATCH p=(N1)-[E1]->(N2) RETURN p;

[0053] Among them, "()" is the identification symbol corresponding to a node, and the identifiers N1 and N2 inside it are variables representing nodes; "[]" is the identification symbol corresponding to a connecting edge, and the identifier E1 inside it is a variable representing a connecting edge; "-[E1]->" indicates that the connecting edge between nodes N1 and N2 is a directed edge E1, the source node is N1, and the destination node is N2. Similarly, it can be known that in a graph query statement, the connecting edge can also be limited to other directions: "(N1)<-[E1]-(N2)" represents a directed edge E1 with the source node being N2 and the destination node being N1; "(N1)-[E1]-(N2)" represents an undirected edge E1 between nodes N1 and N2.

[0054] Executing this graph query can match several subgraphs from the graph data to be queried. The RETURN clause can return the query results in a predefined form. For example, "RETURN p" returns the matched subgraphs in the form of paths. If the graph query statement in the above example is executed on the graph data shown in Attachment Figure 1 it will obtain the query results: A -> B and C -> B.

[0055] In addition, the RETURN clause can also include several query fields. For example, "RETURN N1.name" returns the value of the name attribute corresponding to the N1 node in the subgraphs that meet the query conditions.

[0056] In some examples, it is also possible to perform filtering queries on the labels or attributes of nodes to anchor specific nodes in the graph data. The examples are as follows:

[0057] MATCH p=(N1)-[E1]->(N2) WHERE N1.id = $id RETURN p;

[0058] Due to the chained programming feature of the graph query language, the graph query statement in this example is also equivalent to the following statement:

[0059] MATCH p=(N1:{id:$id})-[E1]->(N2) RETURN p;

[0060] Among them, "{}" is the identification symbol corresponding to the attribute, and the attribute filtering condition is inside it. This graph query statement will match the following subgraph paths in the graph data to be queried: two nodes connected by a directed edge, and the id attribute value of the source node is $id.

[0061] The above is a brief elaboration of some common example queries in graph queries. In actual practice, graph query statements can also include other clauses, such as LIMIT, DISTINCT, CASE, etc. This specification does not give examples and elaborate here.

[0062] As mentioned above, in actual applications, when business data is established and stored in the form of graph data, during the execution of graph queries, due to changes in business data or batch reading of graph data, the graph data may be frequently updated. This dynamic change of graph data usually causes the graph database to perform repeated calculations on the query of graph data. Below, taking the streaming calculation in a typical application as an example, the generation process of this problem will be elaborated.

[0063] In the scenario of streaming computing, data is usually not read in all at once, but is read in batches in a continuous and dynamic manner. For example, along with the generation of business data, new nodes and connection edges are generated, and these new vertex-edge data form incremental graph data, which is gradually read into the graph database in the form of a data stream. Alternatively, the graph database can also pull or receive the vertex-edge data newly generated within a certain time interval (e.g., 1 hour) from the business system at regular intervals (or called a time window), forming incremental graph data. Generally speaking, after the incremental graph data is read into the graph database, it will be temporarily stored in the form of a temporary graph, etc. When certain conditions are met, the incremental graph data will be superimposed on the current full graph data to obtain an updated version of the full graph data.

[0064] On the other hand, some graph queries need to be executed repeatedly multiple times, or the query results are updated in response to the arrival of incremental graph data. For such graph queries, if the updated full graph data is traversed and matched again each time, it will consume a large amount of computing resources, and the query results will also be repeated with historical data, causing duplicate calculations. Especially when dealing with large-scale graph data, this problem is more prominent.

[0065] Figure 2 Shows an execution schematic of a graph query in a typical streaming computing scenario. Referring to the accompanying drawings, Time Window #1 and Time Window #2 are two data read-in nodes that are sequentially generated in the time series. The numbers of the time windows are only used to reflect their relative order and do not represent a specific time window. In actual applications, there can be multiple time windows, and only two of them are used as examples for illustration in this example.

[0066] The graph query statement in the example is: MATCH p=(N1)->(N2)->(N3) RETURN p;

[0067] This query aims to find a non-loop chain path consisting of 3 nodes and 2 consecutive directed edges in the graph data.

[0068] After the graph database updates the streaming data of Time Window #1, the graph query is executed on the graph data, and the execution results are obtained: A->B->D, A->B->E. When Time Window #2 arrives, the streaming data increment therein is updated to the graph database, that is, the node C and the directed edge between node B and C are added. At this time, when the graph query is executed again to update the query results, usually a new complete traversal of all vertex-edges of the full graph is required, and all are recalculated to obtain the execution results A->B->D, A->B->E, A->B->C. This traditional full-volume calculation processing method, although it can ensure the correctness of the execution results, a large amount of duplicate calculations will inevitably generate additional overhead, slow down the execution speed of the graph query, and it is difficult to meet the business scenarios with high real-time requirements.

[0069] In response to this, the inventor found through research that the incremental point and edge data contained in the streaming data will only generate incremental query results that meet the graph query conditions within a subgraph of a limited range, and will not affect the historical query results. This means that the traditional full-scale calculation method will waste a large amount of computing power in the repeated calculation of historical query results. Therefore, such execution overhead can be optimized through technical means. In view of this, the embodiments of this specification propose a scheme for incremental graph query to save the overhead of graph query in the incremental update scenario. The technical solution will be elaborated in detail below.

[0070] Figure 3 Fig. shows a method architecture for performing incremental query on graph data. Among them, the target graph query is a query statement for graph data aiming to match paths with N hops, simply referred to as N-hop target graph query. In addition, it can be understood that the graph data shown in the accompanying drawings is only for illustrative purposes, and the graph data in actual applications will be constructed based on business data, including a large number of nodes and more complex connection relationships. Similarly, the incremental graph data shown in the accompanying drawings is also only for illustrative purposes. In specific practice, the incremental graph data contained in the streaming data is not limited to just one incremental edge and the two incremental nodes it connects.

[0071] Continue to refer to Figure 3 , the edges contained in the incremental graph data are called incremental edges, and the nodes connected by the incremental edges are called incremental nodes. After obtaining the incremental graph data, in combination with the existing graph data and the hop number N of the target graph query, perform N-hop neighbor expansion on the incremental nodes to determine a subgraph, called the extended subgraph. The extended subgraph can cover all incremental changes caused by the incremental graph data that will affect the target graph query results. Therefore, when performing the target graph query, only need to retrieve and match in the extended subgraph to obtain the incremental query results that meet the query conditions. Obviously, in this example method, the execution scope of the target graph query is limited to the extended subgraph with a limited scope, thus avoiding full-graph retrieval of the graph data, realizing incremental query of the graph data, saving computing resources, and improving query efficiency. This query strategy, especially when dealing with incremental queries of large-scale graph data, can significantly reduce unnecessary repeated matching calculations and improve the execution performance of graph queries.

[0072] Following the above technical concept, in Figure 4 , Fig. shows a flowchart of an incremental query method for graph data provided according to an embodiment of this specification. It can be understood that this method can be executed by a query engine in a graph database; the query engine can be implemented by any device, equipment, platform, or device cluster with computing and processing capabilities. Refer to Figure 4, the method at least includes the following steps: S401: Determine a target graph query, which involves an N-hop query. S403: Obtain incremental graph data, which includes several incremental edges and incremental nodes connected by the incremental edges. S405: Starting from the incremental nodes, perform N-hop neighbor expansion to obtain an augmented subgraph. S407: In the augmented subgraph, execute matching for the target graph query to obtain a query result.

[0073] The specific execution manners of the above steps will be described in detail below with reference to the accompanying drawings.

[0074] Step S401: Determine a target graph query, which involves an N-hop query.

[0075] As described above, a target graph query is usually expressed as a graph query statement written in accordance with a graph query language specification (such as GQL), which specifically describes the information (i.e., query result) that a user hopes to retrieve from graph data, including but not limited to, a subgraph, a path, and a node. In this step, a graph query statement input by the user can be received as the target graph query; alternatively, a graph query that has been previously matched by the database system and for which the query result needs to be updated currently can also be determined as the above target graph query.

[0076] The number of hops of a graph query refers to the number of connecting edges that need to be passed from the start node to the end node in the result path obtained from the graph data according to the graph query constraint conditions. The start node refers to the node element that is at the first position in the graph query path constraint conditions, and the end node refers to the node element that is at the last position.

[0077] In an example, the graph query statement: MATCH p=(N1)->(N2) RETURN p; the start node is N1, the end node is N2, and there is 1 connecting edge between the two nodes. Therefore, this graph query is a 1-hop query, that is, N = 1.

[0078] In another example, the graph query statement: MATCH p=(N1)->(N2)->(N3) RETURN p; the start node is N1, the end node is N3, and there are 2 connecting edges between the two nodes. Therefore, this graph query is a 2-hop query, that is, N = 2.

[0079] And so on, it is not difficult to know that in a graph query path constraint sequentially connected by single-direction connecting edges (referring to the node variables on the left side of the connecting edge identifier, all are out-edges or all are in-edges), the number of connecting edge identifiers is the number of hops N of this graph query. However, it should be noted that in some practices, the graph query path constraint may include connecting edge identifiers in different directions (mixed out-edges and in-edges, or undirected edges), such as the graph query statement:

[0080] MATCH p=(N1)->(N2)<-(N3)-(N4) RETURN p;

[0081] In such graph queries, the starting node is the first node element in the path constraint condition, i.e., N1, and the ending node is the last node element, i.e., N4. Each connection edge identifier between the two nodes, regardless of direction, is counted as one hop. Therefore, this example graph query is a 3-hop query, i.e., N = 3.

[0082] In addition, in some equivalent expressions of graph query statements, the value of N can also be determined according to the hop quantifier variable in the path constraint condition. For example, the graph query statement:

[0083] MATCH p=(N1)-[*3]->(N2) RETURN p; or its equivalent expression:

[0084] MATCH p=((N1)-[]->(N2)){3} RETURN p;

[0085] In such graph queries, the number 3 is the hop quantifier, indicating that there are 3 directed edges between the starting node N1 and the ending node N2 in the result path. Therefore, the hop count of the example graph query statement can be determined according to the hop quantifier variable, i.e., N = 3.

[0086] In the introduction of each embodiment of this specification, graph query statements with out-edges as connection edges will be used as examples for elaboration. However, the technical ideas involved can be equivalently inferred and applied to the execution of other graph query statements containing directed / undirected connection edges.

[0087] Next, in step S403: Obtain incremental graph data, which includes several incremental edges and the incremental nodes connected by the incremental edges.

[0088] As mentioned above, in the streaming computing scenario, the newly added graph data is not loaded into the graph database all at once, but is gradually read into the graph database in batches over time. Figure 5 Disclosed is a schematic diagram of streaming acquisition of incremental graph data.

[0089] Referring to the accompanying drawings, as time goes by, new graph data is continuously generated. Under the streaming computing framework, this graph data, as streaming data, is read into the graph database batch by batch. In a typical practice, the incremental graph data is the latest batch of graph data streamed to the graph database. That is to say, when a batch of graph data is read into the graph database, this batch of graph data can be used as incremental graph data for the execution of graph queries. In some specific implementations, the generation of graph data does not have clear batch boundaries. For example, graph data generated along with the update of business data in the production environment does not have a fixed generation frequency and does not have clear batch division conditions. At this time, a fixed time interval can be preset, and each time interval is regarded as a time window. The graph data generated and read within the time window is incremental graph data. Refer to Figure 5 , the graph database storing the graph data obtains the incremental data corresponding to the time window every preset time window, and the incremental graph data is the incremental data of the most recent time window.

[0090] It should be noted that the above uses the generation of new graph data as an example to elaborate on the concept of incremental graph data, but it does not represent a limitation on the application scenario. In actual applications, even if no new graph data is generated, the streaming computing framework can also be applied to the existing graph data, and the graph data can be read into the graph database / graph query engine batch by batch.

[0091] In one example, the graph data targeted by the target graph query is a large-scale graph, and the I / O throughput of the system is not sufficient to support reading it all at once. At this time, streaming computing can also be applied to divide the large-scale graph data into several small batch sub-graph data and read it in gradually.

[0092] In another example, the target graph can also be distributedly stored in different computing devices through graph partitioning techniques (for example, vertex partitioning, edge partitioning). At this time, streaming computing can also be applied to gradually read the sub-graph data stored in different computing devices.

[0093] In short, the application scenarios for generating incremental graph data are diverse and will not be listed one by one in the embodiments of this specification.

[0094] In the incremental graph data, it can include connecting edges, nodes, or a combination of nodes and connecting edges. In the embodiments of this specification, the connecting edges in the incremental graph data are regarded as incremental edges, and the nodes connected by the incremental edges are all regarded as incremental nodes, even if the node itself already exists in the existing graph data.

[0095] Figure 6 Examples of determining incremental edges and incremental nodes in different incremental graph data are given. Refer to Figure 6Left schematic part (Example 1). When both the connecting edges and nodes included in the incremental graph data are new data, they can be directly determined as incremental edges and incremental nodes respectively. Continue to refer to Figure 6 Middle schematic part (Example 2). When the incremental graph data only includes new connecting edges and the nodes are existing nodes in the historical graph data, the new connecting edges can be determined as incremental edges, and all the nodes connected by the incremental edges are regarded as incremental nodes. Similarly, in Figure 6 Right schematic part (Example 3). The incremental connecting edges can be determined as incremental edges, the newly added nodes are determined as incremental nodes, and the other node (existing node) connected by the incremental edges is also determined as an incremental node.

[0096] After obtaining the incremental graph data and determining the incremental nodes, in step S405, starting from the incremental nodes, perform N-hop neighbor expansion to obtain an extended subgraph.

[0097] As mentioned above, for target graph queries, the incremental update of graph data usually only affects nodes and edges within a limited range, rather than the entire graph. By identifying these affected points and edges, the target graph query operation can be limited to the local subgraph range, thereby achieving incremental query for graph data. Retrieving these points and edges that may be affected by the incremental update can be completed by means of the topological structure of the graph data. Therefore, in this step, based on the incremental nodes, combined with the graph data and the hop number N of the target graph query, extract the N-hop neighbor subgraph of the incremental nodes (i.e., the extended subgraph). The N-hop neighbor subgraph not only ensures comprehensive coverage of the points and edges affected by the incremental update but also effectively controls the scale of the subgraph.

[0098] In some implementations, a traditional graph traversal algorithm can be used, supplemented by a limit on the traversal level (i.e., N hops), to achieve the extraction of the extended subgraph. For example, use the Breadth First Search (BFS) algorithm, starting from the incremental nodes, perform N-layer traversal in the graph data (including incremental graph data and existing graph data), and extract the N-hop neighbor subgraph of the incremental nodes as the extended subgraph.

[0099] In some implementation manners, the N-hop expansion of the incremental nodes can be achieved through message propagation to obtain the extended subgraph. The message propagation method has more significant technical advantages for the scenario of distributed graph storage. Because in a distributed scenario, graph data is stored in multiple computing nodes, and traditional graph traversal algorithms need to design an additional traversal synchronization mechanism across computing nodes to adapt to the distributed computing environment. By adopting message propagation technology, the extraction of subgraphs across computing nodes can be effectively achieved.

[0100] Specifically, to extract the N-hop neighbor subgraph, starting from the incremental node, N rounds of message propagation can be performed. A single round of message propagation includes propagating a first message from a first node to its neighbor nodes. Among them, in the first round of message propagation, the incremental node is used as the first node, and in non-first-round message propagations, the node that received the first message in the previous round is used as the first node. Then, the nodes and edges involved in the N rounds of message propagation are classified into the extended subgraph. Figure 7 Schematic diagram showing a method for extracting an extended subgraph in this embodiment.

[0101] Referring to the accompanying drawings, taking the 2-hop neighbor subgraph as an example (i.e., N = 2), after determining the incremental node (the node shown by the dashed line in the accompanying drawings), starting from the incremental node (i.e., the first node), sequentially send EvolveMessage (the aforementioned first message) to its neighbor nodes. EvolveMessage can be an encapsulated message body carrying the node information of the starting node, where the node information can be, for example, a unique identifier such as a node id, which is not specifically limited here. The node that receives the EvolveMessage is used as the starting node in the new round of message propagation and continues to send the EvolveMessage to its neighbor nodes. After such N rounds of message propagation, all the nodes that received the EvolveMessage and the associated connecting edges form an N-hop neighbor subgraph starting from the incremental node, that is, the extended subgraph.

[0102] In the process of implementing message propagation, it is also necessary to avoid the occurrence of propagation loops to prevent infinite loops. In some engineering implementations, the sender id can be carried in the first message, and in non-first-round message propagations, determine the target neighbor node of any first node. The target neighbor node is a neighbor node of the first node, and its node id is different from the sender id carried in the first message received by the first node. Send the first message from the first node to the target neighbor node, where the node id of the first node is used as the sender id.

[0103] That is to say, the target neighbor node is the destination node to which the first node will send the first message in the current round. When determining the target neighbor node of the first node, it is necessary to determine whether the neighbor node has ever sent the first message to the first node. If it has, it means that the neighbor node has been traversed and cannot be used as the target neighbor node of the first node. In some scenarios, the graph data has a multi-hop loop topology structure. When determining the target neighbor node, the condition that the neighbor node id is different from any sender id in all the previous first messages received by the first node can be used as the judgment condition.

[0104] It should be noted that in this engineering implementation, the node ID is used as the unique identifier for the judgment of target neighbor nodes, but it does not represent a specific limitation on the unique identifier; in other engineering implementations, other variables with unique identifier functions can also be used as the judgment basis to determine the target neighbor nodes, and the embodiments of this specification will not give examples one by one for this.

[0105] In addition, in some specific implementations, the target neighbor nodes can also be purposefully screened in combination with the path constraint conditions of the target graph query. For example, in the constraint conditions of the target graph query, all connection edges are in a single direction. Since the message propagation is from the downstream node to the upstream node, the neighbor nodes connected by a specific directed edge can be screened according to the reverse direction of this single direction as the target neighbor nodes.

[0106] In an example, the target graph query is: MATCH p=(N1)->(N2)->(N3) RETURN p; from the reverse direction of path retrieval (from right to left), all the path constraint conditions are incoming edges. In each message propagation round, for the first node, the neighbor nodes connected by the incoming edges of this node can be determined as the target neighbor nodes.

[0107] Correspondingly, in another example, the target graph query is: MATCH p=(N1)<-(N2)<-(N3) RETURN p; from the reverse direction of path retrieval (from right to left), all the path constraint conditions are outgoing edges. In each message propagation round, for the first node, the neighbor nodes connected by the outgoing edges of this node can be determined as the target neighbor nodes.

[0108] The above describes the method for extracting the augmented subgraph. It is worth noting that although the augmented subgraph can limit the retrieval range of the target graph query and improve the graph query efficiency, however, the subgraph extraction still consumes a certain amount of computing resources. In order to effectively improve the end-to-end execution efficiency of the graph query, it is necessary to comprehensively consider and weigh the consumption of computing resources by extracting the augmented subgraph and directly executing the target graph query respectively. In some scenarios, directly executing the target graph query can obtain the query result more efficiently. That is to say, the improvement effect of implementing incremental query on the overall query efficiency is limited, and the target graph query can be directly executed to obtain the full-scale query result.

[0109] For example, in some scenarios, if the starting node in the target graph query is clearly restricted to a uniquely specified node, such that a node can be uniquely determined from the graph data as the starting node of the target graph query, then in this case, there is no need to extract an extended subgraph for the incremental nodes, and instead, the target graph query can be directly executed. In other words, in the embodiments of this specification, preferably, when the target graph query does not contain a constraint condition that restricts the starting node to a uniquely specified node, the above-mentioned extraction of the extended subgraph is performed. The constraint condition can be a constraint on the unique identifier of the node. For example, the restriction on the node id in the WHERE clause; it can also be a unique restriction on other attributes of the node. For example, the restriction on the name attribute of the node in the path constraint: (N1:{name:$name}), and no further examples will be given here.

[0110] In some scenarios, when the target graph query restricts a target path that is less than the low hop count threshold (Th1), the computing resources consumed by directly executing the target graph query will be lower than the resource consumption of "first extracting the extended subgraph and then executing the query". Correspondingly, in some other scenarios, when the target graph query restricts a target path that is greater than the high hop count threshold (Th2), the scale of the extended subgraph extracted based on the hop count N will be close to the scale of the original graph data, and the goal of extracting the extended subgraph to limit the execution scope of the target graph query will be difficult to achieve. Therefore, when the target graph query meets any one of the above two scenarios, the target graph query can be directly executed in the graph with the incremental graph data updated. The above low hop count threshold and high hop count threshold can be thresholds determined based on expert experience. Typically, according to practical experience, the low hop count threshold can be set to 2, and the high hop count threshold can be set to a specific threshold with reference to the scale of the graph data in different scenarios. For example, 5 hops (in some types of graph data, in principle, any two nodes in the graph can be connected through 5 hops). In other words, in the embodiments of this specification, the value of N is greater than or equal to 2 and less than a preset first threshold.

[0111] After obtaining the extended subgraph through the above steps, in step S407, in the extended subgraph, perform the matching for the target graph query to obtain the query result.

[0112] In this step, the target graph query can be executed only in the obtained extended subgraph. Since the extended subgraph has comprehensively covered the points and edges that may be affected by the incremental update, and the number of points and edges in it is much smaller than the number of points and edges in the entire graph, the incremental query implemented through this step will reduce unnecessary repeated calculations and greatly improve the execution efficiency of the target graph query.

[0113] In one implementation, N-hop traversal can be performed in the augmented subgraph according to the query conditions in the target graph query. According to the results of the N-hop traversal, the query results can be determined. In an example, if the query conditions do not contain any restrictions on the starting node, then any node in the augmented subgraph can be used as the starting node for the query to perform matching for the target graph query. In another example, if the query conditions include multi-starting point constraints (for example, filtering by the amount attribute of the node: WHERE amount>30), then the nodes in the augmented subgraph that meet the multi-starting point constraints can be used as the starting nodes for the query to perform matching for the target graph query. In other examples, purposeful N-hop traversal can also be performed according to various other constraint conditions included in the target graph query to further save computing resources, and the embodiments of this specification will not list them one by one.

[0114] In the above implementation, several target paths that meet the query conditions can be retrieved from the augmented subgraph according to the query conditions in the target graph query. Compared with the historical query results for the graph data, these target paths include the incremental query results that meet the query conditions introduced by the incremental graph data, and may also include some target paths that overlap with the historical query results. Therefore, in order to implement incremental query, the query results can be compared with the historical query results, and through set difference calculation (such as set difference operation, target path coverage check, etc.), redundant paths can be removed from the query results to obtain the incremental query results.

[0115] In some engineering implementations, the incremental query results can be filtered out by traversing each target path in the query result set and the historical query result set and comparing them one by one. The time complexity of this engineering implementation is O(n 2 ).

[0116] Since the target paths corresponding to the incremental query results are generated due to the addition of the incremental graph data, each incremental target path will contain at least one incremental edge. The presence or absence of the incremental edge can be used to determine whether the target path belongs to the incremental query results. Based on this, in some engineering implementations, during the N-hop traversal process, in response to traversing through an incremental edge, a dynamic marker, hereinafter referred to as the first marker, can be added in this traversal step. After the N-hop traversal is completed, the paths containing the first marker are filtered out from the traversal paths that meet the query conditions as the result paths, that is, the incremental query results.

[0117] The first flag can be a boolean variable or other forms of flag variables, which is used to identify the existence of incremental edges in the target path. The specific implementation of the first flag is not limited in this embodiment. In the implementation of this project, the first flag is added to the target path along with the execution of the N-hop traversal. Therefore, after the N-hop traversal is completed and the query result is obtained, the target path with the first flag can be filtered out by traversing the query result set, which is the incremental query result. The time complexity of this project implementation is O(n).

[0118] In addition, in some practices, graph data is distributed and stored in several computing nodes. In this scenario, the above project implementation needs to perform cross-node data reading to obtain the basis for judging the incremental query result (the first flag). To optimize this process, according to an implementation, during the N-hop traversal process, after adding the first flag, the first flag can be recorded in each subsequent traversal step. In this way, unnecessary cross-node reading can be reduced and the screening efficiency can be improved during the process of screening the incremental query result. Specifically, when the first flag is passed to the termination node of the target path that meets the query conditions along with the execution of the N-hop traversal, the target path with the first flag can be immediately identified as the incremental query result; conversely, if the first flag has not been added / transmitted when the N-hop traversal reaches the termination node of the target path that meets the query conditions, this target path duplicates the records in the historical query results and does not belong to the incremental query result. The time complexity of this project implementation is O(1).

[0119] The above is an introduction to the main process of an incremental query method for graph data provided in the embodiments of this specification. Although the above embodiments mainly use the graph query statements in the GQL standard and their corresponding graph databases as examples to elaborate on the method process. However, the technical concept embodied therein can also be applied to other graph databases and the execution process of related graph query statements.

[0120] By using the above method provided in the embodiments of this specification, after obtaining the incremental graph data, the extended subgraph that may produce incremental query results can be determined. When performing a graph query, only graph matching that meets the query conditions is performed in the extended subgraph, realizing incremental query for graph data and avoiding full-graph calculation, thereby improving the execution performance of the graph query.

[0121] In this specification, the "first" in terms such as the first node and the first message, and the corresponding "second", "third" (if any) in the text are only for the convenience of distinction and description and do not have any restrictive meaning.

[0122] The above content describes specific embodiments of this specification, and other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than in the embodiments, and still achieve the desired result. Additionally, the processes depicted in the figures do not necessarily have to be performed in the specific order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.

[0123] Figure 8 FIG. is a schematic diagram of an incremental query device for graph data according to an embodiment of this specification. The device 800 is deployed in a computing device, which can be implemented by any device, equipment, platform, device cluster, etc. with computing and processing capabilities. This device embodiment corresponds to Figure 4 the method embodiment shown. The device 800 includes:

[0124] A determination module 801, configured to determine a target graph query, which involves an N-hop query.

[0125] An acquisition module 802, configured to acquire incremental graph data, which includes a number of incremental edges and the incremental nodes connected by the incremental edges.

[0126] An expansion module 803, configured to start from the incremental nodes and perform N-hop neighbor expansion to obtain an expanded subgraph.

[0127] An execution module 804, configured to perform matching for the target graph query in the expanded subgraph to obtain a query result.

[0128] According to an embodiment of another aspect, this specification also provides a computer program product, including a computer program / instructions, which when executed by a processor, implement the steps of the foregoing method in combination with Figure 4 the method.

[0129] According to an embodiment of yet another aspect, this specification also provides a computing device, including a memory and a processor, characterized in that the memory stores executable code, and when the processor executes the executable code, it implements the steps of the foregoing method in combination with Figure 4 the method.

[0130] Those skilled in the art should be able to realize that in the above one or more examples, the functions described in the embodiments of the present invention can be implemented by hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or codes on a computer-readable medium.

[0131] The specific embodiments described above further elaborate on the objectives, technical solutions, and beneficial effects of the embodiments of the present invention. It should be understood that the above description is only the specific embodiments of the embodiments of the present invention and is not used to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made on the basis of the technical solutions of the present invention shall be included within the protection scope of the present invention.

Claims

1. An incremental query method for graph data, comprising: Determine the target graph query, which involves N-hop queries; Obtain incremental graph data, which includes a number of incremental edges and incremental nodes connected by the incremental edges; Taking the incremental node as the starting point, N-hop neighbor expansion is performed to obtain an extended subgraph; In the expanded subgraph, a match is performed for the target graph query to obtain a query result.

2. The method according to claim 1, wherein: The target graph query does not include a constraint condition that limits the starting node to a uniquely specified node.

3. The method according to claim 1, wherein: The value of N is greater than or equal to 2 and less than a preset first threshold.

4. The method according to claim 1, wherein: The incremental graph data is the latest batch of graph data streamed to the graph database.

5. The method according to claim 1, wherein: The graph database storing the graph data obtains incremental data corresponding to the time window every preset time window; the incremental graph data is the incremental data of the most recent time window.

6. The method according to claim 1, wherein: The N-hop neighbor expansion is performed to obtain an extended subgraph, including: Perform N rounds of message propagation, where a single round of message propagation includes propagating a first message from a first node to its neighboring nodes; wherein in a first round of message propagation, the incremental node is the first node, and in a non-first round of message propagation, the node that received the first message in the previous round is the first node; The nodes and edges involved in the N rounds of message propagation are included in the expanded subgraph.

7. The method according to claim 6, wherein: The first message carries a sender ID; The non-first-round message dissemination includes: Determine a target neighbor node of any first node, where the target neighbor node is a neighbor node of the first node, and the node ID is different from the sender ID carried in the first message received by the first node; A first message is sent from the first node to the target neighbor node, wherein the node id of the first node is used as the sender id.

8. The method according to claim 1, wherein: The performing of matching for the target graph query to obtain a query result includes: According to the query condition in the target graph query, perform N-hop traversal in the expanded subgraph; The query result is determined based on the result of the N-hop traversal.

9. The method according to claim 8, wherein: The N-hop traversal includes, in response to traversing through an incremental edge, adding a first marker in the traversal step; The determining of the query result includes selecting a path including the first mark from the traversal paths meeting the query condition as a result path.

10. The method according to claim 9, wherein: After adding the first mark, the method further includes: The first mark is kept and recorded in each subsequent traversal step.

11. A device for incremental query of graph data, the device comprising: A determination module, configured to,determine a target graph query, which involves an N-hop query; The acquisition module is configured to acquire incremental graph data, including a plurality of incremental edges and incremental nodes connected by the incremental edges; An expansion module is configured to perform N-hop neighbor expansion starting from the incremental node to obtain an expanded subgraph; The execution module is configured to perform matching for the target graph query in the expanded subgraph to obtain a query result.

12. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.

13. A computing device comprising a memory and a processor, characterized in that: The memory stores executable codes, and when the processor executes the executable codes, the method according to any one of claims 1 to 10 is implemented.