Heterogeneous collaborative sub-graph matching method for dynamic graph

By adopting the coordinated working mode between CPU and GPU in dynamic graph subgraph matching, the subgraph matching problem of batch update of dynamic graphs is solved, and the problem of long update processing time and insufficient use of GPU resources in the prior art is solved, and the processing efficiency of subgraph matching is improved.

CN120067713AActive Publication Date: 2025-05-30NORTHEASTERN UNIV CHINA
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510535174.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-27
Publication Date
2025-05-30
Estimated Expiration
2045-04-27

AI Technical Summary

Technical Problem

When the prior art processes subgraph matching for batch update dynamic graphs, the update processing time is long, and GPU computing resources are not fully utilized, resulting in low subgraph matching processing efficiency.

Method used

A heterogeneous collaborative subgraph matching method for dynamic graphs is proposed. Through the coordinated work of the CPU and the GPU, the CPU is responsible for batch update operations and mining the dynamic data subgraph area affected by the update. The GPU is responsible for subgraph matching processing of the original data graph and merging the results of the two to improve processing efficiency.

Benefits of technology

Through the CPU/GPU heterogeneous framework, the negative impact of batch dynamic updates on sub-graph matching is avoided, and the processing efficiency of dynamic graph sub-graph matching is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067713A_ABST
    Figure CN120067713A_ABST
Patent Text Reader

Abstract

The invention provides a heterogeneous collaborative sub-graph matching method for a dynamic graph, and belongs to the technical field of sub-graph matching. The method comprises the following steps: acquiring a data graph, a query graph set and a dynamic update sequence; the data graph is stored in a GPU; the data graph and the dynamic updating sequence are stored in a CPU, and the data graph is updated through the dynamic updating sequence; according to the query graph set, performing parallel sub-graph matching by using a data graph in the GPU to obtain a first sub-graph matching result; according to the query graph set and the dynamic update sequence, performing sub-graph matching by using a data graph in the CPU to obtain a second sub-graph matching result; and combining the first sub-graph matching result and the second sub-graph matching result to obtain a final sub-graph matching result, and updating the data graph in the GPU according to the dynamic updating sequence. According to the method provided by the invention, the time required for matching the sub-hypergraphs is greatly shortened.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of subgraph matching, and particularly relates to a heterogeneous collaborative subgraph matching method for dynamic graphs. Background Art

[0002] The graph structure contains vertex and edge information and is widely used to represent complex relationships between various objects. In most real-world applications, the graph information structure changes over time by adding / removing vertices / edges, which is called a dynamic graph. For example, a social network is a typical dynamic graph. The registration and cancellation of new users can be regarded as the addition and deletion of vertices in the graph respectively, and the addition and cancellation of follows can be regarded as the addition and deletion of edges in the graph respectively. How to effectively manage and query dynamic graphs has attracted increasing attention.

[0003] Subgraph matching is one of the most basic problems in graph analysis and has extensive applications in protein interaction network analysis, social network analysis, compound search, etc. Given a query graph and a data graph, subgraph matching is to find all subgraphs in the data graph that are isomorphic to the query graph. Although many solutions to the subgraph matching problem have been proposed, subgraph matching is still an NP-hard problem due to the high cost of the enumeration matching process. Among them, the NP-hard problem refers to those problems that cannot find an exact solution in polynomial time.

[0004] Currently, the main concern for dynamic graphs is to solve the continuous subgraph matching problem. Continuous subgraph matching performs incremental matching for each individual update in a continuous manner. However, in the current data processing environment, the data scale is huge and dynamic updates are usually applied in batches. Retrieving a certain pattern in a dynamically updated graph in batches is of great significance. For example, in a social network, the social network changes in batches all the time. Retrieving the social matching situation of a certain individual is of great significance for social analysis.

[0005] Currently, all subgraph matching algorithms based on the CPU (Central Processing Unit) adopt a filtering and enumeration framework. First, some filtering rules are designed according to the query graph structure to filter out candidate nodes in the data graph that cannot provide the final solution. Then, a data structure is designed to store the remaining candidate matching nodes. In the enumeration stage, a matching order is determined according to the graph structure and candidate statistics. The algorithm uses depth-first traversal along the matching order to find the matching result in the way of expanding nodes. The GPU (Graphic Processing Unit) is a parallel processor with thousands of cores and high-bandwidth memory. At present, many basic graph operations of large graphs are accelerated by using the large-scale parallel computing power of the GPU, such as breadth-first traversal of graphs, shortest path, and minimum spanning tree generation processing. To accelerate subgraph matching, many GPU-based subgraph matching algorithms have been proposed. The existing methods make full use of the parallel computing power of the GPU and can improve the processing efficiency of subgraph matching.

[0006] If the GPU subgraph matching method is used to process the subgraph matching problem for batch-update dynamic graphs, the graph update operation needs to be executed first, and then the subgraph matching method is processed. Since the batch update processing time is long, the GPU computing resources are not fully utilized in the update processing stage, resulting in a low processing efficiency of the subgraph matching method, and the data graph update operation accounts for too much of the total matching time. Summary of the Invention

[0007] Aiming at the deficiencies of the prior art, this application proposes a heterogeneous collaborative subgraph matching method for dynamic graphs. The CPU and GPU execute tasks simultaneously. The CPU mines the dynamic data subgraph area affected by the update for batch update operations, and finds the matching result of the dynamic data subgraph of the query graph for the dynamic data subgraph, that is, finds the incremental subgraph matching. At the same time, the GPU is used to perform subgraph matching processing on the original data graph, and the matching result is filtered to obtain the matching result of the static data subgraph not affected by the update. The two results are combined to obtain the subgraph matching for the dynamic graph. In the offline stage, dynamic update processing is performed on the GPU dynamic graph storage structure, and the impact of batch dynamic updates on subgraph matching is avoided through the CPU / GPU heterogeneous framework, improving the processing efficiency of dynamic graph subgraph matching. This method has strong research and application value.

[0008] In the first aspect, this application proposes a heterogeneous collaborative subgraph matching method for dynamic graphs, including:

[0009] Obtain a data graph, a set of query graphs, and a dynamic update sequence;

[0010] Save the data graph to the GPU;

[0011] Save the data graph and the dynamic update sequence to the CPU, and update the data graph using the dynamic update sequence;

[0012] According to the query graph set, perform parallel subgraph matching using the data graph in the GPU to obtain the first subgraph matching result;

[0013] According to the query graph set and the dynamic update sequence, perform subgraph matching using the data graph in the CPU to obtain the second subgraph matching result;

[0014] Merge the first subgraph matching result and the second subgraph matching result to obtain the final subgraph matching result, and update the data graph in the GPU according to the dynamic update sequence.

[0015] The step of saving the data graph to the GPU includes:

[0016] In the GPU, save the data graph as a static data graph with a four-layer array structure. The static data graph includes: the first-layer array is a label array for storing the label information of the nodes in the data graph; the second-layer array is a neighbor physical address array for recording the physical address information of the neighbor arrays of each node in the data graph; the third-layer array is a node array for recording each node in each data graph. The node array includes two elements, namely the number of neighbor nodes and the size of the neighbor array; the fourth-layer array is a neighbor array, and a neighbor array is independently allocated for each node in the data graph.

[0017] The step of performing parallel subgraph matching using the data graph in the GPU according to the query graph set to obtain the first subgraph matching result includes:

[0018] Find the candidate set of each node in each element of the query graph set in the static data graph;

[0019] Generate the candidate set of each edge according to the candidate set of each node;

[0020] According to the candidate set of each node and the candidate set of each edge, transform each element in the query graph set into an edge generation tree;

[0021] Split an edge generation tree into multiple independent query edges;

[0022] Use the GPU to process multiple independent query edges in parallel, and find the branch candidate results in each independent query edge;

[0023] According to the branch intersection points of each independent query edge, splice each branch candidate result to obtain the first subgraph matching result.

[0024] The step of finding the candidate set of each node in each element of the query graph set in the static data graph includes:

[0025] Determine the query order in the static data graph according to the ascending order of the function values of the first calculation function;

[0026] According to the query order, based on the first filtering rule and the second filtering rule, find the candidate set of each node in each element of the query graph set in the static data graph.

[0027] The first calculation function has the following calculation formula: ;

[0028] where f(u) is the first calculation function, is the number of nodes in the static data graph that have the same label as a certain query node u in the query graph set, and d(u) is the degree of a certain query node u in the query graph set.

[0029] The first filtering rule is a label filtering rule. According to the label information of a certain query node u in the query graph set, generate candidate nodes for the query node u in the static data graph.

[0030] The second filtering rule is a degree filtering rule. According to the degree information of a certain query node u in the query graph set, generate candidate nodes for the query node u in the static data graph.

[0031] Generating the candidate set of each edge according to the candidate set of each node includes:

[0032] For the query edge , starting from the starting query node u, obtain the candidate node array Cand_array of the query node u. Obtain a candidate node v in the candidate node array. Among the neighbor nodes of the candidate node v, find the number of candidate nodes that can be used as u'. Form an array Cand_count with the number of u' candidate nodes. Perform an exclusive prefix process on the array Cand_count to obtain the storage index subscript array of the candidate node v' of u' in the neighbor nodes ; Use the candidate node array Cand_array and the index subscript array to store each query edge to obtain the candidate set of each edge.

[0033] Converting each element in the query graph set into an edge generation tree according to the candidate set of each node and the candidate set of each edge includes:

[0034] Starting from the query point with the largest degree in each query graph in the query graph set, according to the depth-first traversal strategy, use the query edges visited during the traversal as the paths in the generation tree until all the query edges in the candidate set of each edge are visited, and end the traversal process to obtain an edge generation tree.

[0035] Using the query graph set and the dynamic update sequence, perform subgraph matching using the data graph in the CPU to obtain the second subgraph matching result, including:

[0036] Find the area affected by the dynamic update sequence by the maximum hop count of the vertices in the query graph in the query graph set, and use the area as the dynamic data subgraph;

[0037] Determine whether each updated node in the data graph is a candidate node of a certain node in the query graph before and after the update. If so, perform a depth-first traversal of the maximum hop count of the vertices in the corresponding query graph starting from the modified node, and use the subgraph composed of the depth-first traversals of all updated nodes as the incremental data subgraph; if not, skip this updated node and do nothing;

[0038] According to the third filtering rule, the fourth filtering rule, and the fifth filtering rule, find the candidate set of each node in the query graph in the query graph set in the incremental data subgraph;

[0039] Determine the query order of the incremental data subgraph matching according to the ascending order of the function values of the second calculation function;

[0040] Establish an auxiliary storage structure for the candidate set of each node. The auxiliary storage structure is a set of data vertices that find the matching relationship of each query vertex in the data graph after filtering and records whether there is an edge between the data vertices;

[0041] According to the auxiliary storage structure and the backtracking algorithm, find all subgraphs isomorphic to the query graph in the dynamic data subgraph to obtain the second subgraph matching result.

[0042] The third filtering rule is the label filtering rule. According to the label information of the query nodes in the query graph, generate candidate nodes for the query nodes in the incremental data subgraph;

[0043] The fourth filtering rule is the degree filtering rule. According to the degree information of the query nodes in the query graph, generate candidate nodes for the query nodes in the incremental data subgraph;

[0044] The fifth filtering rule is the neighbor label frequency filtering rule, which uses the neighbor nodes of the query nodes in the query graph Filter in the following way: Given a node , if there exists a label , then there exists ;

[0045] where , , if there does not exist a label , then filter out node v, where u is the query node in the query graph, N(u) is the neighbor nodes of u, v is the node in the data graph, C(u) is the matching candidate set of u, L(N(u)) is the labels of the neighbor nodes of u, |N(v, l)| is the number of nodes with label l among the neighbor nodes of v, |N(u, l)| is the number of nodes with label l among the neighbor nodes of u, u’ is another query node in the query graph, and L(N(u’)) is the labels of the neighbor nodes of u’.

[0046] The second calculation function has the following calculation formula: f’(u)=|C’(u) / d(u)|;

[0047] Among them, f’(u) is the second calculation function, C’(u) is the candidate set of each node in the query graph after filtering, and d(u) is the degree size of a certain query node u in the query graph set.

[0048] Beneficial effects:

[0049] This application proposes a heterogeneous collaborative subgraph matching method for dynamic graphs. The CPU and GPU execute tasks simultaneously. The CPU, for batch update operations, mines the dynamic data subgraph area affected by the update, and finds the dynamic data subgraph matching result of the query graph for the dynamic data subgraph, that is, finds the incremental subgraph matching. At the same time, the GPU is used to perform subgraph matching processing on the original data graph, and the matching result is filtered to obtain the matching result of the static data subgraph not affected by the update. The two results are combined to obtain the subgraph matching for the dynamic graph. Through the CPU / GPU heterogeneous framework, the impact of batch dynamic updates on subgraph matching is avoided, and the efficiency of dynamic graph subgraph matching processing is improved. Brief Description of the Drawings

[0050] Figure 1 Flowchart of a heterogeneous collaborative subgraph matching method for dynamic graphs according to an embodiment of the present invention;

[0051] Figure 2 Instance graph query graph Q of the batch dynamic graph subgraph matching instance according to an embodiment of the present invention 1 ;

[0052] Figure 3 Query graph Q of the batch dynamic graph subgraph matching instance according to an embodiment of the present invention 2 ;

[0053] Figure 4 Data graph of the batch dynamic graph subgraph matching instance according to an embodiment of the present invention

[0054] Figure 5 Dynamic graph of the batch dynamic graph subgraph matching instance according to an embodiment of the present invention

[0055] Figure 6 Schematic diagram of the CPU / GPU heterogeneous collaborative subgraph matching process according to an embodiment of the present invention;

[0056] Figure 7 Schematic diagram of the storage structure of the data graph in the GPU according to an embodiment of the present invention;

[0057] Figure 8 Update v in the embodiment of the present invention 2 Schematic diagram of the neighbor array;

[0058] Figure 9 Schematic diagram of the storage structure of the updated data graph in the GPU according to an embodiment of the present invention;

[0059] Figure 10 Schematic diagram of the candidate set of nodes according to an embodiment of the present invention;

[0060] Figure 11 Edge spanning tree and the edge spanning tree in the splitting process according to an embodiment of the present invention;

[0061] Figure 12 Edge spanning tree branch splitting diagram according to an embodiment of the present invention. Detailed implementation manners

[0062] The following combines the drawings and embodiments to further describe in detail the specific implementation manners of the present application. The present application proposes a heterogeneous collaborative subgraph matching method for dynamic graphs, and for the first time proposes a subgraph matching method for batch-updated dynamic graphs. In the face of the batch update of the data graph, a subgraph matching processing task is performed on a certain query graph. This method aims at the subgraph matching problem of dynamic graphs and proposes a parallel subgraph matching of CPU and GPU collaboration. The GPU is used to parallelly match the unchanged part of the data graph, and the CPU is used to match the changed part of the data graph. Then, the two matching results are merged, and finally the data graph in the GPU is updated. The method of the present application greatly reduces the time required for sub-hypergraph matching.

[0063] Embodiment:

[0064] The present application proposes a heterogeneous collaborative subgraph matching method for dynamic graphs, as Figure 1 shown, including:

[0065] Step S1: Obtain a data graph, a query graph set, and a dynamic update sequence;

[0066] In this embodiment, a query graph set , Q 1 , Q 2 , …, Q M are elements in the query graph set, a data graph G and a data graph dynamic update stream sequence are given, where It is an operation of adding or deleting edges in the batch data graph, where is an operation of adding or deleting edges. Subgraph matching for dynamic graphs is performed on the data graph G according to the updated data graph to find the subgraph matching results of the query graphs in the query graph set .

[0067] As Figure 2 shown, the initial data graph is Figure 4 shown, the query graph set , as Figure 2 , Figure 3 shown, for the query graph , the subgraph matching result found is the subgraph composed of the deepened nodes and the bold edges, denoted as = , . The dynamic update sequence , respectively representing adding the edge between and , deleting the edge between and (as shown by the dashed edge), the updated dynamic graph is as Figure 5 shown. At this time, for the query graph , the subgraph matching result found in is , such as the subgraph composed of the deepened nodes, v 4 , v 7 are vertices in the data graph and no subgraph matching result will be generated.

[0068] Step S2: Save the data graph to the GPU;

[0069] In this embodiment, the data graph in the GPU is the data graph before batch update.

[0070] The saving of the data graph to the GPU includes:

[0071] In the GPU, the data graph is saved as a static data graph with a four-layer array structure. The static data graph includes: the first-layer array is a label array for storing the label information of the nodes in the data graph; the second-layer array is a neighbor physical address array for recording the physical address information of the neighbor arrays of each node in the data graph; the third-layer array is a node array for recording each node in the data graph, and the node array includes two elements, namely the number of neighbor nodes and the size of the neighbor array; the fourth-layer array is a neighbor array, and a neighbor array is independently allocated for each node in the data graph.

[0072] Step S3: Save the data graph and the dynamic update sequence to the CPU, and update the data graph using the dynamic update sequence.

[0073] In this embodiment, the flowchart of this embodiment is schematically shown as Figure 6 shown. This embodiment makes full use of the advantages of the CPU and GPU. The GPU is used to perform subgraph matching on the static data graph in parallel, and then the CPU is used to perform subgraph matching on the dynamically changing areas in the data graph. Finally, the subgraph matching results are merged to generate the final match.

[0074] First, it is necessary to construct the static data graph in the GPU. Since the previous GPU parallel subgraph matching algorithms are all for static data graphs, it is difficult to update when applied to the dynamically updated graph in batches. Therefore, this embodiment first proposes a storage structure for the dynamic graph, as Figure 7 shown as a four-layer array storage structure. The first-layer array is the label array, which stores the node label information. The second-layer array is the neighbor physical address array, which records the physical address information of the neighbor array of each node. By accessing the neighbor physical address array, the neighbor array of the node can be found to obtain all neighbor node information of the node. The third-layer array is the node array, and each node contains two elements in the node array, the number of neighbor nodes and the size of the neighbor array. By the number of neighbor nodes and the size of the neighbor array, the usage of the neighbor array can be judged. The fourth-layer array is the neighbor array, and a neighbor array is independently allocated for each node, ensuring the continuous storage of neighbor information. When performing dynamic graph updates, for newly added neighbor nodes, if the neighbor array is full, it is necessary to dynamically allocate neighbor array space, copy the original neighbor nodes to the new neighbor array, and perform the addition operation. For deleting neighbor nodes, the free list is used to dynamically mark the index position of the node to be deleted in the neighbor array to achieve logical deletion.

[0075] For example Figure 5 shown, after the data graph G is dynamically updated, is obtained, and the dynamic update sequence is . For the data node v 1 add a new neighbor node v 2 , for v 2 delete the neighbor node v 3 , add a new neighbor node v 1 , for v 3 delete the neighbor node v 2 . For v 1 node, since the size of the neighbor node array is 1, for adding v 2 , it is necessary to dynamically allocate new neighbor node array space, copy v 0 to the new space, and then add the new neighbor node v 2 .

[0076] For v2 Nodes use an idle list to dynamically record the deletion of v 3 The index position 1 in the neighbor array, etc., for adding new v 1 When adding new v, search for an idle position in the idle list and replace v 3 with v 1 . The specific process is as Figure 7 shown. If the number of newly added neighbors does not exceed the number of the idle list, the overhead caused by space allocation can be avoided.

[0077] For v 3 The processing method is similar to that of v 2 . Store the index position of the neighbor node v 2 in the idle list to complete the logical deletion. The dynamically updated GPU storage structure is as Figure 8 shown. Here, it means that the index position of v 3 is stored in the idle list of v 2 to achieve the logical deletion effect. The storage structure of the updated data graph is as Figure 9 shown.

[0078] Step S4: According to the query graph set, use the data graph in the GPU to perform parallel subgraph matching to obtain the first subgraph matching result, including:

[0079] Step S4.1: Find the candidate set of each node in each element of the query graph set in the static data graph, including:

[0080] Step S4.1.1: Determine the query order in the static data graph according to the ascending order of the function values of the first calculation function;

[0081] In this embodiment, in order to initialize the candidate node set for all query graph nodes, this embodiment uses a heuristic rule to generate a traversal order for the query graph . By observing subgraph matching, it can be known that if the candidate initialization starts from the query node with the fewest candidates, the intermediate result will be the smallest. Since the number of candidates for each node is not known before matching, this paper proposes a calculation function. According to the ascending order of the calculation function values, the query order is determined. According to the query order, according to the following two filtering rules, find the candidate set of each node in the data graph as Figure 10 shown.

[0082] The first calculation function is calculated as follows: ;

[0083] where f(u) is the first calculation function, is the number of nodes in the static data graph that have the same label as a certain query node u in the query graph set, and d(u) is the degree size of a certain query node u in the query graph set.

[0084] Step S4.1.2: According to the query order, based on the first filtering rule and the second filtering rule, find the candidate set of each node in each element of the query graph set in the static data graph.

[0085] The first filtering rule is the label filtering rule. According to the label information of a query node u in the query graph set, generate candidate nodes for the query node u in the static data graph.

[0086] The second filtering rule is the degree filtering rule. According to the degree information of a query node u in the query graph set, generate candidate nodes for the query node u in the static data graph.

[0087] After the first filtering rule and the second filtering rule, obtain the candidate set of each node. The calculation formula is as follows:

[0088] ;

[0089] where C(u) is the candidate set of node u, L(v) is the label of node v, L(u) is the label of query node u, V(G) is the static data graph, d(v) is the degree size of node v, d(u) is the degree size of a query node u in the query graph set, and V(G) is the static data graph.

[0090] Step S4.2: Generate the candidate set of each edge according to the candidate set of each node, including:

[0091] For the query edge , starting from the starting query node u, obtain the candidate node array Cand_array of query node u. Obtain a candidate node v from the candidate node array. Among the neighbor nodes of candidate node v, find the number of candidate nodes that can be used as u'. Form an array Cand_count with the number of candidate nodes of u'. Perform an exclusive prefix process on the array Cand_count to obtain the storage index subscript array of candidate node v' of u' among the neighbor nodes ; Use the candidate node array Cand_array and the index subscript array to store each query edge, and obtain the candidate set of each edge.

[0092] Step S4.3: According to the candidate set of each node and the candidate set of each edge, transform each element in the query graph set into an edge generation tree, including:

[0093] Starting from the query point with the largest degree in each query graph in the query graph set, according to the depth-first traversal strategy, use the query edges visited during the traversal as the paths in the generation tree until all the query edges in the candidate set of each edge are visited, and end the traversal process to obtain an edge generation tree.

[0094] Step S4.4: Split an edge spanning tree into multiple independent query edges;

[0095] Step S4.5: Use GPU to process multiple independent query edges in parallel and find branch candidate results in each independent query edge;

[0096] Step S4.6: According to the branch intersection points of each independent query edge, splice each branch candidate result to obtain the first subgraph matching result.

[0097] In this embodiment, the query graph is transformed into an edge spanning tree, and then each candidate edge is spliced in parallel. Finally, the subgraph matching result on the static graph is generated. Starting from the point with the largest degree in the query graph, according to the depth-first traversal strategy, the query edges visited during the traversal are used as paths in the spanning tree until all query edges are visited, and the traversal process ends. The obtained spanning tree is called an edge spanning tree, denoted as . Figure 2 The spanning tree of the query graph is as Figure 11 shown.

[0098] Using the edge spanning tree of the query graph and the obtained candidate set of query edges, the edges of the spanning tree branches are split into independent query edges, and the query edge candidates are used for connection processing to obtain the candidate set of each branch. For the acquisition of each branch candidate, GPU parallel processing can be used. The candidate nodes at the branch intersections are used for connection to form the matching result set .

[0099] As Figure 11 the edge spanning tree shown, it contains three branches. Branch 1: , Branch 2: and Branch 3: . For Branch 1, it contains two query edges and . Use GPU to process the three branches in parallel to find the branch candidate results. For Branch 1, the candidate results of and need to be connected to obtain the candidate of Branch 1. For the other two branches, only the corresponding edge candidates need to be found. Use the matching nodes at the three branch intersections for branch connection processing to obtain the matching result. The parallel branch splitting process of the edge spanning tree is as Figure 12 shown.

[0100] For Branch 1, the candidate matches of and are respectively and , according to Connect the candidate nodes to obtain candidate for Branch 1 as . For Branch 2, the candidate is , and for Branch 3, the candidates are and . Generate the root node of the tree according to the edges Perform serial connection processing on the candidates of the branches. Connect Branch 1 and Branch 2 to obtain the matching result . According to the candidate, connect with Branch 3 to obtain the final query graph In the query graph G, the matching result is , .

[0101] Step S5: According to the query graph set and the dynamic update sequence, use the data graph in the CPU to perform subgraph matching to obtain the second subgraph matching result, including:

[0102] Step S5.1: Through the maximum hop count of the vertices in the query graphs in the query graph set, find the area affected by the dynamic update sequence, and use the area as the dynamic data subgraph;

[0103] Step S5.2: Determine whether each updated node in the data graph is a candidate node of a certain node in the query graph before and after the update. If so, perform a depth-first traversal of the maximum hop count of the vertices in the corresponding query graph starting from the modified node, and use the subgraph formed by the depth-first traversals of all updated nodes as the incremental data subgraph; if not, skip this updated node and do not perform any operation;

[0104] In this embodiment, through the maximum hop count of the vertices in the query graph, find the area affected by the batch update in the query graph, which is called the dynamic data subgraph. The maximum hop count refers to the maximum number of traversal layers required to traverse the entire query graph starting from a vertex in the query graph. Facing batch updates, we first determine whether each updated node is a candidate node of a certain node in the query graph before and after the update. If so, perform a depth-first traversal of the maximum number of hops of the vertices in its corresponding query graph starting from this node, and use the subgraph formed by the depth-first traversals of all updated nodes as the incremental data subgraph; if not, skip this updated node and do not perform any operation.

[0105] Step S5.3: According to the third filtering rule, the fourth filtering rule, and the fifth filtering rule, find the candidate set of each node in the query graphs in the query graph set in the incremental data subgraph;

[0106] In this embodiment, find the candidate set of each node in the query graph in the incremental data subgraph. In the present invention, according to the following three filtering rules on the incremental data, find the candidate set of each node in the query graph in the data graph.

[0107] The third filtering rule is a label filtering rule. According to the label information of the query nodes in the query graph, candidate nodes are generated for the query nodes in the incremental data subgraph; the calculation formula corresponding to the third filtering rule is the same as that of the first filtering rule.

[0108] The fourth filtering rule is a degree filtering rule. According to the degree information of the query nodes in the query graph, candidate nodes are generated for the query nodes in the incremental data subgraph; the calculation formula corresponding to the fourth filtering rule is the same as that of the second filtering rule.

[0109] The fifth filtering rule is a neighbor label frequency filtering rule, which uses the neighbor nodes of the query nodes in the query graph to filter in the following way: for a given node , if there exists a label , then there exists ;

[0110] where , , if there does not exist a label , then the node v is filtered out, where u is the query node in the query graph, N(u) is the neighbor node of u, v is the node in the data graph, C(u) is the matching candidate set of u, L(N(u)) is the label of the neighbor nodes of u, |N(v, l)| is the number of nodes with label l among the neighbor nodes of v, |N(u, l)| is the number of nodes with label l among the neighbor nodes of u, u' is another query node in the query graph, and L(N(u')) is the label of the neighbor nodes of u'.

[0111] Step S5.4: Determine the query order for the matching of the incremental data subgraph according to the ascending order of the function values of the second calculation function;

[0112] In this embodiment, the query order for subgraph matching is determined. An efficient matching order can reduce redundant calculations and relieve the computing pressure. In this embodiment, a calculation function, that is, the second calculation function, is designed, and the calculation formula is as follows: f'(u) = |C'(u) / d(u)|;

[0113] where f'(u) is the second calculation function, C'(u) is the candidate set of each node in the filtered query graph, and d(u) is the degree size of a certain query node u in the query graph set.

[0114] Step S5.5: Establish an auxiliary storage structure for the candidate set of each node;

[0115] In this embodiment, an auxiliary storage structure is established for the candidate set of each node. The auxiliary structure is a set of data vertices that find the matching relationship of each query vertex in the data graph after filtering, and records whether there is an edge between these data nodes, and these data are obtained through the filtering rules. To speed up the enumeration process, this embodiment establishes an auxiliary index structure. The auxiliary index structure is a tree structure, which is a commonly used structure in the CPU subgraph matching field to accelerate the verification process. The nodes in the auxiliary index structure represent the matching candidate sets in the data graph corresponding to the nodes in the query graph, and the edges in the auxiliary index structure represent whether there is a connected edge in the data graph between the nodes in the candidate sets of different query vertices.

[0116] Step S5.6: According to the auxiliary storage structure and the backtracking algorithm, find all subgraphs in the dynamic data subgraph that are isomorphic to the query graph, and obtain the second subgraph matching result.

[0117] In this embodiment, the backtracking algorithm starts from a node of the target subgraph, attempts to map it to a node of the main graph, and recursively finds a suitable mapping for other nodes of the target subgraph, while checking whether the mapping of the edges meets the structural requirements. If it is found that the mapping is illegal at a certain step, the algorithm will backtrack to the previous step and try other possible mappings until a valid mapping is found or all possible mappings are traversed. After constructing the auxiliary storage structure, the backtracking algorithm can quickly find the matching candidate set corresponding to the next vertex to be recursively called through the auxiliary structure. Enumerate and verify to find the subgraph in the dynamic data subgraph that is isomorphic to the query graph. According to the auxiliary structure and the backtracking algorithm, find all subgraphs in the dynamic data subgraph that are isomorphic to the query graph.

[0118] Step S6: Merge the first subgraph matching result and the second subgraph matching result to obtain the final subgraph matching result, and update the data graph in the GPU according to the dynamic update sequence.

[0119] And updating the data graph in the GPU according to the changed part in the data graph includes: adding neighbor nodes to the storage structure of the static data graph or deleting nodes from the storage structure of the static data graph;

[0120] The process of adding neighbor nodes to the storage structure of the static data graph is as follows:

[0121] If the neighbor array is full, it is necessary to dynamically allocate space for the neighbor array, copy the original neighbor nodes to the neighbor array after the newly added space, and perform the operation of adding neighbor nodes.

[0122] The process of deleting nodes from the storage structure of the static data graph is as follows:

[0123] When deleting a node in the storage structure of a static data graph, in the corresponding position of the neighbor array, the free list is used to dynamically mark the node to be deleted.

[0124] The merging of the first subgraph matching result and the second subgraph matching result includes:

[0125] Step S6.1: Filter the subgraph matching results of the GPU. Since some of the matching results found by the GPU have changed due to batch updates, all subgraph matching results in the GPU subgraph matching results that contain the nodes involved in the batch updates are removed.

[0126] Step S6.2: Merge the subgraph matching results of the CPU and the filtered subgraph matching results of the GPU.

[0127] Step S6.3: Update the storage of the data graph in the GPU. After the subgraph matching is completed, the CPU generates the storage structure of the data graph after batch updates and passes it to the GPU.

[0128] In this embodiment, after merging the subgraph matching results of the CPU and the filtered subgraph matching results of the GPU, the storage of the data graph in the GPU is updated.

[0129] The embodiments in this application are all described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. The key points of each embodiment are the differences from other embodiments.

[0130] The protection scope of this application is not limited to the above embodiments. Obviously, those skilled in the art can make various changes and deformations to the present disclosure without departing from the scope and spirit of the present disclosure. If these changes and deformations belong to the scope of the claims of the present disclosure and their equivalent technologies, the intention of the present disclosure also includes these changes and deformations.

Claims

1. A heterogeneous collaborative subgraph matching method for dynamic graphs, characterized in that: include: Get data graphs, query graph collections, and dynamically update sequences; Save the data graph to the GPU; The data graph and the dynamic update sequence are saved in the CPU, and the data graph is updated using the dynamic update sequence; According to the query graph set, use the data graph in the GPU to perform parallel subgraph matching to obtain a first subgraph matching result; According to the query graph set and the dynamic update sequence, subgraph matching is performed using the data graph in the CPU to obtain a second subgraph matching result; Merging the first subgraph matching result with the second subgraph matching result to obtain a final subgraph matching result, and updating the data graph in the GPU according to the dynamic update sequence; The method of performing parallel subgraph matching using the data graph in the GPU according to the query graph set to obtain a first subgraph matching result includes: Find the candidate set for each node in each element of the query graph set in the static data graph; Generate a candidate set for each edge based on the candidate set of each node; According to the candidate set of each node and the candidate set of each edge, each element in the query graph set is converted into an edge spanning tree; Split an edge spanning tree into multiple independent query edges; Use GPU to process multiple independent query edges in parallel and find branch candidate results in each independent query edge; According to the branch intersection of each independent query edge, each branch candidate result is spliced ​​to obtain the first subgraph matching result; The method of performing subgraph matching using the data graph in the CPU according to the query graph set and the dynamic update sequence to obtain a second subgraph matching result includes: Find the area affected by the dynamic update sequence update by querying the maximum hop number of the vertices in the query graph in the query graph set, and use the area as the dynamic data subgraph; Determine whether each updated node in the data graph is a candidate node of a node in the query graph before and after the update. If so, perform a depth-first traversal of the maximum number of hops of the corresponding vertex in the query graph starting from the modified node, and use the subgraph composed of the depth-first traversal of all updated nodes as the incremental data subgraph; if not, skip this update node and do not perform any operation; According to the third filtering rule, the fourth filtering rule and the fifth filtering rule, finding a candidate set for each node in the query graph in the query graph set in the incremental data subgraph; Determine the query order of the incremental data subgraph matching according to the function value of the second calculation function in ascending order; Establishing an auxiliary storage structure for each candidate set of nodes, wherein the auxiliary storage structure is a set of data vertices that have a matching relationship in the data graph for each query vertex after filtering and records whether there are edges between the data vertices; According to the auxiliary storage structure and the backtracking algorithm, all subgraphs in the dynamic data subgraph that are isomorphic to the query graph are found, and the second subgraph matching result is obtained.

2. A heterogeneous collaborative subgraph matching method for dynamic graphs according to claim 1, characterized in that: Saving the data graph to the GPU includes: In the GPU, the data graph is saved as a static data graph with a four-layer array structure, and the static data graph includes: the first layer array is a label array, which is used to store the label information of the nodes in the data graph; the second layer array is a neighbor physical address array, which is used to record the physical address information of the neighbor array of each node in the data graph; the third layer array is a node array, which is used to record each node in the data graph, and the node array includes two elements, namely the number of neighbor nodes and the size of the neighbor array; the fourth layer array is a neighbor array, and a neighbor array is independently allocated to each node in the data graph.

3. The method for matching heterogeneous collaborative subgraphs for dynamic graphs according to claim 1, characterized in that: The step of finding a candidate set of each node in each element of the query graph set in the static data graph includes: Determine the query order in the static data graph according to the order of the function values ​​of the first calculation function from small to large; According to the query order, based on the first filtering rule and the second filtering rule, a candidate set of each node in each element of the query graph set is found in the static data graph.

4. The method for matching heterogeneous collaborative subgraphs for dynamic graphs according to claim 3, characterized in that: The first calculation function is calculated as follows: ; Among them, f(u) is the first calculation function, is the number of nodes in the static data graph that have the same label as a query node u in the query graph set, and d(u) is the degree of a query node u in the query graph set.

5. The method for heterogeneous collaborative subgraph matching for dynamic graphs according to claim 3, characterized in that: The first filtering rule is a label filtering rule, which generates a candidate node for a query node u in the static data graph according to the label information of a query node u in the query graph set; The second filtering rule is a degree filtering rule, which generates a candidate node for the query node u in the static data graph according to the degree information of a query node u in the query graph set.

6. The method for matching heterogeneous collaborative subgraphs for dynamic graphs according to claim 1, characterized in that: The step of generating a candidate set for each edge according to the candidate set for each node includes: For query edge , starting from the starting query node u, get the candidate node array Cand_array of the query node u, get a candidate node v in the candidate node array, and find the neighbor nodes of the candidate node v In the example, find the number of candidate nodes that can be used as u', form an array Cand_count with the number of candidate nodes for u', perform exclusive prefix processing on the array Cand_count, and obtain the neighbor nodes The storage index subscript array of the candidate node v' of u'; use the candidate node array Cand_array and the index subscript array to store each query edge to obtain the candidate set of each edge.

7. The method for matching heterogeneous collaborative subgraphs for dynamic graphs according to claim 1, characterized in that: The step of converting each element in the query graph set into an edge spanning tree according to the candidate set of each node and the candidate set of each edge includes: Starting from the query point with the largest degree in each query graph in the query graph set, according to the depth-first traversal strategy, the query edges visited in the traversal are used as the paths in the spanning tree until all the query edges in the candidate set of each edge are visited, the traversal process is terminated, and an edge spanning tree is obtained.

8. The method for matching heterogeneous collaborative subgraphs for dynamic graphs according to claim 1, characterized in that: The third filtering rule is a label filtering rule, which generates candidate nodes for the query nodes in the incremental data subgraph according to the label information of the query nodes in the query graph; The fourth filtering rule is a degree filtering rule, which generates candidate nodes for the query nodes in the incremental data subgraph according to the degree information of the query nodes in the query graph; The fifth filtering rule is a neighbor label frequency filtering rule, which uses the neighbor nodes of the query node in the query graph. Filter by: Given a node If there is a tag , then there exists ; in , If a tag does not exist , then filter out the node v, where u is a query node in the query graph, N(u) is u's neighbor node, v is a node in the data graph, C(u) is u's matching candidate set, L(N(u)) is the label of u's neighbor node, |N(v,l)| is the number of nodes with label l in v's neighbor node, |N(u,l)| is the number of nodes with label l in u's neighbor node, u' is another query node in the query graph, and L(N(u')) is the label of u''s neighbor node; The second calculation function is calculated as follows: f'(u)=|C'(u) / d(u)|; Among them, f'(u) is the second calculation function, C'(u) is the candidate set of each node in the filtered query graph, and d(u) is the degree of a query node u in the query graph set.

Citation Information

Patent Citations

  • GPU axis subgraph matching method based on coding tree

    CN113204552A

  • Method, device and equipment for data search

    CN115408427A

  • Dynamic graph increment matching method and device for decomposition sorting on large-scale query graph

    CN115757846A

  • Large graph sub-graph matching method and system based on multiple GPUs (Graphics Processing Unit)

    CN115827924A

  • Dynamic graph increment subgraph matching method based on neighborhood security compression

    CN116484062A