Sub-graph inclusion prediction method based on multi-hop sub-graph matching consistency

By introducing multi-hop consistency and MSMC models into sub-graph inclusion relation prediction, the wrong prediction problem caused by relaxing topological constraints in the prior art is solved, efficient and accurate sub-graph inclusion relation prediction is achieved, and the model is improved in the migration and applicability.

CN120011740APending Publication Date: 2025-05-16NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411973077.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-30
Publication Date
2025-05-16

AI Technical Summary

Technical Problem

The prior art has problems with mispredictions in sub-graph inclusion relational prediction, especially due to the selective limitations and high computational overhead caused by relaxing topological constraints.

Method used

A subgraph inclusion prediction method based on the matching consistency of multi-hop subgraphs is proposed. By constructing an MSMC model, using the neighborhoods of different hop numbers of multi-layer GNN encoding nodes, combined with an MLP projector, the encoding results are projected into the embedding space, and the multi-hop consistency overhead is calculated to predict the subgraph inclusion relationship.

Benefits of technology

It effectively reduces the error prediction caused by relaxing topological constraints, so that it is possible to determine whether the data graph contains a query graph without testing the isomorphism of the subgraph, and the convergence speed is faster, with good migration and applicability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011740A_ABST
    Figure CN120011740A_ABST
Patent Text Reader

Abstract

The invention discloses a sub-graph inclusion prediction method based on multi-hop sub-graph matching consistency, and the method comprises the steps: 1, model construction: constructing an MSMC model, the MSMC model comprises a coding module, a mapping module and a prediction module, the coding module comprises a plurality of layers of GNNs for coding neighborhoods of different hop counts of each node in an input graph, the mapping module comprises an MLP projector, and the prediction module comprises an MLP projector; the projection module is used for projecting representation of each layer of a coding result to an embedding space, and the prediction module is used for calculating multi-hop consistency overhead between nodes and predicting sub-graph inclusion relations; 2, training the MSMC model by using the training data set, and obtaining a sub-graph inclusion prediction model after the training is completed; and step 3, obtaining a to-be-predicted input graph pair, inputting the to-be-predicted input graph pair into the sub-graph inclusion prediction model, and predicting a sub-graph inclusion relation. The method has the advantages of being simple in implementation method, small in calculation overhead, high in prediction efficiency and precision, high in mobility and applicability and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of graph structure data retrieval, and in particular to a subgraph inclusion prediction method based on multi-hop subgraph matching consistency. Background Art

[0002] Graph structured data can describe the interactions between entities and is therefore widely used in various practical applications such as network analysis, bioinformatics, knowledge graphs, etc. Subgraph search is one of the most basic query and mining operations in the above applications. Given a data graph D and a query graph Q, the task of subgraph search is to determine whether D contains Q, that is, the subgraph containment relationship problem.

[0003] Methods for determining subgraph inclusion can be divided into three main categories: (1) Subgraph matching algorithms, which aim to find all bijective mappings from the query graph to the subgraphs of the data graph. This imposes strict constraints on the graph topology and labels, and this problem is considered to be NP-complete. (2) Feature-based methods, in order to improve the search speed, do not directly test the subgraph isomorphism condition on the entire graph. Instead, they decompose the query graph and the data graph into a set of smaller features, namely subgraphs, and then predict the inclusion of subgraphs by comparing the number of features. (3) Neighborhood-based methods, which bypass the subgraph isomorphism test by relaxing the topological constraints; they aggregate the neighborhood information of each node and then match nodes based on the proximity of the neighborhood, thereby limiting the graph comparison to a reasonable scale. (4) Match nodes by the label distribution of the k-hop neighborhood and match by maximizing the sum of node matches. Feature-based and neighborhood-based methods convert the graph into multidimensional vectors, which are then stored as indices for subsequent tasks. However, the above methods predict subgraph inclusion without finding specific matches. Their goal is to prune inappropriate candidates as much as possible to reduce the search space. Therefore, it is necessary to further use the subgraph matching algorithm to enumerate all matching results.

[0004] In the prior art, researchers have proposed a neighborhood-based method that attempts to use neural networks to predict subgraph inclusion relations. For example, subgraph relations are considered as partial order relations, or node-induced subgraphs are projected into a continuous feature space, and then the features of all query-data substructure pairs are compared, or edge-aligned subgraph matching solutions based on the neighborhood information of each edge are used for graph retrieval, which estimates the edge distance between the query graph and a series of corpus graphs by learning an optimal edge alignment matrix, and makes predictions based on the distance score. The above neighborhood-based methods match graphs by matching nodes. The key idea is that if a data node u matches a query node v, then u's k-hop neighborhood must contain more information than v's k-hop neighborhood. Any node pair that meets this observation is considered a potential match, and such a relationship should also be preserved in the embedding space; therefore, in the embedded representation, u's embedding value in each dimension is greater than v's embedding value. Although the above neighborhood-based methods can successfully perform subgraph searches in large networks with up to 10 million nodes and perform well in subgraph inclusion relationship prediction or graph retrieval tasks, relaxing topological constraints may lead to incorrect predictions, thus there will be selectivity limitations. In addition, for a connected graph with N vertices, the number of edges ranges from N-1 to N(N-1) / 2. Compared with the node-aligned method, the edge-aligned method will also incur greater computational overhead. Summary of the invention

[0005] The technical problem to be solved by the present invention is: in response to the problems existing in the prior art, the present invention provides a subgraph inclusion prediction method based on multi-hop subgraph matching consistency, which is simple to implement, has low computational overhead, high prediction efficiency and accuracy, and strong portability and applicability. It can effectively reduce the erroneous predictions caused by relaxing topological constraints, so that it is possible to determine whether a data graph contains a query graph without testing subgraph isomorphism.

[0006] In order to solve the above technical problems, the technical solution proposed by the present invention is:

[0007] A subgraph inclusion prediction method based on multi-hop subgraph matching consistency, comprising the following steps:

[0008] Step 1: Model construction: Construct an MSMC model, which includes an encoding module, a mapping module and a prediction module. The encoding module includes a multi-layer GNN, which is used to encode the neighborhoods of different hops of each node in the input graph pair respectively so as to meet the multi-hop internal consistency. The mapping module includes an MLP projector, which is used to project the representation of each layer in the encoding result to the embedding space to convert each node in the input graph pair into a series of multi-hop embeddings. The prediction module is used to calculate the multi-hop consistency cost between the nodes of the input graph pair according to the projection result output by the mapping module to measure the relative importance of each hop, and predict the subgraph inclusion relationship between the input graph pairs according to the calculated multi-hop consistency cost;

[0009] Step 2: Model training: Use multiple groups of positive examples and negative examples to form a training data set and train the MSMC model, wherein the positive examples are sample graph pairs for which multi-hop consistency is established, that is, the sample graph pairs satisfy the subgraph inclusion relationship, and the negative examples are sample graph pairs for which multi-hop consistency is not established, that is, the sample graph pairs do not satisfy the subgraph inclusion relationship. After the training is completed, a subgraph inclusion prediction model is obtained;

[0010] Step 3: Subgraph inclusion prediction: obtain the input graph pair to be predicted, input it into the subgraph inclusion prediction model, and predict the subgraph inclusion relationship between the input graph pairs.

[0011] Further, judging whether multi-hop consistency holds includes: for each node in the data graph And query each node in the graph If (v,u) matches, that is, for any k, Then determine each isomorphic to The subgraph of the input graph, that is, the multi-hop consistency between the data graph D and the query graph Q holds, otherwise the multi-hop consistency does not hold. represent the nodes in the input graph pair (D, Q), represents the set of nodes in the data graph D in the input graph pair, represents the set of nodes in the query graph Q in the input graph pair, φ is the embedding function used to convert the graph into a multidimensional vector, and is the k-hop subgraph induced by v and u, i.e., the k-hop neighborhood of the corresponding node.

[0012] Furthermore, the multi-hop internal consistency is: for the input connected graph G, if for each In each hop, v Consistently subgraph, then the connected graph G is determined to satisfy the multi-hop internal consistency.

[0013] Furthermore, the encoding module learns the attributes of each node in the input graph and aggregates information from the neighborhood of each node through each layer of the GNN, and adds the aggregated neighborhood information to the embedding output of the previous layer to update the k-hop embedding of this layer. The model stores and reuses the intermediate output as the embedding of the neighborhood with different hops at each layer, expressed as:

[0014]

[0015] Among them, Encoder i (v) represents the encoding result of node v, Ψ is the update function, □ is the aggregation function, yes The embedded representation of is the information aggregated from the neighbors of node v.

[0016] Furthermore, a multi-hop consistency cost function is used in the prediction module to calculate the multi-hop consistency cost between the nodes of the input graph, and the calculation expression of the multi-hop consistency cost function is:

[0017]

[0018] in, is the multi-hop consistency overhead between node v and node u, represent the nodes in the input graph pair (D, Q), represents the set of nodes in the data graph D in the input graph pair, represents the set of nodes in the query graph Q in the input graph pair, are learnable weights, represents the cost of each hop, φ(·) represents the After encoding by the encoding module, multi-hop embedding is obtained, and then the projector MLP is used to map the representation of each layer to the embedding function of the embedding space.

[0019] Furthermore, in step 2, the maximum boundary loss function is used as the objective function for training during the training process. The calculation expression of the objective function is:

[0020]

[0021] in, is a learnable boundary, E + is a set of positive examples, E - is a set of negative examples.

[0022] Furthermore, the subgraph inclusion relationship between the predicted input graph pairs includes: if it is determined that each query node in the query graph Q matches at least one data node in the data graph D, then the data graph D is considered to contain the query graph Q; otherwise, it is determined that the data graph D does not contain the query graph Q.

[0023] A subgraph inclusion prediction device based on multi-hop subgraph matching consistency, comprising:

[0024] A model building module, used to build an MSMC model, the MSMC model includes an encoding module, a mapping module and a prediction module, the encoding module includes a multi-layer GNN, which is used to respectively encode the neighborhoods of different hops of each node in the input graph pair so that they conform to multi-hop internal consistency, the mapping module includes an MLP projector, which is used to project the representation of each layer in the encoding result to the embedding space, so as to convert each node in the input graph pair into a series of multi-hop embeddings, the prediction module is used to calculate the multi-hop consistency cost between the nodes of the input graph pair according to the projection result output by the mapping module to measure the relative importance of each hop, and predict the subgraph inclusion relationship between the input graph pairs according to the calculated multi-hop consistency cost;

[0025] A model training module is used to use multiple groups of positive examples and negative examples to form a training data set and train the MSMC model, wherein the positive examples are sample graph pairs for which multi-hop consistency is established, that is, the sample graph pairs satisfy the subgraph inclusion relationship, and the negative examples are sample graph pairs for which multi-hop consistency is not established, that is, the sample graph pairs do not satisfy the subgraph inclusion relationship. After the training is completed, a subgraph inclusion prediction model is obtained;

[0026] The subgraph inclusion prediction module is used to obtain the input graph pairs to be predicted, input them into the subgraph inclusion prediction model, and predict the subgraph inclusion relationship between the input graph pairs.

[0027] A computer device comprises a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to perform the above method.

[0028] A computer-readable storage medium storing a computer program, wherein the computer program implements the above method when executed by a processor.

[0029] Compared with the prior art, the advantages of the present invention are: the present invention adopts a multi-hop subgraph matching consistency MSMC method to predict subgraph inclusion relations, predicts subgraph inclusion by quantifying the multi-hop consistency overhead of graph pairs, and implements subgraph inclusion prediction using a neighborhood-based subgraph inclusion relationship measurement method, which can ensure that the inclusion of node-induced subgraph pairs within and between graphs remains consistent at each hop, effectively reducing erroneous predictions caused by relaxing topological constraints, so that it is possible to determine whether a data graph contains a query graph without testing subgraph isomorphism, and compared with traditional neural network-based prediction methods, it has a faster convergence speed and good portability and applicability. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] Figure 1 It is a schematic diagram of error prediction using a traditional method when nodes do not match in a specific application embodiment.

[0031] Figure 2 is another example schematic diagram of error prediction in a specific application embodiment.

[0032] Figure 3 It is a schematic diagram of some cases that meet the definition of Lemma 1 in a specific application embodiment.

[0033] Figure 4 It is a schematic diagram of the principle of the multi-hop consistency and MSMC model adopted in this embodiment.

[0034] Figure 5 It is a schematic diagram of the performance comparison results of various methods in neighborhood-based inclusion prediction in a specific application embodiment of the present invention.

[0035] Figure 6 It is a schematic diagram of another part of the performance comparison results of various methods on neighborhood-based inclusion prediction in a specific application embodiment of the present invention.

[0036] Figure 7 It is a schematic diagram of the performance results of different variants of MSMC on six real-world datasets in specific application examples.

[0037] Figure 8 It is a schematic diagram of the results of the subgraph alignment application in the experiment.

[0038] Fig. 9 It is a schematic diagram of the pruning ability results in the experiment. DETAILED DESCRIPTION

[0039] The present invention is further described below in conjunction with the accompanying drawings and specific preferred embodiments, but the protection scope of the present invention is not limited thereby.

[0040] The existing subgraph matching technology based on neighborhood information uses an inductive bias method, that is, the subgraph induced by the matched data graph node must contain the corresponding part of the query graph node. Any node pair that meets this characteristic is considered a potential match. However, this characteristic relaxes the strict topological constraints of subgraph isomorphism and is therefore insufficient to effectively eliminate false predictions. Figure 1 As shown in the figure, nodes v and u do not match, while the number of nodes contained in their 3-hop neighborhoods are [green: 2, blue: 3, apricot: 1] and [green: 5, blue: 4 apricot: 8] respectively. The embedding value of node u in each dimension is greater than that of node v. If the traditional subgraph inclusion prediction method is used, it will be incorrectly predicted as a match.

[0041] In order to solve the above problems, the present invention proposes a Multi-hop Subgraph Matching Consistency (MSMC) method based on neighborhood subgraph distance metric to achieve subgraph inclusion prediction. Different from the traditional method that only compares node embeddings on k-hops, the MSMC method of the present invention requires that the intra-graph and inter-graph inclusion of node embeddings on each hop be consistent, namely: (1) Multi-hop inter-graph consistency: If a data node u matches a query node v, then the k-hop subgraph induced by u Should be included in each hop (2) Multi-hop graph consistency: For any node in graph G, G k G should be included in each hop (k-1) The present invention uses the above multi-hop consistency as an inductive bias design to implement the MSMC model, and uses the model to convert a series of hop-based neighborhoods of a node into Projecting into an embedding space that maintains multi-hop consistency, the distance between nodes in the embedding space can reflect their actual inclusion relationship, and then by calculating the multi-hop consistency overhead between the nodes of the input graph pair to measure the relative importance of each hop, the above-mentioned neighborhood-based subgraph inclusion relationship is effectively measured, thereby realizing the prediction of the subgraph inclusion relationship between the input graph pairs. At the same time, the present invention can avoid costly subgraph isomorphism testing and improve the efficiency and accuracy of prediction.

[0042] To facilitate understanding, the relevant technical background of the present invention is first introduced by way of examples.

[0043] A connected graph can be represented as in, is a set of vertices, and the label of each vertex v is recorded as ε G is a set of edges. Given a data graph D, D′ is a subgraph of D that satisfies and and For each node, there is l D′(v) = l D (v). represents the k-hop subgraph induced by v, i.e., the hop-based neighborhood of v.

[0044] Subgraph isomorphism: If there exists a bijective function So that (1) l Q (v) = l D′ (f(v)), and (2) e(f(v),f(v′))∈ε D′ , then Q is isomorphic to the subgraph D′ of D. f is called an isomorphism mapping, (Q,D′) is a graph isomorphism, and (Q,D) is a subgraph isomorphism.

[0045] Subgraph inclusion relation: Given Q and D, the subgraph inclusion relation is used to determine whether there is a mapping from Q to D as defined above. Unlike subgraph matching, the subgraph inclusion relation checks the existence of a mapping without necessarily finding all mappings.

[0046] The feature and neighborhood-based subgraph inclusion relationship prediction method can usually be decomposed into two steps: (1) using the defined embedding function The graph is converted into a multidimensional vector, where □ is a D-dimensional embedding space, which can be continuous or discrete; (2) The vector representation of the graph predicts the subgraph inclusion relationship. Feature-based methods generate embedded representations by calculating features (i.e., subgraphs), while neighborhood-based methods aggregate the neighborhood information of each node in the graph through the shortest path between nodes. Neighborhood information can be regarded as a special type of feature and is not topologically constrained; therefore, if D contains Q, it must contain all the features of Q, and their embeddings can reflect this.

[0047] Subgraph inclusion relation prediction: If D contains Q, then the vector output by the embedding function φ(·) satisfies φ(D) in all dimensions d. d ≥φ(Q) d .

[0048] According to the above definition, the topological constraints of subgraph isomorphism of traditional neighborhood-based methods can be further relaxed to match the following properties of node pairs:

[0049] Property 1. If D contains Q, then for and If G u Include G in k hops v , then φ(G u )d≥φ(G v )d, which indicates that u and v are potential matches.

[0050] Subgraph inclusion cost function: If nodes v and u match, their embeddings should satisfy Property 1 to reflect the subgraph inclusion relationship in the embedding space. Therefore, given a node pair and And the embedding function φ(·), the cost function of the traditional neighborhood-based subgraph inclusion relationship prediction is:

[0051]

[0052] Among them, φ(G v ) d >φ(G u ) d This means that in the dth dimension G v Not included in G u , which contradicts Property 1, so v and u are not matched and such node pairs should be penalized. This cost function can be extended to edge-induced subgraphs. Based on such a function, existing methods eliminate inappropriate node pairs and then predict subgraph inclusion relations based on the sum of node matches, node votes, or graph transfer costs.

[0053] Misprediction: Neighborhood-based methods can circumvent the expensive subgraph isomorphism test by relaxing the topological constraints of subgraph isomorphism; these methods generate node embeddings by aggregating information from the node’s k-hop neighborhood. For a node pair (v,u), if their embeddings satisfy Property 1, i.e., φ(G u ) d ≥φ(G v ) d , then it is considered as a potential match. However, relaxing the constraints also brings negative effects: (1) These methods only capture the neighborhood distribution of nodes; for example, Figure 2 Although (v,u) is not a match, And when considering their 2n-1 hop neighborhood distribution (i.e., odd number of hops), v is easily matched with u; (2) When k increases, nodes in larger subgraphs naturally have more neighbors than nodes in smaller subgraphs. Therefore, larger subgraphs can gather more information accordingly. Even if they do not match, their embeddings may meet Property 1, such as Figure 1 As shown. The simplest solution to the above problem is to find a suitable k; however, the properties of graphs vary, and it is difficult to find such a k for all possible cases. Although the application of propagation factors, sampling and normalization methods can alleviate the former situation, it is difficult to make up for the latter problem. The above shortcomings limit the selectivity of the neighborhood-based methods in the prior art and hinder many applications that require high accuracy.

[0054] Accurate subgraph inclusion prediction is very important in subgraph search, which can significantly reduce the search space and thus improve the search speed. In order to reduce the prediction errors caused by relaxing the topological constraints of subgraph isomorphism, the present invention proposes multi-hop consistency by mining the mutual relationship between subgraph isomorphic node pairs.

[0055] Any subgraph isomorphism pair meets the following consistency: (i) Multi-hop internal consistency: Given a connected graph G, for each In each hop, v Consistently ii) Multi-hop external consistency: Given Q is isomorphic to a subgraph in D, for each and each If (v,u) matches, then isomorphic to A subgraph of . Figure 4 The above multi-hop consistency is demonstrated in .

[0056] This example further proposes Lemma 1: as well as And is not equal to v, the length of the shortest path between (v,u) in D is less than or equal to the shortest path length of (v,u) in D′, that is, for v∈D′, the distance to its neighbor in D′ may be farther than in D. Figure 3 Some of the cases that satisfy the above definition are shown.

[0057] To prove the correctness of multi-hop consistency, assume that the query graph Q is a connected graph. According to the definition, v is the set of nodes whose distance v is less than or equal to k. induced subgraph, so for any node in Q, multi-hop internal consistency naturally holds. If Q is isomorphic to D′, they should share equivalent topological structures and label distributions. According to Lemma 1 above, v’s neighbors may be farther away from v in D′ than in D, so v’s neighbors that were originally in the k-hop neighborhood in D may appear in the (k+i)-hop neighborhood in D′, where i∈N + ; Therefore, the missing neighbors in D′ will make the k-hop subgraph of v in D′ a subgraph of the k-hop subgraph of v in D, which can be extended to Q, thus proving that multi-hop external consistency also holds.

[0058] In this embodiment, when performing subgraph inclusion prediction based on multi-hop consistency, assuming that a given graph G, Include For any k, it should satisfy Assume that D contains Q, and for any match And any k should satisfy

[0059] Because: (1) When k is large enough, the query node v When the entire query Q is covered, determine The subgraph inclusion relation in D is equivalent to predicting the subgraph inclusion relation of Q in D. (2) For small k, this consistency degenerates to predicting k-hop common subgraphs; however, it can still be used as a powerful graph similarity measure. Therefore, by finding a match for each v in D, the proposed MSMC can predict subgraph inclusion.

[0060] Assume that given a node pair (v,u), where and Using multi-hop consistency, we can predict the subgraph inclusion of the induced subgraph at each hop. Figure 1 For example, nodes v and u do not match, however, their representations generated by 3-hop aggregation meet attribute 1, so (v,u) will be predicted as a match. The metric in this embodiment can detect the situation that does not meet the consistency of subgraph inclusion at 1-hop, because the label counts of v and u are [green: 1, blue: 1] and [green: 1, apricot: 1] respectively. In this way, this embodiment uses multi-hop consistency to filter out false positives that attribute 1 cannot detect. In theory, when k is large enough, Cover the entire Q, then judge The inclusion of a subgraph in D is equivalent to the prediction of the inclusion of Q in D. Otherwise, multi-hop consistency degenerates into determining k-hop common subgraphs, which is still a powerful graph similarity metric. This embodiment further proposes a solution when k is not large enough.

[0061] For a given data graph D and query graph Q, multi-hop consistency can be used to predict the subgraph inclusion of node pairs (v,u), where and If every query node matches at least one data node, D is considered to contain Q; otherwise, it can be safely judged that D does not contain Q without testing subgraph isomorphism.

[0062] Multi-hop consistency as inductive bias: Compared with property 1, multi-hop consistency constructs a subgraph inclusion space with two strong constraints, namely multi-hop internal consistency and multi-hop external consistency. The prediction of the inclusion relationship of the induced subgraph of v and u is not only affected by their k-hop subgraph embedding, but also by their (k-1)-hop subgraph embedding. Violation of multi-hop consistency under any number of hops will cause v and u to not become a matching node pair. Therefore, by using intra-graph consistency and inter-graph consistency as inductive biases, it helps the model MSMC to converge quickly and achieve higher accuracy and good generalization ability on less training data.

[0063] like Figure 1 As shown, the steps of the subgraph inclusion prediction method based on multi-hop subgraph matching consistency proposed in this embodiment include:

[0064] Step 1: Model construction: Construct the MSMC model, which includes the encoding module, mapping module and prediction module. The encoding module is a multi-layer GNN that encodes the neighborhood of each node in the input graph with different hops based on multi-hop consistency. In order to make the encoding result conform to the multi-hop internal consistency, the mapping module is an MLP projector, which is used to convert the encoding result The representation of each layer in is projected into the embedding space to transform each node in the input graph pair into a series of multi-hop embeddings M = {M 1 , ..., M k The prediction module is used to calculate the multi-hop consistency cost between the nodes of the input graph according to the projection results output by the mapping module to measure the relative importance of each hop, and predict the subgraph inclusion relationship between the input graphs according to the calculated multi-hop consistency cost.

[0065] Specifically, the processing flow of the MSMC model in this embodiment consists of two stages: encoding stage and projection stage. In the first stage, MSMC encodes the neighborhoods of each node with different hop counts in the input graph. To meet the multi-hop internal consistency; the second stage projects the encoding results into the embedding space, in which the possibility of the subgraph inclusion relationship of the node pair is quantitatively calculated by verifying the multi-hop external consistency.

[0066] like Figure 4 As shown in Figure 1, the left side shows a multi-hop consistency example and the right side shows the MSMC model. The MSMC model consists of a k-layer GNN and an MLP, which transforms each node in the input graph into a series of multi-hop embeddings. The intermediate output of each layer in the model is regarded as a representation of a subgraph of the input graph. In this example, Q is contained in D, so for a matching node pair (v,u), the multi-hop consistency cost calculated according to equation (3) should be equal to 0.

[0067] In order to balance accuracy and speed, the MSMC model in this embodiment learns the attributes of each node in the input graph through each layer in the GNN, aggregates the information from the neighborhood of each node, and adds the aggregated neighborhood information to the embedding output of the previous layer to update the k-hop embedding, and stores and reuses the intermediate output as the embedding of the neighborhood with different hop numbers in each layer, that is:

[0068]

[0069] Among them, Encoder i (v) represents the encoding result of node v, Ψ is the update function, □ is the aggregation function, yes The expression, is the information aggregated from the neighbors of node v.

[0070] To ensure multi-hop internal consistency, this embodiment uses a sum function when aggregating and updating so that scale information is not lost when aggregating neighborhoods, and updates the k-hop embedding by adding the aggregated neighborhood information to the embedding output of the previous layer.

[0071] In the MSMC model of this embodiment, at the end of encoding, a multi-scale multi-hop representation of each node is obtained, that is, a multi-hop embedding Then a projector MLP is used to map the representation of each layer to the embedding space, and the entire model can be viewed as an embedding function φ(·).

[0072] Assume D contains Q, and and Matching, then it should always be satisfied at each hop number i and each dimension d Based on equation (1), this embodiment formalizes the multi-hop consistency cost function as follows:

[0073]

[0074] in, is the multi-hop consistency overhead between node v and node u, represent the nodes in the input graph pair (D, Q), represents the set of nodes in the data graph D in the input graph pair, represents the set of nodes in the query graph Q in the input graph pair, are learnable weights, represents the cost per hop. The use of square operations can reduce the sensitivity of the function to subtle differences while increasing the penalty for large differences. φ(·) represents the cost of After encoding by the encoding module, multi-hop embedding is obtained, and then the projector MLP is used to map the representation of each layer to the embedding function of the embedding space.

[0075] Step 2: Model training: Use multiple groups of positive examples and negative examples to form a training data set and train the MSMC model, where the positive examples are sample graph pairs for which multi-hop consistency is established, that is, the sample graph pairs satisfy the subgraph inclusion relationship, and the negative examples are sample graph pairs for which multi-hop consistency is not established, that is, the sample graph pairs do not satisfy the subgraph inclusion relationship. After the training is completed, a subgraph inclusion prediction model is obtained.

[0076] Specifically, in the present embodiment, the maximum boundary loss function can be used as the objective function for training during the training process, and its calculation expression is:

[0077]

[0078] in, is a learnable boundary, E + is a set of positive examples, E - is a set of negative examples.

[0079] During the training process, this embodiment uses the maximization of the boundary loss function as the objective function for training, and takes minimizing the multi-hop consistency overhead between matching node pairs as the training goal, that is, for matching pairs, the multi-hop consistency overhead calculated by the MSMC model should be minimized, otherwise, the total overhead should be at least higher than the preset threshold γ.

[0080] Step 3: Subgraph inclusion prediction: Get the input graph pair to be predicted, input it into the subgraph inclusion prediction model, and predict the subgraph inclusion relationship within the input graph or between input graph pairs.

[0081] The trained subgraph inclusion prediction model can predict the subgraph inclusion relationship between the input graph pair (data graph D and query graph Q) in real time. and each For any k, we have Then judge isomorphic to The subgraph of the input graph, i.e., the data graph D

[0082] The multi-hop consistency between the query graph Q holds, otherwise it does not hold. represent the nodes in the input graph pair (D, Q), represents the set of nodes in the data graph D in the input graph pair, Represents the set of nodes in the query graph Q in the input graph pair.

[0083] When this embodiment predicts the subgraph inclusion relationship between input graphs, if it is determined that each query node in the query graph Q matches at least one data node in the data graph D, then the data graph D is considered to contain the query graph Q, otherwise it is determined that the data graph D does not contain the query graph.

[0084] This embodiment aims to predict the subgraph inclusion relationship problem, adopts the multi-hop subgraph matching consistency (MSMC) method, and uses the neighborhood-based subgraph inclusion relationship measurement method to achieve subgraph inclusion prediction, so that the node-induced subgraph pair inclusion within and between graphs is consistent at each hop, so that binary prediction is performed to determine whether the data graph contains the query graph without testing the subgraph isomorphism condition. Compared with the traditional neural network-based prediction method, the MSMC model of this embodiment converges faster and achieves the highest performance level. It is also applicable to problems such as network alignment or pruning, has good portability and applicability, and can be applied to multiple fields such as subgraph search.

[0085] In order to verify the effectiveness of the present invention, an experimental test was carried out on the above method of the present invention.

[0086] Experimental settings: All models are trained on a server equipped with an Intel Core i9-10900K CPU, a GeForce RTX 2080Ti GPU, or a server with three GeForce RTX 3090 GPUs and an Intel Xeon Silver 4210 CPU.

[0087] Datasets: Experiments were conducted on nine datasets from different fields: 1) COX2, a labeled dataset of chemical compounds; 2) ENZYMES; 3) PROTEINS; 4) D&D, a protein labeling dataset in bioinformatics; 5) MSRC21, a real-world image dataset, with labels representing the relationship between superpixels; 6) IMDB-BINARY, an unlabeled movie collaboration dataset collected from social networks; 7) AIDS, a structural dataset of anti-HIV compounds; 8) PTC-FM; 8) PTC-FR; 9) PTC-MM, three datasets of compound structures with different carcinogenic effects on mice. In this example, each dataset is divided into a training set and a test set with a ratio of 8:2, and positive and negative examples are sampled from the divided datasets.

[0088] This example first evaluates the performance of the MSMC of the present invention in predicting neighborhood-based subgraph inclusion and compares it with other benchmark models. Then, the performance of MSMC is evaluated when k is large enough or when k is insufficient.

[0089] Benchmark models: This embodiment selects five models from related fields, all of which generate similarity or distance scores for a given input graph pair. These models are: subgraph inclusion prediction model (1) NeuroMatch and graph similarity learning methods, which represent subgraph inclusion relations as partial order relations; (2) SimGNN, which emphasizes important nodes for specific similarity indicators; (3) GOTSim, which captures neighborhood context and regards the similarity calculation problem between graph pairs as a minimum conversion cost estimation problem; (4) ISONET, which performs alignment matching based on the neighborhood information of edges; (5) Prune4Sed uses an attention mechanism to use the information of the query graph to guide the pruning process of the data graph, thereby approximating the subgraph edit distance.

[0090] The experimental results are as follows Figure 5 , 6 As shown, Figure 5 (a) corresponds to the results in the COX2 dataset, (b) corresponds to the results in the D&D dataset, (c) corresponds to the results in the PROTEINS dataset, (d) corresponds to the results in the IMDB-BINARY dataset, and (e) corresponds to the results in the MSRC 21 dataset; Figure 6 (a) corresponds to the results of the ENZYMES dataset, (b) corresponds to the results of the AIDS dataset, (c) is the result on the PTC-MM dataset, (d) is the result on the PTC-FR dataset, and (e) is the result on the PTC-MM dataset. From the comparison results, it can be seen that the MSMC of the present invention surpasses all benchmark models and achieves state-of-the-art performance on all six datasets. SimGNN performs well on five datasets and is the second best model. Its good performance is attributed to the generated node pair comparison histogram features, which provide information about the node pair feature distribution and graph size. The graph matching network GOTSim experienced a performance decline on PROTEINS, probably because the difference in query data size provides GOTSim with more opportunities to find the best graph transformation. Although the graph pairs are similar, they do not match. When the number of nodes or edges reaches 100, the alignment calculation of the GOTSim and ISONET experiments cannot be completed because it is too expensive; therefore, this embodiment conducts additional experiments on datasets with smaller graphs, and the results can be seen in Figure 5 , 6. GOTSim maintains good performance when the query data size does not differ significantly, which verifies the speculation about GOTSim. For similar reasons, ISONET generally performs worse than expected. NeuroMatch's performance degrades on D&D and PROTEINS, which also verifies that neighborhood-based methods relax topological constraints and lead to prediction errors, which becomes particularly prominent when the neighborhood of the data node is larger than the query node. The unstable performance of Prune4SED may be due to its high reliance on the attention mechanism.

[0091] In order to evaluate the performance of MSMC when k is large enough or insufficient, this embodiment first samples 20,000 pairs of graphs and randomly selects a node v to ensure that its k-hop neighborhood can cover the entire query graph, and calculates the cost of the pairwise subgraph inclusion relationship between v and each data node; if its minimum cost exceeds the threshold, v cannot match any node in the data graph, so D does not contain Q; in the latter case, this embodiment samples 20,000 pairs of graphs without checking k-hop coverage. If D contains Q, each query node should be able to find at least one match in the data graph, otherwise, D does not contain Q. The results show that when the k-hop neighborhood of the selected node covers the entire query graph, the neighborhood-based inclusion detection can already provide accurate subgraph inclusion prediction. Even if k is not large enough, the MSMC of the present invention can still perform subgraph inclusion prediction with an accuracy of up to 96.81% in a short time.

[0092]

[0093] Table 1 shows the accuracy of the MSMC of the present invention and the average running time (seconds) for each pair of graphs on subgraph inclusion prediction. “Neighborhood” means that MSMC predicts inclusion prediction based on the neighborhood, and “Full-graph” represents the actual subgraph inclusion prediction.

[0094] Table 2: Consistency when performing multi-hop reasoning on graph data

[0095]

[0096] Table 3: Migration performance results

[0097]

[0098] This example first constructs the following variants of MSMC: (1) k-pred MSMC: only the embedding generated by the last layer of the model is used for prediction; (2) k-th MSMC: the computation of subgraph inclusion overhead is applied to the representation generated by the last layer of the model, and multi-hop embedding is used for prediction; (3) CE-MSMC: cross entropy loss is used for training, and multi-hop embedding is used for prediction. Ablation experiments are performed to verify which part of the model contributes most to the performance, whether multi-hop consistency holds, and whether multi-hop consistency can effectively reduce mispredictions.

[0099] The experimental results are as follows Figure 7 As shown, Figure 7 (a) corresponds to the results on the COX2 dataset, (b) corresponds to the results on the D&D dataset, (c) corresponds to the results on the PROTEINS dataset, (d) is the result on the IMDB-BINARY dataset, (e) is the result on the MSRC 21 dataset, and (f) corresponds to the result on the ENZYMES dataset. The results show that CE-MSMC is the second best performing model on the four datasets, including PROTEINS, but converges the slowest in most cases. Except for COX2, k-th MSMC performs the worst, presumably because only applying the subgraph inclusion cost to the last layer and using multi-hop embedding does not help the model autonomously learn multi-hop external consistency, but on COX2, the difference in neighborhood size between query nodes and data nodes is the least obvious. k-pred MSMC outperforms CE-MSMC on both datasets, but performs worst on IMDB-BINARY. Compared with k-th MSMC, the other three models verify that multi-hop internal consistency does help alleviate the prediction error problem caused by neighborhood size differences and accelerates convergence. k-pred MSMC outperforms CE-MSMC on both datasets, indicating that multi-hop external consistency can improve prediction performance.

[0100] In order to verify whether the multi-hop external consistency holds, this embodiment samples 4096 pairs of graphs. For each pair of graphs, the output of each layer is used for prediction. The results in Table 2 show that MSMC achieves a high level of multi-hop external consistency on three real-world datasets. Although the performance of MSRC 21 drops by about 7.3% at the sixth layer, the scale of multi-hop embedding continuously improves the prediction ability of MSMC.

[0101] In order to verify the transferability, this example uses seeds different from those used in training and samples 4096 pairs of images for each dataset. The results are shown in Table 3. Figure 7Compared with the models trained separately in , the MSMC trained with COX2 maintains high performance on the five datasets and only performs poorly on IMDB-BINARY, which may be because the density of IMDB-BINARY is much higher than that of other datasets (the node-to-edge ratio is 4.88), while the ratios of other datasets range from 1.05 to 2.56.

[0102] The application of the present invention in network alignment is further verified. Network alignment is a multi-classification task. For each query node, the MSMC of the present invention ranks the multi-hop consistency cost of the node pair and finds its top k matches. This embodiment extracts 10,000 pairs of matching graphs from each data set, samples different query and data node ratios, and performs the graph alignment task when the query graph and the data graph are of equal size, otherwise, performs the subgraph alignment task.

[0103] The calculation of subgraph inclusion overhead may not be suitable for graph alignment tasks, because for any query graph Q and data graph D′, D(Q,D′)=0 does not necessarily mean that Q and D′ are isomorphic. Therefore, this embodiment replaces the calculation of subgraph inclusion overhead with the calculation of graph matching overhead: Δ(v,u)=D(v,u)+D(u,v), and the multi-hop consistency overhead function is adjusted accordingly.

[0104] As shown in Table 4, the multi-hop consistency overhead of the modified graph matching is slightly better than the original multi-hop consistency overhead on each dataset, and the latter can provide high-quality graph alignment at Hit@3. Figure 8 As shown, Figure 8 (a) corresponds to the result on the COX2 dataset, (b) corresponds to the result on the ENZYMES dataset, and (c) corresponds to the result on the MSRC 21 dataset. It can be seen from the results that the Hit@k of the MSMC of the present invention decreases with the increase of the ratio of the query graph to the data graph size, but still remains at a satisfactory level. This is because nodes with larger neighborhoods tend to have more potential matches, so enumerating matches becomes difficult.

[0105] Table 4: Results of graph alignment

[0106]

[0107] In this embodiment, the pruning ability of MSMC is further verified by adopting a simple iterative refinement function. MSMC is used to embed the computational graph and save it as an index, thereby pruning nodes and edges that are unlikely to match in the enumeration matching. Fig. 9 shows the average proportion of nodes and edges reduced in the data graph, where Fig. 9(a) corresponds to the results on the COX2 dataset, (b) corresponds to the results on the ENZYMES dataset, and (c) corresponds to the results on the MSRC 21 dataset. It can be seen from the results that after each iteration, this embodiment recreates a subgraph D′ using the top 5 lowest-cost data nodes for each query node, and generates embeddings for the new D′ in the next iteration. The first iteration has the best refinement effect, and the process usually converges at the fifth iteration and can reduce the size of the data graph by at least 35% with an accuracy of up to about 90%.

[0108] In summary, the present invention introduces multi-hop consistency to measure the subgraph inclusion prediction overhead and the neighborhood-based subgraph inclusion model MSMC, and predicts the subgraph inclusion relationship by quantifying the multi-hop consistency overhead of graph pairs. It can effectively reduce the prediction errors caused by relaxing topological constraints, is easy to implement, and can be flexibly applied to applications such as performing network alignment and pruning unmatched nodes in the enumeration process.

[0109] The subgraph inclusion prediction device based on multi-hop subgraph matching consistency in this embodiment includes:

[0110] A model building module, used to build an MSMC model, the MSMC model includes an encoding module, a mapping module and a prediction module, the encoding module includes a multi-layer GNN, which is used to respectively encode the neighborhoods of different hops of each node in the input graph pair so as to meet the multi-hop internal consistency, the mapping module includes an MLP projector, which is used to project the representation of each layer in the encoding result to the embedding space, so as to convert each node in the input graph pair into a series of multi-hop embeddings, the prediction module is used to calculate the multi-hop consistency cost between the nodes of the input graph pair according to the projection result output by the mapping module to measure the relative importance of each hop, and predict the subgraph inclusion relationship between the input graph pairs according to the calculated multi-hop consistency cost;

[0111] A model training module is used to use multiple groups of positive examples and negative examples to form a training data set and train the MSMC model, wherein the positive examples are sample graph pairs for which multi-hop consistency is established, that is, the sample graph pairs satisfy the subgraph inclusion relationship, and the negative examples are sample graph pairs for which multi-hop consistency is not established, that is, the sample graph pairs do not satisfy the subgraph inclusion relationship. After the training is completed, a subgraph inclusion prediction model is obtained;

[0112] The subgraph inclusion prediction module is used to obtain the input graph pairs to be predicted, input them into the subgraph inclusion prediction model, and predict the subgraph inclusion relationship between the input graph pairs.

[0113] The subgraph inclusion prediction apparatus based on multi-hop subgraph matching consistency in this embodiment corresponds to the subgraph inclusion prediction method based on multi-hop subgraph matching consistency in the above-mentioned embodiment, and will not be described in detail here.

[0114] This embodiment further provides a computer device, including a processor and a memory, wherein the memory is used to store a computer program, and the processor is used to execute the computer program to perform the above method.

[0115] It is understandable that the above method of this embodiment can be executed by a single device, such as a computer or server, etc., and can also be applied to a distributed scenario and completed by multiple devices in cooperation with each other. In the case of a distributed scenario, one of the multiple devices can only execute one or more steps in the above method of this embodiment, and multiple devices interact to complete the above method. The processor can be implemented in the form of a general-purpose CPU, a microprocessor, an application-specific integrated circuit, or one or more integrated circuits, etc., for executing related programs to implement the above method of this embodiment. The memory can be implemented in the form of a read-only memory ROM, a random access memory RAM, a static storage device, and a dynamic storage device. The memory can store an operating system and other applications. When the above method of this embodiment is implemented by software or firmware, the relevant program code is stored in the memory and called and executed by the processor.

[0116] This embodiment further provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above method is implemented.

[0117] Those skilled in the art will appreciate that the above-mentioned embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes. The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, may be implemented by computer program instructions. These computer program instructions may be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the functions in the process. Figure 1 A process or multiple processes and / or boxes Figure 1These computer program instructions can also be stored in a computer-readable memory that can guide a computer or other programmable data processing device to work in a specific way, so that the instructions stored in the computer-readable memory produce a product including an instruction device, which implements the functions specified in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide for implementing the process in the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0118] The above is only a preferred embodiment of the present invention, and does not limit the present invention in any form. Although the present invention has been disclosed as a preferred embodiment, it is not intended to limit the present invention. Therefore, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the content of the technical solution of the present invention shall fall within the scope of protection of the technical solution of the present invention.

Claims

1. A subgraph inclusion prediction method based on multi-hop subgraph matching consistency, characterized in that the steps include: Step 1: Model construction: Construct an MSMC model, which includes an encoding module, a mapping module and a prediction module. The encoding module includes a multi-layer GNN, which is used to encode the neighborhoods of different hops of each node in the input graph pair respectively so as to meet the multi-hop internal consistency. The mapping module includes an MLP projector, which is used to project the representation of each layer in the encoding result to the embedding space to convert each node in the input graph pair into a series of multi-hop embeddings. The prediction module is used to calculate the multi-hop consistency cost between the nodes of the input graph pair according to the projection result output by the mapping module to measure the relative importance of each hop, and predict the subgraph inclusion relationship between the input graph pairs according to the calculated multi-hop consistency cost; Step 2: Model training: Use multiple groups of positive examples and negative examples to form a training data set and train the MSMC model, wherein the positive examples are sample graph pairs for which multi-hop consistency is established, that is, the sample graph pairs satisfy the subgraph inclusion relationship, and the negative examples are sample graph pairs for which multi-hop consistency is not established, that is, the sample graph pairs do not satisfy the subgraph inclusion relationship. After the training is completed, a subgraph inclusion prediction model is obtained; Step 3: Subgraph inclusion prediction: obtain the input graph pair to be predicted, input it into the subgraph inclusion prediction model, and predict the subgraph inclusion relationship between the input graph pairs.

2. The subgraph inclusion prediction method based on multi-hop subgraph matching consistency according to claim 1, characterized in that: Determining whether multi-hop consistency is established includes: for each node in the data graph And query each node in the graph If (v,u) matches, that is, for any k, Then determine each isomorphic to The subgraph of the input graph, that is, the multi-hop consistency between the data graph D and the query graph Q holds, otherwise the multi-hop consistency does not hold. represent the nodes in the input graph pair (D, Q), represents the set of nodes in the data graph D in the input graph pair, represents the set of nodes in the query graph Q in the input graph pair, φ is the embedding function used to convert the graph into a multidimensional vector, and is the k-hop subgraph induced by v and u, i.e., the k-hop neighborhood of the corresponding node.

3. The subgraph inclusion prediction method based on multi-hop subgraph matching consistency according to claim 1, characterized in that: The multi-hop internal consistency is: for the input connected graph G, if for each In each hop, v Consistently subgraph, then the connected graph G is determined to satisfy the multi-hop internal consistency.

4. The subgraph inclusion prediction method based on multi-hop subgraph matching consistency according to claim 1, characterized in that: The encoding module learns the attributes of each node in the input graph and aggregates the information from the neighborhood of each node through each layer of GNN, and adds the aggregated neighborhood information to the embedding output of the previous layer to update the k-hop embedding. The intermediate output is stored and reused in each layer as the embedding of the neighborhood with different hop numbers. The expression is: Among them, Encoder i (v) represents the encoding result of node v, Ψ is the update function, □ is the aggregation function, yes The embedded representation of is the information aggregated from the neighbors of node v.

5. The subgraph inclusion prediction method based on multi-hop subgraph matching consistency according to claim 1, characterized in that: The prediction module uses a multi-hop consistency cost function to calculate the multi-hop consistency cost between the nodes of the input graph. The calculation expression of the multi-hop consistency cost function is: in, is the multi-hop consistency overhead between node v and node u, represent the nodes in the input graph pair (D, Q), represents the set of nodes in the data graph D in the input graph pair, represents the set of nodes in the query graph Q in the input graph pair, are learnable weights, represents the cost of each hop, φ(·) represents the After encoding by the encoding module, multi-hop embedding is obtained, and then the projector MLP is used to map the representation of each layer to the embedding function of the embedding space.

6. The subgraph inclusion prediction method based on multi-hop subgraph matching consistency according to any one of claims 1 to 5, characterized in that: Step 2 During the training process, the multi-hop consistency overhead between the matched node pairs is minimized and the maximum boundary loss function is used as the objective function for training. The calculation expression of the objective function is: in, is a learnable boundary, E + is a set of positive examples, E - is a set of negative examples.

7. The subgraph inclusion prediction method based on multi-hop subgraph matching consistency according to any one of claims 1 to 5, characterized in that: The subgraph inclusion relationship between the predicted input graph pairs includes: if it is determined that each query node in the query graph Q matches at least one data node in the data graph D, then the data graph D is considered to contain the query graph Q; otherwise, it is determined that the data graph D does not contain the query graph Q.

8. A subgraph inclusion prediction device based on multi-hop subgraph matching consistency, characterized in that: include: A model building module, used to build an MSMC model, the MSMC model includes an encoding module, a mapping module and a prediction module, the encoding module includes a multi-layer GNN, which is used to respectively encode the neighborhoods of different hops of each node in the input graph pair so as to meet the multi-hop internal consistency, the mapping module includes an MLP projector, which is used to project the representation of each layer in the encoding result to the embedding space, so as to convert each node in the input graph pair into a series of multi-hop embeddings, the prediction module is used to calculate the multi-hop consistency cost between the nodes of the input graph pair according to the projection result output by the mapping module to measure the relative importance of each hop, and predict the subgraph inclusion relationship between the input graph pairs according to the calculated multi-hop consistency cost; A model training module is used to use multiple groups of positive examples and negative examples to form a training data set and train the MSMC model, wherein the positive examples are sample graph pairs for which multi-hop consistency is established, that is, the sample graph pairs satisfy the subgraph inclusion relationship, and the negative examples are sample graph pairs for which multi-hop consistency is not established, that is, the sample graph pairs do not satisfy the subgraph inclusion relationship. After the training is completed, a subgraph inclusion prediction model is obtained; The subgraph inclusion prediction module is used to obtain the input graph pairs to be predicted, input them into the subgraph inclusion prediction model, and predict the subgraph inclusion relationship between the input graph pairs.

9. A computer device comprising a processor and a memory, wherein the memory is used to store a computer program, wherein: The processor is configured to execute the computer program to perform the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the method according to any one of claims 1 to 7 is implemented.