Sub-graph matching cache query method and device, and vertex query method and device
By caching and reusing the intersection results in the sub-graph matching process, the problem of large update overhead in the prior art is solved and the matching efficiency is improved.
Patent Information
- Application Number
- CN202510268574.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-07
- Publication Date
- 2025-06-24
AI Technical Summary
The existing subgraph matching algorithm does not retain intermediate results, resulting in large update overhead and unable to effectively utilize the previously updated calculation results.
Reduce update overhead by cached intersection results and re-reuse them. The specific method includes obtaining the normal extended point dependency set of the current query vertex, calculating the intersection result of the vertex, and cacheing and reusing it to reduce the repeated intersection process.
By caching and reusing the intersection results, the update overhead and repeated intersection processes are reduced, and the efficiency of sub-graph matching is improved.
Smart Images

Figure CN120196656A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of data processing, and particularly relates to a cache query method, a vertex query method and a device for subgraph matching.
Background Art
[0002] With the development of mobile applications, a large amount of complex data is dynamically created at high speed. This data contains valuable information and can be easily modeled as a graph. It is very important to effectively analyze this data with a dynamic graph structure.
[0003] Dynamic subgraph matching is an important research topic in network analysis, aiming to efficiently identify subgraph patterns that meet specific structures and attributes from a dynamic network. Dynamic subgraph matching has wide applications in scenarios such as social network analysis, financial transaction monitoring, and the Internet of Things. Its effectiveness and real-time nature are of great significance for anomaly detection, pattern recognition, and trend prediction.
[0004] The continuous subgraph matching (CSM) algorithm for dynamic graphs is a current research hotspot. Some subgraph matching algorithms do not rely on any auxiliary data to accelerate the search. The IncMat algorithm updates the subgraph centered on the edges and performs subgraph matching on the retrieval part; the MultiView algorithm treats each query as an independent view and focuses on how to merge multiple views. The Graphflow algorithm significantly reduces the search space by applying worst-case optimal join. However, these subgraph matching algorithms do not retain intermediate results and cannot utilize the computational results of previous updates, so they are easily affected by the overhead of recomputation.
Summary of the Invention
[0005] To solve the technical problem that existing subgraph matching algorithms do not retain intermediate results, resulting in large update overhead, the present invention provides a cache query method, a vertex query method and a device for subgraph matching, which reduce the update overhead by caching the intersection results and reusing them.
[0006] The first aspect of the embodiment of the present invention provides a cache query method for subgraph matching, including:
[0007] Obtain the normal extended point dependency set of the current query vertex, where the current query vertex is a vertex in the matching order corresponding to the query graph, has a label Ls, and the normal extended point dependency set of the current query vertex includes vertices u1 to u n , and the candidate solutions for the vertices u1 to u n matching in the data graph are vertices v1 to v n , n≥1, and is a positive integer;
[0008] Obtain the set of v1 neighbors with the label Ls of the vertex v1 according to the data graph;
[0009] If n≥2 and the set of v1 neighbors is not empty, then determine whether there is an intersection result of the cached v corresponding to the vertex v2, where the 1~2,Ls intersection result is the intersection result of the sets of v1 to v2 neighbors with the label Ls from vertex v1 to v2; 1~2,Ls If there is, obtain the
[0010] intersection result; 1~2,Ls intersection result;
[0011] If not, obtain the 1~2,Ls intersection result according to the data graph;
[0012] Traverse the set of normal extension point dependencies of the current query vertex to obtain and save the 1~2,Ls intersection result to the 1~n,Ls intersection result;
[0013] Output the 1~n,Ls intersection result as the target candidate set corresponding to the current query vertex.
[0014] The second aspect of the embodiments of the present invention provides a vertex query method for subgraph matching, including:
[0015] If the set of normal extension point dependencies of the current query vertex is not empty and the set of frozen extension point dependencies of the current query vertex is empty, then perform a cache query according to the set of normal extension point dependencies of the current query vertex by using the cache query method for subgraph matching described in the first aspect above to obtain the target candidate set corresponding to the current query vertex.
[0016] The third aspect of the embodiments of the present invention provides a subgraph matching device, including
[0017] A preprocessing module, configured to preprocess the data graph in response to graph updates to obtain a query graph, a matching order, a set of normal extension point dependencies, and a set of frozen extension point dependencies;
[0018] A search and matching module, configured to obtain the target candidate set corresponding to the current query vertex according to the query graph, the matching order, the set of normal extension point dependencies, and the set of frozen extension point dependencies, so as to search for a subgraph of the data graph that matches the query graph;
[0019] A cache query module, configured to execute the cache query method for subgraph matching described in the first aspect above, and output the target candidate set corresponding to the current query vertex to the search and matching module.
[0020] A fourth aspect embodiment of the present invention provides a computer device, which includes at least one connected processor, a memory, and a transceiver. Among them, the memory is used to store program codes, and the processor is used to call the program codes in the memory to execute the steps of the vertex query method described in the second aspect above.
[0021] A fifth aspect of the embodiments of the present invention provides a computer storage medium, which includes instructions that, when running on a computer, cause the computer to execute the steps of the vertex query method for subgraph matching described in the second aspect above.
[0022] A sixth aspect of the embodiments of the present invention provides a computer program product, which includes instructions that, when executed by a processor, implement the steps of the vertex query method for subgraph matching described in the second aspect above.
[0023] Compared with the related art, in the embodiments provided by the present invention, by caching the intersection results of the vertex neighbors of each layer, saving the intermediate results, reducing the update overhead, reusing the intersection results, reducing a large number of repeated intersection processes, and accelerating the matching process.
Description of the Drawings
[0024] Figure 1 A schematic diagram of subgraph matching provided by an embodiment of the present invention;
[0025] Figure 2 A schematic diagram of the label distribution of point 2 in the data graph provided by an embodiment of the present invention;
[0026] Figure 3 A schematic diagram of the label distribution coverage of point 2 in the data graph G for the query graph q provided by an embodiment of the present invention;
[0027] Figure 4 A schematic diagram of an intersection result during the subgraph matching provided by an embodiment of the present invention;
[0028] Figure 5 A schematic diagram of the flow of the vertex query method for subgraph matching provided by an embodiment of the present invention;
[0029] Figure 6 A schematic diagram of the label classification of vertices 2 and 3 in the data graph G provided by an embodiment of the present invention;
[0030] Figure 7 A schematic diagram of the construction of the query point dependency set of the query graph Q provided by an embodiment of the present invention;
[0031] Figure 8 A schematic diagram of the flow of the cache query method for subgraph matching provided by an embodiment of the present invention;
[0032] Figure 9Schematic diagram of the cache structure of the forest structure provided by the embodiment of the present invention;
[0033] Figure 10A Schematic diagram of query graph Q1 provided by the embodiment of the present invention;
[0034] Figure 10B Schematic diagram of data graph G provided by the embodiment of the present invention;
[0035] Figure 11 Schematic diagram of the virtual structure of the subgraph matching device provided by the embodiment of the present invention;
[0036] Figure 12 Schematic diagram of the hardware structure of the server provided by the embodiment of the present invention.
Detailed implementation manners
[0037] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.
[0038] In the prior art, the IncMat algorithm, the MultiView algorithm, and the Graphflow algorithm do not retain intermediate results and cannot utilize the previously updated calculation results, resulting in huge update overhead.
[0039] Other algorithms construct data structures or indexes related to queries during updates to accelerate matching. SJ-Tree and TriC maintain tree-like structures and cache partially matched subqueries. However, such methods consume a large amount of time and space. TurboFlux and SymBi maintain intermediate results of queries based on the data center to reduce the search space, while Rapidflow accelerates batch subgraph matching by creating a global index. However, the data structures of these algorithms are closely related to the queries, and the space overhead is at least proportional to the product of the sizes of the queries and the graphs. Therefore, when evaluating a large number of queries, the space usage will increase sharply and may reach an unacceptable level.
[0040] Existing CSM methods either have huge update overhead or a large number of repeated intersection processes. Therefore, the present invention provides a cache query method, a vertex query method, and a device for subgraph matching that caches intersections and reuses them.
[0041] First, some terms related to the present invention will be explained:
[0042] Graph: A graph is a triple G = (V, E, L). Where V represents the set of vertices, E is the set of edges. L is a labeling function that assigns labels to each vertex and edge. For ease of description, in the present invention, the labels of vertex u and edge (u, v) are represented by L(u) and L(u, v) respectively, and the sets of different vertex (edge) labels are represented by L(V) and L(E). We use N G (V) to represent the neighbor set of v in G, and degr(v) to represent the corresponding average degree. For ease of expression, we can also use V G , E G , L G to represent the vertex set, edge set, and labeling function of G respectively.
[0043] Update Streams and Dynamic Graphs: A dynamic graph is defined as a tuple (G0, Δ). Where G o is an initial graph, and Δ is an update stream defined on G0. Δ is essentially a sequence of update operations Δ = {O1, O2,...}. Each O t is a triple <op, v1, v2>. Where op is the update type indicating the insertion / deletion of edge <v1, v2>. We use op = + and op = - to represent insertion and deletion respectively. We use G t to represent the data graph formed by applying the operations {O1, O2,..., O t} on G0. For ease of expression, we also use (op, e) to represent <op, v1, v2>, where e is the edge <v1, v2>.
[0044] Subgraph Matching: Consider a data graph G = (V, E, L) and a query graph Q = (V Q , E Q , L Q ). A subgraph g of G is called a match (isomorphism) of Q. If and only if there exists a bijective mapping Q from V g to V such that the following conditions are satisfied:
[0045]
[0046] A data vertex v matches a query vertex u if and only if L(v) = L Q (u). And a data edge e = (v1, v2) matches a query edge ∈ = (u1, u2) if and only if v1 (v2) matches u1 (u2) and L Q (u1, u2) = L(v1, v2). When the context is clear, we can also call F a match of Q because the bijective mapping is equivalent to the target subgraph. As Figure 1 shown, Figure 1A schematic diagram of subgraph matching provided by an embodiment of the present invention.
[0047] Definition of the continuous subgraph matching problem: Consider a dynamic graph and a set of query graphs Q = {Q1, Q2, …, Q k}. We process each update operation in sequence Δ on G0. For each update operation o t-1 applied on G t , the algorithm returns the corresponding new (inserted) / expired (deleted) matches for each query Q t in Q caused by o j .
[0048] Label distribution: For a vertex v, the label statistics LD(v) of all the edges e and neighbors u connected to it is the label distribution of vertex v. Figure 2 LD(2) shows the label distribution of point 2 in the data graph provided by an embodiment of the present invention.
[0049] Label distribution coverage: For vertices u, v (v, u). If each bit of the label distribution LD(u) (LD(v)) of u (v) is greater than or equal to the label distribution LD(v) (LD(u)) of v (u). Then it is said that u (v) achieves label distribution coverage for v (u). Figure 3 LD(2) shows the label distribution coverage of point 2 in the data graph G for the query graph Q provided by an embodiment of the present invention.
[0050] Definition of matching order and left / right neighbors: Given a query graph Q, the matching order is denoted as Φ, which is a sequence of all query vertices: u1, u2, …, u k (where k = |V(Q)|). For the vertex u i in Φ, u j , if u i is a neighbor of u j in query Q and i < j (or i > j), then u i is called the left (or right) neighbor of u j . We use LN(u i )(RN(u i )) to represent the set of left (right) neighbors of u i . Use q i to represent the sub-query guided by the first i vertices in Φ.
[0051] Definition of candidate set / intersection result: Consider a query Q, a data graph G, and the corresponding matching order. When matching to u i , the matching of its left neighbor has been completed. The candidate set of u i , denoted as is the intersection of the neighbors of the corresponding vertices of its left neighbor, which is also called the intersection result. Figure 4 This is a schematic diagram of an intersection result in the subgraph matching process provided by an embodiment of the present invention. The query graph Q is matched along the matching order {vertex 1, vertex 2, vertex 3} and has been matched to vertex 3 with label C. Its left neighbors are vertices 1 and 2. The vertices corresponding to vertices 1 and 2 in the data graph G are vertices 2 and 3. By taking the intersection of the neighbors with label C of vertices 2 and 3 in the data graph G, the Figure 4 shown intersection result is obtained, that is, the candidate set {vertex 4, vertex 5} for the match of vertex 3 in the query graph Q in the data graph G.
[0052] The present invention provides a vertex query method for subgraph matching. The vertex query method for subgraph matching will be described below from the perspective of a subgraph matching device. The subgraph matching device can be a server or a service unit in the server, and no specific limitation is made.
[0053] Please refer to Figure 5 , which is a schematic flowchart of the vertex query method for subgraph matching provided by an embodiment of the present invention. The vertex query method includes:
[0054] 101. If the normal expansion point dependency set of the current query vertex is not empty and the frozen expansion point dependency set of the current query vertex is empty, perform a cache query according to the normal expansion point dependency set of the current query vertex to obtain the target candidate set corresponding to the current query vertex.
[0055] In this embodiment, the current query vertex is the vertex currently being matched in the matching order. The vertices before the current query vertex in the matching order have been matched, and the vertices or candidate solutions matched by them in the data graph G have been added to the target mapping table (Mapping) corresponding to the query graph Q. It should be noted that when the length of the Mapping reaches the length of the point set of the query graph Q, it means that all vertices in the query graph Q have been matched. Directly expand all isolated points in the query graph Q into the Mapping and directly return the result, that is, the subgraph matching is completed.
[0056] Before the matching, the data graph G is preprocessed as follows.
[0057] Label distribution construction:
[0058] Determine the label distribution corresponding to each vertex in the data graph and the label distribution corresponding to each vertex of each query graph in the query graph set. Specifically, for all vertices v of the data graph G and all query graphs Q corresponding to the data graph G, a corresponding label distribution LD(v) is established.
[0059] Construction based on the data structure of the neighbors of the vertex:
[0060] Classify the neighbors of vertices according to the data graph by labels, specifically including:
[0061] Obtain the number of vertices and the number of labels according to the data graph;
[0062] Traverse the vertices and the neighbors of the vertices;
[0063] Obtain the labels of the neighbors of the vertex to construct a two-dimensional label array of the vertex, where the subscript of the first dimension is the label of the neighbor, and the second dimension stores the neighbors with the same label. Figure 6 This is a schematic diagram of the neighbor label classification of vertices 2 and 3 in the data graph G provided by the embodiment of the present invention.
[0064] Matching order construction:
[0065] Generate corresponding matching orders for all edges of each query graph Q. Those skilled in the art can select existing matching order generation strategies according to specific needs. Use the two vertices of the starting edge as the first two vertices in the matching order. For the other vertices in the matching order, there are three selection sorting strategies:
[0066] 1. Select the vertex that is most connected to the current vertex in the query graph Q in the matching order and add it after the current vertex in the matching order.
[0067] 2. If there are vertices with the same number of connections to the current vertex in the query graph Q in the matching order, select the vertex with a larger degree and add it after the current vertex in the matching order. Here, the degree of a vertex in the query is the number of edges connected to this vertex.
[0068] 3. If the number of connections to the current vertex in the query graph Q is equal and their degrees are also equal, select the vertex with a smaller number and add it after the current vertex in the matching order. Here, the number is determined by the data. The vertices in the query graph Q have their own numbers, and the numbers are actually simple 1, 2, 3, 4, ….
[0069] Thus, according to the above strategies, corresponding matching orders can be established for all edges e in all query graphs Q corresponding to the data graph G.
[0070] Query point dependent list construction:
[0071] Determine the query point dependent set corresponding to each vertex in the matching order. Specifically, in the matching order, when the following two conditions are met, put vertex v into the query point dependent set of vertex u:
[0072] 1. The position of vertex u in the matching order is after vertex v;
[0073] 2. There is an edge connecting vertex u and vertex v in the query graph Q. Figure 7 This is a schematic diagram for constructing the query point dependency set of the query graph Q provided by an embodiment of the present invention.
[0074] Construction of the attributes of vertices in the matching order:
[0075] Assign attributes to the vertices in the matching order. There are three attributes in total, namely normal vertex, freeze vertex, and isolated vertex. Specifically, the corresponding attributes can be assigned to each vertex in the query matching order through the following methods:
[0076] 1. For vertex v at the t-th position in the matching order {t} , if vertex v {t} appears in the query point dependency set of vertex v {t+1} at the (t + 1)-th position, then vertex v {t} is a normal vertex; v {t}
[0077] 2. If vertex v {t} does not appear in the query point dependency set of vertex v {t+1} at the (t + 1)-th position, but appears in the query point dependency set of vertex v {j} at the j-th position, and j > t + 1, then vertex v {t} is a freeze expansion point;
[0078] 3. If vertex v {t} does not appear in the query point dependency sets of the vertices after it in the order, then vertex v {t} is an isolated vertex.
[0079] Construction of the normal expansion point dependency set (normalDependentList) and the freeze expansion point dependency set (searchExpandVertexList):
[0080] Determine the normal expansion point dependency set and the freeze expansion point dependency set of each vertex in the matching order transaction according to the attributes corresponding to each vertex in the query point dependency set. Specifically, the normal vertices in the query point dependency set corresponding to the vertex are put into the normal expansion point dependency set corresponding to the vertex, and the freeze vertices are put into the freeze expansion point dependency set corresponding to the vertex.
[0081] Construction of the query graph index (Find QGraph Map):
[0082] Establish the query graph indexes corresponding to all query edges in the query graph set. Specifically, for each triple corresponding to each edge in all query graphs Q as the key value; to query the number of the query graph Q and the combination of the numbers of the corresponding edges in the query graph Q <q index , edge index > as the value; traverse all the query edges in all the query graphs to establish a query graph index. When there is an updated edge in the data graph, all relevant query graphs and corresponding edges can be quickly found based on the query graph index for subsequent searches.
[0083] Index graph update (indexUpdate)
[0084] When the target edge is updated in the data graph, the index of the data graph needs to be updated simultaneously. indexUpdate is mainly about the addition and / or deletion of edges on the data graph G t above, where the addition and / or deletion of edges will cause the modification of the adjacent edges of the two vertices of the edge and the change of the label distribution. The update is performed as follows:
[0085] When op = + (that is, when adding an edge), for the data graph G t the label distribution of the affected vertex u(v) on it is incremented at the position where index = v(u), and the label L(e) position of the added edge e also needs to be incremented.
[0086] When op = - (that is, when deleting an edge), for the data graph G t the label distribution of the affected vertex u(v) on it is decremented at the position where index = v(u), and the label L(e) position of the added edge e also needs to be decremented.
[0087] Specifically, the increment operation: when an edge is inserted in the data graph, the label distributions of the neighbors of the two vertices of the inserted edge will change, so its label distribution needs to be incremented by 1 at the label positions of the corresponding new neighbors, and the label position of the corresponding new edge also needs to be incremented by 1; the decrement operation: when an edge is deleted in the data graph, the label distributions of the neighbors of the two vertices of the inserted edge will change, so its label distribution needs to be decremented by 1 at the label positions of the corresponding new neighbors, and the label position of the corresponding new edge also needs to be decremented by 1.
[0088] Please refer to Figure 8 , which is the flow schematic diagram of the cache query method for subgraph matching provided by the embodiment of the present invention. During the subgraph matching process, for the current query vertex, when the subgraph matching device determines that the normal expansion point dependency set of the current query vertex is not empty and the frozen expansion point dependency set of the current query vertex is empty, cache query is performed according to the normal expansion point dependency set of the current query vertex to obtain the target candidate set of the current query vertex. The cache query method specifically includes:
[0089] 1011. Obtain the normal extension point dependency set of the current query vertex. The current query vertex is a vertex in the matching order corresponding to the query graph and has a label Ls. The normal extension point dependency set of the current query vertex includes vertices u1 to u n , vertices u1 to u n . The candidate solutions matched in the data graph for vertices v1 to v n , n ≥ 1 and is a positive integer;
[0090] 1012. Obtain the v1 neighbor set with label Ls of vertex v1 according to the data graph;
[0091] 1013. If n ≥ 2 and the v1 neighbor set is not empty, then determine whether there is a cached v 1~2,Ls intersection result corresponding to vertex v2, where the v 1~2,Ls intersection result is the intersection result of the v1 to v2 neighbor sets with label Ls of vertices v1 to v2;
[0092] 1014. If there is, then obtain the v 1~2,Ls intersection result;
[0093] 1015. If there is no such result, then obtain the v 1~2,Ls intersection result according to the data graph;
[0094] 1016. Traverse the normal extension point dependency set of the current query vertex to obtain and save the v 1~2,Ls intersection result to the v 1~n,Ls intersection result;
[0095] 1017. Output the v 1~n,Ls intersection result as the target candidate set corresponding to the current query vertex.
[0096] It can be understood that since both the current query vertex and the normal point have connecting edges in the query graph Q, the vertices matched by the current query vertex in the data graph G must have connecting edges with the vertices matched by the normal point in the data graph G. Therefore, the v 1~n intersection result is the union of the candidate solutions of the vertices matched by the current query vertex in the data graph G.
[0097] It can also be understood that the order of vertices u1 to u n is obtained by sorting. The candidate solutions matched by vertices u1 to u n in the data graph may be one or more. The subgraph matching device outputs the target candidate set based on an actual input of different candidate solutions according to a normal extension point dependency set.
[0098] In this embodiment, obtaining the v 1~2 intersection result according to the data graph may specifically include:
[0099] Obtain the set of v2 neighbors with label Ls of vertex v2 according to the data graph;
[0100] Take the intersection of the v1 neighbor set and the v2 neighbor set to obtain v 1~2,Ls Intersection result.
[0101] In some embodiments, if n > 2, the v 1~t,Ls Intersection result can be obtained through the following steps, where n >= t > 2 and t is a positive integer:
[0102] Determine whether there is a vertex v t Corresponding to the v 1~t,Ls Intersection result, where n >= t > 2 and t is a positive integer;
[0103] If it exists, obtain the v 1~t,Ls Intersection result;
[0104] If it does not exist, obtain the vertex v t With label Ls of v t Neighbor set;
[0105] Take the intersection of the v t Neighbor set and the v 1~t-1,Ls Intersection result to obtain the v 1~t,Ls Intersection result.
[0106] In this embodiment, the cache query method further includes:
[0107] If n = 1 or the v1 neighbor set is empty, output the v1 neighbor set as the target candidate set corresponding to the current query vertex.
[0108] That is, if there is only one vertex element in the normal extended point dependency set of the current query vertex, there is no need to take the intersection, and the v1 neighbor set is the target candidate set of the current query vertex. When the v1 neighbor set is an empty set, the intersection with any set must be empty. Therefore, the empty v1 neighbor set can be directly output as the target candidate set of the current query vertex.
[0109] In some embodiments, the cache query method further includes:
[0110] If the v p Neighbor set is empty, output the v p Neighbor set as the target candidate set of the current query vertex, where n >= p > 2 and p is a positive integer.
[0111] In this embodiment, please refer to Figure 9 , which is a schematic diagram of the cache structure of the forest structure provided by the embodiment of the present invention. The v 1~2,Ls Intersection result to v 1~n,LsThe intersection results are cached and managed in a forest structure. The forest structure has vertices as root nodes, labels as children of vertices, and intersection results as children of labels, v 1~2,Ls The intersection results to v 1~n,Ls The intersection results are separately saved to v with vertex v1 as the root node and label Ls as the parent node 1~2,Ls TreeNode to v 1~n,Ls In the tree node
[0112] It should be noted that each update of the data graph will invalidate the intersection results saved in the previous update, so the cache needs to be cleared. To speed up the clearing process, the subgraph matching device can use the root node set to save the root nodes of the forest. By traversing this set v1Set, the root nodes, that is, the vertices in the data vertex layer, can be quickly found. Only the pointers of all TreeNode* saved in tree level1 need to be cleared, and at the same time, the head pointer in the TreeNodePool is set to 0 to invalidate all nodes. At this time, no node space needs to be released, improving the efficiency. When a certain treeNode is used, the data in the node is cleared. Specifically, it includes the following steps:
[0113] Traverse the root node set v1Set to obtain the vertex v1 in the root node layer;
[0114] Traverse the label layer saved by vertex v1;
[0115] Clear the TreeNode* array saved by the label;
[0116] Repeat the above steps until the root node set v1Set is traversed.
[0117] 102. If the normal extension point dependency set of the current query vertex is not empty and the frozen extension point dependency set of the current query vertex is not empty, then intersect the neighbor sets with label Ls of the candidates matched by the vertices in the normal extension point dependency set of the current query vertex in the data graph to obtain the first candidate set.
[0118] 103. Intersect the neighbor sets with label Ls of the candidates matched by the vertices in the frozen extension points of the current query vertex in the data graph to obtain the second candidate set.
[0119] 104. Intersect the first candidate set and the second candidate set to obtain the target candidate set corresponding to the current query vertex.
[0120] In this embodiment, when the subgraph matching device determines that the normal expansion point dependency set of the current query vertex is not empty and the frozen expansion point dependency set of the current query vertex is not empty, it can obtain some candidate solutions according to the normal expansion point dependency set and obtain some candidate solutions according to the frozen expansion point dependency set. Since the current query vertex has edges connected to both the normal points and the frozen points in the query graph Q, the vertex in the data graph G that the current query vertex matches must have edges connected to both the vertices in the data graph G that the normal points and the frozen points match. Therefore, the intersection of the first candidate set and the second candidate set is the set of candidate solutions for the vertex in the data graph G that the current query vertex matches.
[0121] Taking the intersection of the neighbor sets with label Ls of the candidate solutions in the data graph of the vertices in the normal expansion point dependency set of the current query vertex to obtain the first candidate set may specifically include:
[0122] Obtaining the neighbor sets with label Ls of vertices v1 to v n from the data graph;
[0123] If n = 1, the neighbor set of v1 is the first candidate set;
[0124] If n ≥ 2, taking the intersection of the neighbor sets with label Ls of vertices v1 to v n to obtain the first candidate set.
[0125] It should be noted that in the matching order, for the previous query vertex relative to the current query vertex, if the previous query vertex is a normal point and there are multiple candidate solutions in its corresponding target candidate set, then the multiple candidate solutions corresponding to the previous query vertex need to be expanded one by one to obtain the first target candidate set corresponding to the current query vertex; for the next query vertex relative to the current query vertex, if the current query vertex is a normal point and there are multiple candidate solutions in its corresponding target candidate set, then the multiple candidate solutions corresponding to the current query vertex need to be expanded one by one to query the next query vertex.
[0126] In this embodiment, taking the intersection of the neighbor sets with label Ls of the candidate solutions in the data graph of the vertices in the frozen expansion points of the current query vertex to obtain the second candidate set may specifically include:
[0127] Obtaining the neighbor sets with label Ls of the candidate solutions in the data graph of vertices f1 to fm in the data graph, the frozen expansion point set includes vertices f1 to fm, m >= 1 and is a positive integer;
[0128] If m > 1, taking the intersection of the neighbor sets with label Ls of candidate solutions f1-cand to fm-cand to obtain the second candidate set, candidate solution f1-cand is any candidate solution of vertex f1, and candidate solution fm-cand is any candidate solution of vertex fm;
[0129] Traverse the candidate solutions corresponding to vertices f1 to fm to obtain all second candidate sets;
[0130] If m = 1, the neighbor set with label Ls of the candidate solution corresponding to vertex f1 is the second candidate set.
[0131] It can be understood that vertices f1 to fm are vertices in the matching order, whose positions are before the current query vertex and are connected to the current query vertex in the query graph. Therefore, when matching along the matching order, vertices f1 to fm have been matched, and the target candidate sets corresponding to vertices f1 to fm are obtained respectively, that is, the set of candidate solutions of vertices f1 to fm. However, since vertices f1 to fm are frozen points, they are not actually expanded, but only virtually expanded, and the target candidate sets corresponding to vertices f1 to fm are added to the mapping. At this time, when the current query vertex needs to be queried, the frozen points F1 to Fm need to be expanded.
[0132] The vertices in the data graph G that the current query vertex matches must be connected to all the vertices in the data graph G that vertices f1 to fm match and have the label Ls. When traversing the candidate solutions corresponding to the frozen points, if each of vertices f1 to fm has only one candidate solution, the number of second candidate sets is also one. The intersection of the m neighbor sets with label Ls of candidate solutions f1-cand to fm-cand can be used to obtain the second candidate set.
[0133] If there is a vertex among vertices f1 to fm that has at least two candidate solutions, then the number of the second candidate sets is at least two. Each time, a second candidate set is obtained by taking the intersection of the neighbor sets with label Ls of one candidate solution for each frozen point. For example, if the number of frozen points is two, namely vertices f1 and f2, the candidate solutions corresponding to vertex f1 are f1-C1 and f1-C2, and the candidate solutions corresponding to vertex f2 are f2-C1 and f2-C2. Then, the intersection of the neighbor sets with label Ls of f1-C1 and f2-C1 is taken to obtain a second candidate set A1, the intersection of the neighbor sets with label Ls of f1-C1 and f2-C2 is taken to obtain a second candidate set A2, the intersection of the neighbor sets with label Ls of f1-C2 and f2-C1 is taken to obtain a second candidate set A3, and the intersection of the neighbor sets with label Ls of f1-C2 and f2-C2 is taken to obtain a second candidate set A4. That is, by traversing the candidate solutions of vertices f1 and f2, a total of four second candidate sets are obtained. Further, if the obtained second candidate sets A1 and A2 are empty sets, it means that f1-C1 surely cannot match the frozen point f1. Therefore, by expanding the intersection step, vertices that match the frozen point can be further queried. Similarly, by taking the intersection of the first candidate set and the second candidate set, vertices that match the frozen point can also be further queried.
[0134] It can also be understood that there may be a connection between two vertices fx and fy among vertices f1 to fm, and there is a corresponding relationship between the candidate solutions of one vertex and those of the other vertex. For example, there is an edge connection, and x < y. Then, when vertex fx takes a certain candidate solution, the candidate solutions of vertex fy may change accordingly. For example, if the number of frozen points is two, namely vertices f1 and f2, when the candidate solution corresponding to vertex f1 is f1-C1, it is confirmed that the candidate solution corresponding to vertex f2 is f2-C1, and when the candidate solution corresponding to vertex f1 is f1-C2, it is confirmed that the candidate solution corresponding to vertex f2 is f2-C2. Then, the intersection of the neighbor sets with label Ls of f1-C1 and f2-C1 is taken to obtain a second candidate set, and the intersection of the neighbor sets with label Ls of f1-C2 and f2-C2 is taken to obtain a second candidate set. That is, by traversing the candidate solutions of vertices f1 and f2, a total of two second candidate sets are obtained.
[0135] It should be noted that there is no order of precedence for obtaining the first candidate set and obtaining the second candidate set. They can be executed successively, can be executed simultaneously, and no specific limitation is made here.
[0136] 105. If the normal expansion point dependency set of the current query vertex is empty and the frozen expansion point dependency set of the current query vertex is not empty, then the intersection of the neighbor sets with label Ls of the candidate solutions that match the vertices in the frozen expansion points of the current query vertex is taken to obtain the second candidate set as the target candidate set corresponding to the current query vertex.
[0137] In this embodiment, when the sub-graph matching device determines that the normal extension point dependency set of the current query vertex is empty and the frozen extension point dependency set of the current query vertex is not empty, it can obtain the target candidate set according to the frozen extension point dependency set. Since the normal extension point dependency set is empty, it means that there is no normal point in the query graph Q that has an edge connection with the current query vertex. It is only necessary to query and obtain the vertices in the data graph G that have an edge connection with the candidate solution of the frozen point and have the label Ls. The specific method has been described above and will not be elaborated here.
[0138] In some embodiments, the vertex query method further includes:
[0139] Performing label filtering on the first two vertices in the matching order corresponding to the query graph.
[0140] It should be noted that when the target edge is updated in the data graph, the target edge corresponding triple According to the triple Find the corresponding value in the FindQGraph Map, that is, find the query graph containing the triple and the query edge on the query graph that matches the triple In other words, the target edge can be regarded as a candidate solution of the query edge Taking the query edge as the starting edge, and taking the two vertices u , u i , u j of the query edge as the first two vertices in the matching order. When searching in the Find QGraph Map, a batch of query graphs will be found. The triple may also match different query edges of each query graph, and each query edge will generate a different matching order.
[0141] Therefore, by performing label filtering on the first two vertices in the matching order, that is, performing label filtering on the two vertices of the query edge and the two vertices of the target edge, query edges that do not match the target edge can be excluded, thereby filtering out unnecessary matching orders and accelerating the matching speed.
[0142] The principle of the test is that if L(u i ) = L(v i ) and L(u j ) = L(v j ) (L(u i ) = L(v j ) and L(u j ) = L(v i )) are satisfied, then u i and v i , u j and vj (u i and v j , u j and v i ) perform a label coverage test. The above has already elaborated in detail on the labels and label distribution coverage during the noun description, and specific details will not be repeated here.
[0143] In some embodiments, the vertex query method further includes:
[0144] If the normal expansion point dependency set of the current query vertex is not empty and the frozen expansion point dependency set of the current query vertex is empty, then perform label filtering on the vertices in the target candidate set;
[0145] If the normal expansion point dependency set of the current query vertex is not empty and the frozen expansion point dependency set of the current query vertex is not empty, then perform label filtering on the vertices in the first candidate set;
[0146] If the normal expansion point dependency set of the current query vertex is empty, the frozen expansion point dependency set of the current query vertex is not empty, and the current query point is a normal point, then perform label filtering on the vertices in the target candidate set.
[0147] It can be understood that performing label filtering on the vertices in the target candidate set and the first candidate set can also exclude the vertices that do not match the current query vertex, thereby accelerating the matching speed.
[0148] When the normal expansion point dependency set of the current query vertex is not empty and the frozen expansion point dependency set of the current query vertex is empty, the subgraph matching device determines that the cache query method can be executed, obtains the target candidate set from the cache, and needs to perform label filtering on the vertices in the target candidate set.
[0149] When the normal expansion point dependency set of the current query vertex is not empty and the frozen expansion point dependency set of the current query vertex is not empty, the subgraph matching device obtains the first candidate set according to the normal expansion point dependency set, obtains the second candidate set according to the frozen expansion point dependency set, performs label filtering on the first candidate set, and then takes the intersection of the first candidate set and the second candidate set. Performing label filtering on the first candidate set before the intersection can effectively reduce the first candidate set, with great benefits. And taking the intersection of the label-filtered first candidate set and the second candidate set is equivalent to performing label filtering on the second candidate set, so there is no need to perform label filtering on the target candidate set after the intersection.
[0150] If the normal expansion point dependency set of the current query vertex is empty and the frozen expansion point dependency set of the current query vertex is not empty, the subgraph matching device will obtain the second candidate set as the target candidate set according to the frozen expansion point dependency set. At this time, if the current query vertex is a normal point, it means that the current query vertex exists in the normal point expansion dependency set of the subsequent query vertex, and it is necessary to expand the target candidate set corresponding to the current query vertex to query the subsequent query vertex. Therefore, label filtering is performed on the target candidate set corresponding to the current query vertex to speed up the matching process. If the current query vertex is a frozen point, it means that the current query vertex does not exist in the query point dependency set of the subsequent query vertex, that is, to query the subsequent query vertex, it does not need to rely on the current query vertex, and the current query vertex is not actually expanded. Therefore, there is no need to perform label filtering on the current query vertex.
[0151] For the sake of easy understanding, the following will be described in combination with a specific application scenario Figure 10A and Figure 10B as follows:
[0152] When the target edge ΔG is inserted into the data graph, a new round of subgraph matching search is triggered. At this time, the label distributions of the two vertices v2 and v3 in the data graph G are updated first. The specific update method has been described in detail above and will not be elaborated here.
[0153] After that, look for a query graph containing the target triple <A, edgeLabel, B> in the Find QGraph Map corresponding to the data graph. Suppose the query graph Q1 and the edge e=(u0, u1) shown in Figure 10A are found.
[0154] After that, judge whether the label distribution of vertex v2 can cover the label distribution of vertex u0 in the query graph Q1 and whether the label distribution of vertex v3 can cover the label distribution of vertex u1. At this time, both can cover (u0 matches v2, u1 matches v3), so the search continues on the query graph Q1.
[0155] At this time, since the basis for continuing the search in the query graph Q1 is the matching of the <u0, u1> edge, starting from the <u0, u1> edge, search for the next query point in the matching Order corresponding to the query graph Q1. At this time, the matchingOrder is {u0, u1, u2, u3, u4}, and the Mapping at this time is {v2, v3}. According to the matching Order, it can be determined that the current query vertex is vertex u2. The query point dependency set of vertex u2 is {u1}, where the normal extension point dependency set is {u1}, and the frozen extension point dependency set is an empty set. The subgraph matching device confirms that the cache query method can be executed, and searches for vertices with label C among the neighbors of vertex v3 in the data graph (because the label of vertex u2 is C), and then obtains the neighbor set {v6, v7, v8} of v3. Since the normal extension point dependency set {u1} has only one element, the neighbor set {v6, v7, v8} of v3 is the target candidate set {v6, v7, v8} corresponding to vertex u2. Perform label filtering on it and find that all three points pass, so all candidate solutions are retained in the target candidate set {v6, v7, v8}.
[0156] The next query vertex in the matching order is vertex u3, with label D. The normal extension point dependency set of vertex u3 is {u2}, and the frozen extension point dependency set is an empty set. Vertex u2 is a normal point, and the target candidate set {v6, v7, v8} of vertex u2 needs to be actually expanded.
[0157] Add v6 to Mapping to get {v2, v3, v6}, execute the cache query method. The vertex corresponding to the normal extension point dependency set {u2} is {v6}, and obtain the neighbor set {v 11 , v 12} of v6 with label D in the data graph as the target candidate set of vertex u3, and perform label filtering on it to retain all candidate solutions.
[0158] The next query vertex in the matching order is vertex u4, with label D. The normal extension point dependency set of vertex u4 is {u2, u3}, and the frozen extension point dependency set is an empty set. Since vertex u3 is a normal point, the target candidate set {v 11 , v 12} of vertex u3 needs to be actually expanded.
[0159] Add v 11 to Mapping to get {v2, v3, v6, v 11}, execute the cache query method. The vertices corresponding to the normal extension point dependency set {u2, u3} are {v6, v 11}, and obtain v6 and v 11If the intersection of the neighbor sets with label D is an empty set, then v 11 is deleted from Mapping.
[0160] Add v 12 to Mapping as {v2, v3, v6, v 12}, execute the cache query method, and the vertices corresponding to the normal extension point dependency set {u2, u3} are {v6, v 12}. Obtain the intersection of the neighbor sets with label D of v6 and v 12 . At this time, a series of matches corresponding to vertex v6 in the matching order have been completed, and the possibility of matching between vertex v6 and vertex u2 has been excluded. Therefore, v6 and v 12 are deleted from Mapping.
[0161] Subsequently, vertices v7 and v8 are added to Mapping in sequence, and the same query method as vertex v6 is adopted, which will not be elaborated here. Finally, neither vertex v6 nor v7 is successfully matched with vertex u2, and vertex v8 is successfully matched with vertex u2. Eventually, the length of Mapping is equal to the number of points in the query graph Q1, and {v2, v3, v8, v 12 , v 13} and {v2, v3, v8, v 13 , v 12} are obtained, and the subgraph formed by the data graph G is isomorphic to the query graph Q1.
[0162] Compared with the related technology, in the embodiments provided by the present invention, by caching the intersection results of the vertex neighbors of each layer, saving the intermediate results, reducing the update overhead, reusing the intersection results, reducing a large number of repeated intersection processes, and accelerating the matching process. Classify the neighbors of the data graph according to the vertex labels, making it faster to filter the neighbors and speed up the intersection process. A cache intersection management based on the forest structure is proposed, which is convenient for querying and managing the intersection results and improves the efficiency.
[0163] The present invention is described above from the perspective of the vertex query method. Next, the present invention will be described from the perspective of the subgraph matching device.
[0164] Please refer to Figure 11 , which is a schematic virtual structure diagram of the subgraph matching device provided by the embodiment of the present invention. The subgraph matching device 300 includes:
[0165] A preprocessing module 301, configured to preprocess the data graph in response to a graph update to obtain a query graph, a matching order, a normal extension point dependency set, and a frozen extension point dependency set;
[0166] The search and matching module 302 is configured to obtain a target candidate set corresponding to the current query vertex according to the query graph, the matching order, the normal expansion point dependency set, and the frozen expansion point dependency set, so as to search for a subgraph of the data graph that matches the query graph;
[0167] The cache query module 303 is configured to execute the above cache query method and output a target candidate set corresponding to the current query vertex to the search and matching module;
[0168] The forest structure maintenance module 304 is configured to clear the cache in response to data graph updates.
[0169] In some embodiments, the preprocessing module 301 is specifically configured for label distribution construction, data structure construction based on the neighbors of vertices, matching order construction, query point dependency set construction, normal expansion point dependency set and frozen expansion point dependency set construction, query graph index construction, and index graph update. The specific methods have been described above and will not be elaborated here.
[0170] In some embodiments, the search and matching module 302 is specifically configured to:
[0171] If the normal expansion point dependency set of the current query vertex is not empty and the frozen expansion point dependency set of the current query vertex is empty, perform a cache query according to the normal expansion point dependency set of the current query vertex to obtain a target candidate set corresponding to the current query vertex;
[0172] If the normal expansion point dependency set of the current query vertex is not empty and the frozen expansion point dependency set of the current query vertex is not empty, intersect the neighbor sets with label Ls of the candidate solutions matched by the vertices in the normal expansion point dependency set of the current query vertex in the data graph to obtain a first candidate set.
[0173] Intersect the neighbor sets with label Ls of the candidate solutions matched by the vertices in the frozen expansion points of the current query vertex in the data graph to obtain a second candidate set.
[0174] Intersect the first candidate set with the second candidate set to obtain a target candidate set corresponding to the current query vertex.
[0175] If the normal expansion point dependency set of the current query vertex is empty and the frozen expansion point dependency set of the current query vertex is not empty, intersect the neighbor sets with label Ls of the candidate solutions matched by the vertices in the frozen expansion points of the current query vertex in the data graph to obtain a second candidate set as the target candidate set corresponding to the current query vertex.
[0176] In some embodiments, the search and matching module 302 is specifically further configured to:
[0177] Obtain vertices v1 to v according to the data graph nThe neighbor set with label Ls;
[0178] If n = 1, the neighbor set of v1 is the first candidate set;
[0179] If n ≥ 2, then intersect the neighbor sets of vertices v1 to v n with label Ls to obtain the first candidate set.
[0180] In some embodiments, the search and matching module 302 is further specifically configured to:
[0181] Obtain the neighbor set with label Ls of the candidate solutions matched by vertices f1 to fm in the data graph. The frozen expansion point set includes vertices f1 to fm, where m >= 1 and is a positive integer;
[0182] If m > 1, intersect the neighbor sets with label Ls of candidate solutions f1-cand to fm-cand to obtain the second candidate set. Candidate solution f1-cand is any candidate solution of vertex f1, and candidate solution fm-cand is any candidate solution of vertex fm;
[0183] Traverse the candidate solutions corresponding to vertices f1 to fm to obtain all second candidate sets;
[0184] If m = 1, the neighbor set with label Ls of the candidate solution corresponding to vertex f1 is the second candidate set.
[0185] In some embodiments, the search and matching module 302 is further specifically configured to:
[0186] If the normal expansion point dependency set of the current query vertex is not empty and the frozen expansion point dependency set of the current query vertex is empty, then perform label filtering on the vertices in the target candidate set;
[0187] If the normal expansion point dependency set of the current query vertex is not empty and the frozen expansion point dependency set of the current query vertex is not empty, then perform label filtering on the vertices in the first candidate set;
[0188] If the normal expansion point dependency set of the current query vertex is empty, the frozen expansion point dependency set of the current query vertex is not empty, and the current query point is a normal point, then perform label filtering on the vertices in the target candidate set.
[0189] In some embodiments, the forest structure maintenance module 304 is specifically configured to:
[0190] Traverse the root node set v1Set to obtain the vertex v1 in the root node layer;
[0191] Traverse the label layer saved by vertex v1;
[0192] Clear the array of TreeNode* saved by the tag;
[0193] Repeat the above steps until the root node set v1Set is traversed completely.
[0194] Figure 12 FIG. is a schematic structural diagram of the server of the present invention. The server 400 in this embodiment includes at least one processor 401, at least one network interface 404 or other user interfaces 403, a memory 405, and at least one communication bus 402. The server 400 optionally includes a user interface 403, including a display, a keyboard or a pointing device. The memory 405 may include a high-speed RAM memory, and may also include a non-volatile memory, such as at least one disk memory. The memory 405 stores execution instructions. When the server 400 runs, communication occurs between the processor 401 and the memory 405. The processor 401 invokes the instructions stored in the memory 405 to execute the vertex query method for subgraph matching described above. The operating system 406 includes various programs for implementing various basic services and processing tasks according to the hardware.
[0195] The server provided by the embodiment of the present invention can execute the technical solutions of the embodiments of the above vertex query method. The implementation principles and technical effects are similar, and will not be elaborated here.
[0196] The embodiment of the present invention also provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computer, it implements the method flow related to the subgraph matching device in any of the above method embodiments. Correspondingly, the computer can be the above subgraph matching device.
[0197] The embodiment of the present invention also provides a computer program or a computer program product including a computer program. When the computer program is executed on a certain computer, it will cause the computer to implement the method flow related to the subgraph matching device in any of the above method embodiments. Correspondingly, the computer can be the above subgraph matching device.
[0198] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product.
[0199] The above-described embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A cache query method for subgraph matching, characterized in that: include: Get the normal extension point dependency set of the current query vertex, the current query vertex is a vertex in the matching order corresponding to the query graph, has a label Ls, and the normal extension point dependency set of the current query vertex includes vertices u1 to u n , the vertices u1 to u n The candidate solutions matched in the data graph are vertices v1 to v n , n≥1, and is a positive integer; Obtain a set of v1 neighbors of the vertex v1 with the label Ls according to the data graph; If n ≥ 2 and the neighbor set of v1 is not empty, determine whether there is a cached v corresponding to vertex v2. 1~2,Ls The intersection result, where the v 1~2,Ls The intersection result is the intersection result of the neighbor sets of vertices v1 to v2 with the label Ls; If it exists, get the v 1~2,Ls Intersection results; If it does not exist, then obtain the v according to the data graph 1~2,Ls Intersection results; Traverse the normal extension point dependency set of the current query vertex to obtain and save the v 1~2,Ls Intersection result to v 1~n,Ls Intersection results; Output the v 1~n,Ls The intersection result is the target candidate set corresponding to the current query vertex.
2. The method according to claim 1, characterized in that The method further comprises: If n≥3, determine whether there is a vertex v t The corresponding v 1~t,Ls The intersection result, n≥t≥3, is a positive integer; If it exists, get the v 1~t,Ls Intersection results; If it does not exist, then obtain the vertex v according to the data graph t The v with the label Ls t Neighbor Set; The v t Neighbor set and v 1~t-1,Ls The intersection results are intersected to obtain the v 1~t,Ls Intersection results.
3. The method according to claim 1, characterized in that The method further comprises: If n=1 or the v1 neighbor set is empty, the v1 neighbor set is output as the target candidate set corresponding to the current query vertex.
4. The method according to claim 1, characterized in that The v 1~2,Ls The intersection result to the v 1~n,Ls The intersection result is cached in a forest structure, where the vertex is the root node, the label is the child node of the vertex, and the intersection result is the child node of the label. 1~2,Ls The intersection result to the v 1~n,Ls The intersection result is saved in a tree node with the vertex v1 as the root node and the label Ls as the parent node.
5. A vertex query method for subgraph matching, characterized in that: include: If the normal extension point dependency set of the current query vertex is not empty and the frozen extension point dependency set of the current query vertex is empty, then a cache query is performed based on the normal extension point dependency set of the current query vertex using the subgraph matching cache query method as described in any one of claims 1-4 to obtain the target candidate set corresponding to the current query vertex.
6. The method according to claim 5, characterized in that Also includes: If the normal extension point dependency set of the current query vertex is not empty and the frozen extension point dependency set of the current query vertex is not empty, performing intersection of the vertices in the normal extension point dependency set of the current query vertex and the neighbor sets with the label Ls of the candidate solutions matched in the data graph to obtain a first candidate set; Intersecting the vertices in the frozen extension points of the current query vertex with the neighbor sets of the candidate solutions matched in the data graph and having the label Ls, to obtain a second candidate set; Intersecting the first candidate set with the second candidate set to obtain a target candidate set corresponding to the current query vertex; If the normal extension point dependency set of the current query vertex is empty and the frozen extension point dependency set of the current query vertex is not empty, then the neighbor sets with the label Ls of the candidate solutions matching the vertices in the frozen extension points of the current query vertex in the data graph are intersected to obtain a second candidate set as the target candidate set corresponding to the current query vertex.
7. The method according to claim 6, characterized in that The step of performing intersection of the vertices in the normal extension point dependency set of the current query vertex and the neighbor sets with the label Ls of the candidate solutions matched in the data graph to obtain the first candidate set includes: According to the data graph, the vertices v1 to v are obtained. n The set of neighbors with label Ls; If n=1, the v1 neighbor set is the first target candidate set; If n ≥ 2, move the vertices v1 to v n The neighbor sets with label Ls are intersected to obtain the first target candidate set.
8. The method according to claim 6, characterized in that The step of performing intersection of the neighbor sets with the label Ls of the candidate solutions matching the vertices in the frozen extension points of the current query vertex in the data graph to obtain a second candidate set includes: Obtaining, according to the data graph, a neighbor set with a label Ls of candidate solutions matching vertices f1 to fm in the data graph, wherein the frozen extension point dependency set of the current query vertex includes the vertices f1 to fm, and m≥1 is a positive integer; If m≥2, perform intersection of neighbor sets with label Ls from candidate solutions f1-cand to fm-cand to obtain the second candidate set, wherein the candidate solution f1-cand is any candidate solution of the vertex f1, and the candidate solution fm-cand is any candidate solution of the vertex fm; Traverse the candidate solutions corresponding to the vertices f1 to fm to obtain all the second candidate sets; If m=1, the neighbor set with label Ls of the candidate solution corresponding to the vertex f1 is the second candidate set.
9. The method according to claim 6, characterized in that Also includes: If the normal extension point dependency set of the current query vertex is not empty and the frozen extension point dependency set of the current query vertex is empty, label filtering is performed on the vertices in the target candidate set; If the normal extension point dependency set of the current query vertex is not empty and the frozen extension point dependency set of the current query vertex is not empty, then perform label filtering on the vertices in the first candidate set; If the normal extension point dependency set of the current query vertex is empty, the frozen extension point dependency set of the current query vertex is not empty, and the current query vertex is a normal point, label filtering is performed on the vertices in the target candidate set.
10. The method according to claim 5, characterized in that The method further comprises: Obtaining the number of vertices and the number of labels according to the data graph; Traverse the vertex and the neighbors of the vertex; Get the labels of the neighbors of the vertex to construct a two-dimensional label array of the vertex. The subscript of the first dimension is the label of the neighbor, and the second dimension stores the neighbors with the same label.
11. A sub-graph matching device, characterized in that: include: A preprocessing module, used to preprocess the data graph in response to graph updates to obtain a query graph, a matching order, a normal extension point dependency set, and a frozen extension point dependency set; A search and matching module, configured to obtain a target candidate set corresponding to a current query vertex according to the query graph, the matching order, the normal extension point dependency set, and the frozen extension point dependency set, so as to search for a subgraph of the data graph that matches the query graph; A cache query module, used to execute the cache query method as described in any one of claims 1-4, and output the target candidate set corresponding to the current query vertex to the search and matching module.