Sub-graph matching coverage set method and device, storage medium and computer program product
By using forward intersection pruning, full coverage pruning and multi-extension methods in the sub-graph matching process, the problem of redundant calculation in sub-graph matching is solved, and a more efficient matching process and more accurate results are achieved.
Patent Information
- Application Number
- CN202510343626.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-22
- Publication Date
- 2025-07-18
AI Technical Summary
The prior art has serious problems with redundant computing in the sub-graph matching process, especially the lack of an effective pruning mechanism, which leads to inefficient computing efficiency and waste of resources.
The forward intersection pruning method, full coverage pruning method and multi-extended method are used to pre-calculate candidate sets of unmatched vertices and identify covered vertices, prune redundant branches in advance, and reduce unnecessary calculations.
It significantly reduces redundant calculations in sub-graph matching, improves search efficiency, ensures a more accurate matching process, and allows matching multiple isolated vertices to be simultaneously matched, reduces backtracking operations, and improves matching efficiency.
Smart Images

Figure CN120336586A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical fields of data analysis and data mining, and in particular, to a subgraph matching covering set method, apparatus, storage medium, and computer program product. Background Art
[0002] As a complex data analysis technology, subgraph matching has extensive applications in complex network analysis, pattern recognition, and graph databases. Traditional subgraph matching usually requires calculating all matching subgraphs, which is time-consuming and inefficient. Moreover, a single vertex may appear in multiple matches, resulting in a large amount of redundant calculation. For example, in the prior art: there is a lack of an effective pruning mechanism and it is impossible to exclude branches that cannot produce matches in advance; only the candidate set of a single vertex is calculated, and the candidate sets of multiple unmatched vertices cannot be considered simultaneously; due to the backtracking search method, a large number of redundant matching results will be generated.
[0003] Therefore, there is an urgent need to provide a subgraph matching covering set method, apparatus, storage medium, and computer program product to reduce the redundant calculation in subgraph matching.
[0004] Application Content
[0005] Embodiments of the present application provide a subgraph matching covering set method, apparatus, storage medium, and computer program product.
[0006] In a first aspect, embodiments of the present application provide a subgraph matching covering set method, the method including:
[0007] Receiving a given query graph to obtain a target subgraph in a data graph;
[0008] Obtaining a query point set including all query points and a matching order including all query points according to the query graph;
[0009] According to the matching order, starting from the first query point, calculating a candidate set of the current query point according to the data graph;
[0010] Obtaining a right neighbor of the current query point in the matching order, and a candidate set of each right neighbor;
[0011] Matching a data point for the current query point according to the candidate set of the current query point;
[0012] Updating the candidate set of the right neighbor of the current query point according to the data point matched by the current query point; and
[0013] Performing a forward intersection pruning method and / or a full coverage pruning method during the matching process,
[0014] wherein, the forward intersection pruning method includes:
[0015] Determine whether there is an empty set in the candidate set of the updated right neighbor;
[0016] If there is an empty set, determine that the matched data point is invalid, directly prune the current branch, and invalidate the update of the candidate sets of all right neighbors caused by the matching of this data point. All candidate sets of right neighbors return to the unupdated state. When the current query point is not the first query point, backtrack to the previous query point to rematch the data point.
[0017] The full-coverage pruning method includes:
[0018] Determine whether the current query point reaches or exceeds the pruning depth. Among them, the pruning depth corresponds to a specific query point. Before this specific query point becomes the current query point, the candidate sets of one or more query points have not been calculated. And when this specific query point becomes the current query point, after updating the candidate set of the right neighbor of this specific query point, at this time, the candidate sets of all query points have been calculated;
[0019] If the current query point reaches or exceeds the pruning depth, determine whether the matched data points of the current query point and the data points in the candidate sets of the query points after the current query point are all included in the coverage point set; and
[0020] If the matched data points of the current query point and the data points in the candidate sets of the query points after the current query point are all included in the coverage point set, prune from the data point matched by the current query point this time.
[0021] In one embodiment, updating the candidate set of the right neighbor of the current query point according to the data point matched by the current query point includes:
[0022] For the uninitialized right neighbor, calculate the candidate set of this right neighbor according to the data point matched by the current query point; and
[0023] For the right neighbor with a non-empty original candidate set, take the intersection of the original candidate set of this right neighbor and the newly calculated candidate set, and replace the original candidate set with the intersection result as the updated candidate set of this right neighbor.
[0024] In one embodiment, during the matching process, execute the forward intersection pruning method and the full-coverage pruning method. When it is determined that there is no empty set in the candidate set of the updated right neighbor, execute the operation of determining whether the current query point reaches or exceeds the pruning depth.
[0025] In one embodiment, the pruning depth d satisfies the following conditions:
[0026] Any and
[0027] exist
[0028] wherein, u d is the query point corresponding to the pruning depth d, and LN(u j ) refers to the set of left neighbors of the query point u j .
[0029] In one embodiment, in the matching order, the isolation point is placed at the end part, wherein the isolation point refers to a query point without a right neighbor, and the method further includes:
[0030] For each query point that is an isolation point, one by one, use the unmatched data points in the candidate set of the isolation point to replace the matched data points of the isolation point on the current branch, form a new solution and update the covered point set,
[0031] wherein, before replacing the matched data points of the isolation point, the unmatched data points do not exist on the current branch.
[0032] In one embodiment, the method further includes:
[0033] For each query point that is an isolation point, one by one, use the unmatched data points in the candidate set of the isolation point to replace the matched data points of the isolation point on the current branch, form a new solution and update the covered point set,
[0034] wherein, before replacing the matched data points of the isolation point, the unmatched data points do not exist on the current branch.
[0035] In one embodiment, the method further includes:
[0036] Judge whether the current query point reaches or exceeds the expansion depth, and the multi-expansion depth is the position where the first isolation point appears in the matching order; and
[0037] If the current query point reaches or exceeds the expansion depth, perform the operation of, for each query point that is an isolation point, one by one, using the unmatched data points in the candidate set of the isolation point to replace the matched data points of the isolation point on the current branch.
[0038] In a second aspect, the present application further provides a subgraph matching coverage set device, and the subgraph matching coverage set device includes:
[0039] A memory configured to store computer-executable instructions;
[0040] A processor configured to run computer-executable instructions stored in the memory to implement the subgraph matching coverage set method described above.
[0041] In a third aspect, the present application also provides a storage medium storing computer-executable instructions, which, when run by a processing unit, implement the subgraph matching coverage set method described above.
[0042] In a fourth aspect, the present application also provides a computer program product including computer-executable instructions, which, when executed by a processing unit, implement the subgraph matching coverage set method described above.
[0043] Compared with the prior art, the subgraph matching coverage set method, device, storage medium, and computer program product provided in the embodiments of the present application use matching coverage to replace full coverage, saving search time and improving search efficiency. Further, (1) by identifying covered vertices, redundant branches are effectively pruned, unnecessary calculations are reduced, and the algorithm efficiency is improved. (2) The candidate set of unmatched vertices is pre-computed to enhance the pruning effect and ensure a more accurate matching process. (3) Allowing multiple isolated vertices to be matched simultaneously reduces backtracking operations and significantly improves the matching efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0044] To more clearly illustrate the technical solutions in the embodiments of the present application, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings, where:
[0045] Figure 1 is a schematic diagram of a case of network attack, the result set of full coverage, and the result set of matching coverage exemplified in an embodiment of the present application.
[0046] Figure 2 is a flowchart of the subgraph matching coverage set method provided in an embodiment of the present application;
[0047] Figure 3 is a schematic query graph and data graph provided in an embodiment of the present application;
[0048] Figure 4 is provided in an embodiment of the present application according to Figure 2 a schematic diagram of the query point set and matching order obtained from the shown query graph;
[0049] Figure 5 is in an embodiment of the present application according to Figure 2 the shown data graph andFigure 3 Schematic diagram of a search tree obtained by performing a matching search in the shown matching order
[0050] Figure 6 It is a schematic diagram of another matching order and another data graph provided in an embodiment of the present application
[0051] Figure 7 is Figure 1 Schematic diagram of the refined process of the shown method
[0052] Figure 8 In an embodiment of the present application, according to Figure 6 Schematic diagram of a search tree obtained by performing a matching search according to the shown matching order and data graph
[0053] Figure 9 It is a schematic diagram of the hardware structure of a subgraph matching cover set device provided in an embodiment of the present application Specific embodiments
[0054] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application
[0055] The present application aims to propose the concept of match cover and the subgraph matching cover set method, device, storage medium and computer program product for realizing match cover
[0056] As a complex data analysis technology, subgraph matching has extensive applications in complex network analysis, pattern recognition, and graph databases. Traditional subgraph matching usually requires calculating all matching subgraphs, which is time-consuming and inefficient. Moreover, a single vertex may appear in multiple matches, resulting in a large amount of redundant calculations. However, in some cases, it is only necessary to return a matching cover set. For example, in terms of network attacks, please refer to Figure 1 shown Figure 1 Part (a) of Figure 1 shows a representative network attack example, where an attacker launches a DDos attack on the victim's machine Figure 1Part (c) of FIG. 1 shows the complete 8 matching sets (i.e., the full coverage result set {g1, g2, ..., g8}), which basically overlap with each other. In fact, only a subset of these matching sets needs to be returned, and any attack, robot program, and victim are included in this subset, such as Figure 1 As shown in part (d) of the figure, two of the eight matching sets {g1, g2} include all relevant vertices, and for each attacker or robot program, there is at least one match in the two matching sets {g1, g2} as evidence of its role in the DDos attack. Therefore, for a query graph, there is a situation where it is not necessary to find all its matching sets in a data graph, but only to find some matching sets.
[0057] In the example above, Figure 1 Part (c) shows the full coverage result set. Figure 1 Part (d) shows the result set of matching coverage. For the same query graph and data graph, the result set of full coverage includes all solutions, while the result set of matching coverage only includes partial solutions, that is, partial matching subgraphs, which together include the data points of all solutions. Since full coverage requires finding all solutions, while matching coverage only requires finding valid partial solutions, matching coverage can effectively reduce redundant calculations and reduce time costs.
[0058] The embodiment of the present application provides a subgraph matching cover set method, which includes a forward intersection pruning method, a full coverage pruning method and / or a multiple extension method. The forward intersection pruning method refers to pre-calculating the candidate set of unmatched vertices (i.e., query points), and judging in advance whether the current search branch will generate an empty candidate set (i.e., an empty set), so as to perform a safe pruning operation. The full coverage pruning method refers to using the coverage points in the obtained matching results to pre-prune the redundant matching branches that may be generated later, wherein, for a query result set R obtained for a data graph G and a query graph Q, any vertex belonging to the query result set R is called a coverage point. The multiple extension method refers to optimizing the isolated points in the matching order, i.e., query points without right neighbors, using the multiple extension method, and replacing the matched data points with unmatched data points for the same query point to obtain all the coverage points, thereby accelerating the matching process. The forward intersection pruning method, the full coverage pruning method and the multiple extension method are used to significantly reduce the redundant calculations in subgraph matching.
[0059] See also Figure 2 and Figure 3 As shown, Figure 2 A schematic diagram of a flow chart of a subgraph matching covering set method provided in an embodiment of the present application, Figure 3A query graph and a data graph for an example. The target subgraph matching the query graph is obtained from the data graph by using the subgraph matching covering set method. The method includes the following operations.
[0060] In operation 101, a given query graph is received to obtain a target subgraph in a data graph.
[0061] Please refer to Figure 3 as shown, where part (a) is an example query graph 10 and part (b) is an example data graph 20. The query graph 10 contains 6 query points u1 - u6, and the data graph 20 contains 100 data points v1 - v 100 in total.
[0062] In operation 103, all n query point sets and the matching order of the n query points are obtained according to the query graph.
[0063] Please refer to Figure 4 as shown, where Figure 4 part (a) shows all query point sets Φ = {u1, u2, u3, u4, u5, u6} obtained according to the query graph 10, Figure 4 and part (b) is a matching order 30 of all query points u1 - u6 obtained according to Figure 3 the query graph 10 shown.
[0064] In operation 105, according to the matching order, starting from the first query point u1, the candidate set of the current query point u i is calculated according to the data graph.
[0065] For example, for Figure 3-4 the query points u1 - u6 in the example, the candidate set of query point u1 in the data graph 20 can be obtained as Pu1 = {v1, v5}.
[0066] In operation 107, the right neighbor of the current query point ui in the matching order and the candidate set of each right neighbor u j (j > i) are obtained.
[0067] For example, for query point u1, its right neighbors are (u2, u3, u6), and the original candidate sets of right neighbors u2, u3, and u6 are not initialized.
[0068] In operation 109, according to the candidate set of the current query point u i a data point v is matched for the current query point u i .
[0069] For example, for query point u1, data point v1 is matched.
[0070] In operation 111, update the current query point u according to the data point v matched by the current query point ui. i The right neighbor u j Of the candidate set.
[0071] For example, according to the data point v1 matched by the query point u1, update the candidate sets of the right neighbors u2, u3, and u6 of u1 according to the neighbors of v1. The candidate set of the right neighbor u2 is Pu2 = {v4, v6}, the candidate set of the right neighbor u3 is Pu3 = {v3}, and the candidate set of the right neighbor u6 is Pu6 = {v8}.
[0072] Operation 111 further includes:
[0073] For a right neighbor whose original candidate set is not initialized, calculate the candidate set of this right neighbor according to the data point v matched by the current query point ui, and use it as the updated candidate set of this right neighbor;
[0074] For a right neighbor whose original candidate set is a non-empty set, take the intersection of the original candidate set of this right neighbor and the newly calculated candidate set, and replace the intersection result with the original candidate set as the updated candidate set of this right neighbor.
[0075] In operation 113, determine whether there is an empty set in the candidate set of the updated right neighbor. If there is an empty set, the process proceeds to 114; if there is no empty set, the process proceeds to operation 120.
[0076] In operation 114, determine that the data point matched by the current query point is invalid, prune the current branch, and the update of the candidate set of the right neighbor caused by the matching of this data point is invalid. The candidate set of the right neighbor returns to the unupdated state, and the process proceeds to operation 115.
[0077] In operation 115, determine whether the data points in the candidate set of the current query point have been traversed and matched. If so, the process proceeds to operation 116; if not, the process proceeds to operation 117, that is, match another data point in the candidate set of the current query point u i Of the candidate set for the current query point u i Match another data point.
[0078] In operation 116, determine whether the current query point u i Is the first query point in the matching order. If the current query point is the first query point, the process ends; if the current query point is not the first query point, the process proceeds to operation 118.
[0079] In operation 117, match another data point in the candidate set of the current query point u i Of the candidate set for the current query point u i Match another data point, and the process returns to operation 111. The other data point is a data point that has not been matched with the current query point before.
[0080] In operation 118, the data points are re-matched by backtracking to the previous query point, that is, the value of i is changed to (i - 1), backtracking to the previous query point, changing the previous query point to the current query point, selecting an unmatched data point from its candidate set for matching, and the process returns to operation 111.
[0081] For example, for the Figure 2-3 example in, according to the matching order, starting from u1 and ending at u6, select one of the data points from the candidate set of the current query point u i (where i is any natural number from 1 to 6) for matching:
[0082] Select the matching data point v1 in sequence from the candidate set Pu1 = {v1, v5} of u1, that is, u1--v1;
[0083] Select the matching data point v4 in sequence from the candidate set Pu2 = {v4, v6} of u2, that is, u2--v4;
[0084] Match the data point v3 from the candidate set Pu3 = {v3} of u3, that is, u3--v3;
[0085] Match the data point v5 from the candidate set Pu4 = {v1, v5} of u4 according to the label constraint, that is, u4--v5;
[0086] Match the data point v6 from the candidate set Pu5 = {v4, v6} of u5 according to the label constraint, that is, u5--v6; and
[0087] Select the matching data point v8 from the candidate set Pu6 = {v8} of u6.
[0088] In this way, starting from u1, data points are matched for each current query point in sequence. As the matching progresses, branches of the search tree are gradually formed. For example, after the matching from u1 to u6 is completed, the finally obtained coverage point set is {v1, v4, v3, v5, v6, v8}, and one of the branches 51 of the search tree 50 is obtained (see Figure 5 shown).
[0089] In this embodiment, operations 111-114 are pruned using the forward intersection pruning method, that is, the candidate sets of the query points that have not been matched are calculated in advance. For example, the candidate sets of the right neighbors of each query point are calculated in advance, so as to achieve early filtering and safe pruning.
[0090] Specifically, during the matching process, for the current query point u iFor each neighbor, the candidate sets of other query points that have not been matched yet are calculated in advance through forward intersection (right intersection), thereby filtering out the points that do not meet the conditions and reducing the subsequent calculation amount. In addition, the forward intersection method not only calculates the candidate set of the query point u (i+1) (i.e., the query point that is next to the current query point u i in the matching order), but also calculates the candidate sets of all right neighbors u j , ensuring more efficient pruning.
[0091] Specifically, when the matching reaches a certain query point, after matching the data at this query point and updating the candidate sets of its right neighbors, if there is a right neighbor with an empty candidate set, it means that the data points in the data graph that match this right neighbor have been occupied in the previous matching process, and there are no matching data points for this right neighbor. The branch starting from this query point is directly pruned to avoid redundant calculations.
[0092] Based on this, using the forward intersection method can effectively reduce the search space and improve the matching efficiency due to its ability to prune in advance.
[0093] In operation 120, it is judged whether the current query point reaches or exceeds the pruning depth. If the current query point reaches or exceeds the pruning depth, the process proceeds to 121. If the current query point does not reach the pruning depth, the process jumps to operation 126.
[0094] In operation 121, it is judged whether all the data points matched by the current query point and the data points in the candidate sets of the query points after the current query point are included in the covered point set.
[0095] The pruning depth corresponds to a specific query point. Before this specific query point becomes the current query point, the candidate sets of one or more query points have not been calculated. When this specific query point becomes the current query point, after updating the candidate sets of the right neighbors of this specific query point, all the candidate sets of the query points have been calculated at this time.
[0096] Specifically, in one embodiment, for a specific matching order, the query point corresponding to the pruning depth d is u d , and the pruning depth d satisfies the following conditions:
[0097] Any and
[0098] There exists
[0099] wherein, u d is the query point corresponding to the pruning depth d, and LN(u j ) refers to the set of left neighbors of the query point u j .
[0100] That is to say, in a matching order, the query point u d The query point u on the right j At least one left neighbor exists among the query points on the left of the query point u d According to the forward intersection pruning method, when the enumeration depth (i.e., the position of the current query point reached by sequential matching in the matching order) reaches the pruning depth d, the candidate sets of all query points have been calculated. Because if the candidate set is empty, the empty set pruning has been triggered. For example, in Figure 3-4 In the example given, the pruning depth d is 2, that is, when matching the point u2, the candidate sets of all query points have been calculated. If the data points already matched by the current query point and the data points in the candidate sets of the query points after the current query point are all included in the covering point set, the process proceeds to operation 123; if the data points already matched by the current query point and the data points in the candidate sets of the query points after the current query point are not all included in the covering point set, the process proceeds to operation 126.
[0101] In operation 123, full-coverage pruning is triggered to avoid re-matching and searching starting from the data points matched by the current query point this time, and the process returns to operation 115.
[0102] For example, in the case shown in Figure 5 Assume that the current query point is u2, v4 has been matched before the current query point u2, and the data point matched for u2 this time is v6. When it is found that the data points already matched by the current query point u2, namely v4 and v6, and the data points in the candidate sets of the subsequent query points u3 - u6 are all included in the covering point set {v1, v4, v3, v5, v6, v8}, full-coverage pruning is triggered, that is, terminating the re-matching of data points at the current query point u2 and the subsequent matching of data points of the subsequent query points caused thereby.
[0103] Operations 120 - 123 perform pruning using the full-coverage pruning method, that is, by determining whether the data points already matched by the current query point and the data points in the candidate sets of the subsequent query points are all included in the obtained covering point set. If so, the current branch can be pruned to avoid further calculation, thereby achieving the purpose of reducing redundant calculation and accelerating the formation of the covering set.
[0104] In one implementation, in the matching order generated in operation 101, the isolation points are placed at the end part of the matching order, as shown in Figure 6 Part (a) of Figure 6 Part (a) of is a schematic diagram of the matching order of another example, Figure 6 Part (b) of is another data graph. In this implementation, the isolation point is the query point u without a right neighbor i , for example, in Figure 6In part (a), since query points u4 and u6 have no right neighbors, they are regarded as isolated points and placed at the end of the matching order, that is, u4 and u5 are swapped, so that the isolated points u4 and u5 are both placed at the end of the matching order.
[0105] Please continue to refer to Figure 1 , in one embodiment, the sub-graph matching coverage set method of the present application further includes the following operations.
[0106] In operation 126, it is judged whether the current query point reaches or exceeds the expansion depth. If so, the process proceeds to operation 127; if not, the process proceeds to operation 129.
[0107] Specifically, in this embodiment, the multi-expansion depth is the position where the first isolated point appears in the matching order.
[0108] In operation 127, for each query point u that is an isolated point i , one by one, use the unmatched data points in the candidate set of the isolated point to replace the currently matched data point of the isolated point on the current branch of the search tree, form a solution and update the coverage point set, and the process returns to operation 115.
[0109] Among them, before replacing the matched data point of the isolated point, the unmatched data point does not exist on the current branch.
[0110] Please refer to Figure 7 As shown, operation 127 further includes the following operations.
[0111] In operation 701, judge the number of isolated points. In the case of only one isolated point, the process proceeds to operation 703; in the case of two or more isolated points, the process proceeds to operation 705.
[0112] In operation 703, one by one, use the unmatched data points in the candidate set of the isolated point to judge whether the unmatched data point already exists on the current branch (that is, judge whether the unmatched data point has been occupied by other query points). In the case where the unmatched data point does not exist on the current branch, replace the currently matched data point of the isolated point with the unmatched data point on the current branch, form a solution and update the coverage point set, and the process returns to operation 115. In this embodiment, the coverage point set is a set containing all the matched data points.
[0113] In the case where the unmatched data point exists on the current branch, the process returns to operation 115.
[0114] In operation 705, in the case of two or more isolated points, judge whether there is at least one group of isolated points with the same label among the isolated points.
[0115] In operation 707, in the case where there are no isolation points with the same label among the isolation points, each unmatched data point in the candidate set of each isolation point is used one by one to determine whether the unmatched data point already exists on the current branch. If the unmatched data point does not exist on the current branch, then on the current branch, the unmatched data point replaces the currently matched data point of the corresponding isolation point, a solution is formed, and the covered point set is updated, and the process returns to operation 115.
[0116] In operation 709, in the case where there is at least one group of isolation points with the same label among the isolation points, a group of isolation points with the same label is selected from the isolation points to form an isolation point group. Combinations are formed using the unmatched data points in the candidate set of each isolation point in the isolation point group. The data points within the combination are used to replace the currently matched data points of the corresponding isolation points on the current branch, a solution is formed, and the covered point set is updated, and the process returns to operation 115.
[0117] Among them, each combination formed satisfies the following conditions:
[0118] 1. The data points in the combination form a one-to-one correspondence with the isolation points in the isolation point group. That is, the data points in the combination are composed of one unmatched data point selected from the candidate set of each isolation point in the isolation point group;
[0119] 2. The data points in the combination are all different from each other;
[0120] 3. The data points in the combination do not exist on the current branch.
[0121] In operation 129, the next query point in the matching order after the current query point is changed to the current query point, that is, the value of i is changed to (i + 1), and the process returns to 109.
[0122] Please refer to Figure 8 as shown, for Figure 6 the matching order shown in part (a) of Figure 6 to perform a matching search in the data graph shown in part (b) of Figure 6 to form a schematic diagram of the branches of the search tree. Since the query points u4 and u6 are isolation points without right neighbors, u4 and u6 are placed at the end of the matching order. The candidate set of u4 is Pu4 = {v1, v2, v8}; the candidate set of u6 is Pu6 = {v9, v 10 , ..., v 100}. At this time, on the branch 81 of the search tree 80, the unmatched data points v2, v8 in the candidate set of u4 are respectively used to replace the currently matched data point v1, and the unmatched data points v 10 , v 11 ,..., v 100Replace the currently matched data point v9 to form a new solution and update the set of covered points. It should be noted that according to Figure 6 the data graph shown in part (b) of
[0123] Operation 127 uses the multi-expansion method to expand the set of covered points, so that the updated set of covered points contains all data points corresponding to any query point. In addition, when some isolated points have the same label, the multi-expansion method matches multiple isolated query points (isolated points) at the same time, reducing the backtracking operation and greatly improving the matching efficiency.
[0124] It should be noted that although the introduction of operations 101-129 is carried out based on Figure 1 the order shown, in fact, some operations can be carried out simultaneously, or the order of some operations can be swapped. For example, the forward intersection pruning and the full coverage pruning are carried out alternately. In this way, until all query points are matched to form a solution. In this process, if any link meets the pruning condition, pruning as described in operations 114 and 123 is carried out.
[0125] The embodiment of the present application also provides a subgraph matching coverage set device. The subgraph matching coverage set device can be a computer, a server, a data center, etc., which is not limited here. Please refer to Figure 9 shown, which is the hardware structure diagram of the subgraph matching coverage set device for running the method of the present application. The device 90 includes a processor 901, a memory 903, a communication bus 905, at least one network interface 907 or a user interface 909. The user interface 909 can be, for example, a display, a keyboard or a pointing device. The memory 903 can be a high-speed RAM memory or a non-volatile memory, such as at least one disk memory. The memory 903 stores execution instructions. When the processor 901 runs, communication occurs between the processor 901 and the memory 903. The processor 901 calls the instructions stored in the memory 903, that is, the subgraph matching coverage set device 100, to execute the above-mentioned subgraph matching coverage set method. The device 90 also includes an operating system 911, and the operating system 911 contains various programs for implementing various basic services and processing hardware-based tasks.
[0126] For the device 90 provided in the embodiments of the present application, the processor 901 may perform the operations included in the above-mentioned subgraph matching covering set method to implement the prediction of a short time series based on memory enhancement. The implementation principle and technical effects are similar to those of the subgraph matching covering set method introduced in the foregoing embodiments, and will not be elaborated here specifically.
[0127] The embodiments of the present application also provide a computer-readable storage medium. The computer-readable storage medium stores computer-executable instructions, which may be, for example, the memory 903. The computer-executable instructions can enable a processing unit (such as the processor 901) to execute the subgraph matching covering set method described in the above embodiments. The implementation principle and technical effects are similar to those of the subgraph matching covering set method introduced in the foregoing embodiments, and will not be elaborated here.
[0128] The embodiments of the present application also provide a computer program product. The computer program product includes computer-executable instructions. When the computer-executable instructions are executed by a processing unit (such as the processor 901), the subgraph matching covering set method described in the above embodiments is implemented. The implementation principle and technical effects are similar to those of the subgraph matching covering set method introduced in the foregoing embodiments, and will not be elaborated here.
[0129] In summary, compared with the prior art, the subgraph matching covering set method, device, storage medium, and computer program product provided by the embodiments of the present application use matching coverage to replace full coverage, saving search time and improving search efficiency. Further, (1) by identifying covered vertices, redundant branches are effectively pruned, unnecessary calculations are reduced, and the algorithm efficiency is improved. (2) The candidate set of unmatched vertices is pre-computed to enhance the pruning effect and ensure a more accurate matching process. (3) Allowing multiple isolated vertices to be matched simultaneously reduces backtracking operations and greatly improves the matching efficiency.
[0130] The above are only the embodiments of the present application, and do not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the content of the specification and drawings of the present application, or directly or indirectly applied in other related technical fields, shall be equally included in the patent protection scope of the present application.
Claims
1. A subgraph matching covering set method, characterized in that, The method includes: Receiving a given query graph to obtain a target subgraph in a data graph; Obtaining a query point set containing all query points and a matching order containing all query points according to the query graph; According to the matching order, starting from the first query point, calculating the candidate set of the current query point according to the data graph; Obtaining the right neighbor of the current query point in the matching order and the candidate set of each right neighbor; Matching a data point for the current query point according to the candidate set of the current query point; Updating the candidate set of the right neighbor of the current query point according to the data point matched by the current query point; and Performing a forward intersection pruning method and / or a full coverage pruning method during the matching process, wherein, the forward intersection pruning method includes: Judging whether there is an empty set in the candidate set of the updated right neighbor; If there is an empty set, it is judged that the matched data point is invalid, the current branch is directly pruned, and the update of the candidate sets of all right neighbors caused by the matching of this data point is invalid, and the candidate sets of all right neighbors are returned to the unupdated state. When the current query point is not the first query point, backtrack to the previous query point to rematch the data point, The full coverage pruning method includes: Judging whether the current query point reaches or exceeds the pruning depth, wherein the pruning depth corresponds to a specific query point. Before this specific query point becomes the current query point, the candidate sets of one or more query points have not been calculated, and when this specific query point becomes the current query point, after updating the candidate sets of the right neighbors of this specific query point, at this time, the candidate sets of all query points have been calculated; If the current query point reaches or exceeds the pruning depth, judging whether the matched data points of the current query point and the data points in the candidate sets of the query points after the current query point are all included in the coverage point set; and If the matched data points of the current query point and the data points in the candidate sets of the query points after the current query point are all included in the coverage point set, pruning from the data point matched by the current query point this time.
2. The method according to claim 1, characterized in that, The updating the candidate set of the right neighbor of the current query point according to the data point matched by the current query point includes: For the uninitialized right neighbor, calculating the candidate set of this right neighbor according to the data point matched by the current query point; and For the right neighbor with a non-empty original candidate set, taking the intersection of the original candidate set of this right neighbor and the newly calculated candidate set, and replacing the original candidate set with the intersection result as the updated candidate set of this right neighbor.
3. The method according to claim 2, wherein During the matching process, the forward intersection pruning method and the full coverage pruning method are performed. When it is judged that there is no empty set in the candidate set of the updated right neighbor, the operation of judging whether the current query point reaches or exceeds the pruning depth is performed.
4. The method according to claim 1, characterized in that, The pruning depth d satisfies the following conditions: Any and There exists Among them, u d is the query point corresponding to the pruning depth d, and LN(u j ) refers to the set of left neighbors of the query point u j .
5. The method according to claim 3, wherein In the matching order, the isolation point is placed at the end part, where the isolation point refers to a query point without a right neighbor. The method further includes: For each query point that is an isolation point, one by one, use the unmatched data points in the candidate set of the isolation point to replace the matched data points of the isolation point on the current branch. Wherein, before replacing the matched data points of the isolation point, the unmatched data points do not exist on the current branch.
6. The method according to claim 5, wherein Further included: In the case where there is at least one group of isolation points with the same label among the isolation points, select a group of isolation points with the same label from the isolation points to form an isolation point group, use the unmatched data points in the candidate set of each isolation point in the isolation point group to form a combination, and use the data points in this combination to replace the previous matched data point of the corresponding isolation point on the current branch to form a new solution and a new set of covered points. Wherein, the combination satisfies: The data points in the combination form a one-to-one correspondence with the isolation points in the isolation point group; The data points in the combination are different from each other; and The data points in the combination do not exist on the current branch.
7. The method according to claim 6, characterized in that, Further included: Judge whether the current query point reaches or exceeds the expansion depth, and the multi-expansion depth is the position where the first isolation point appears in the matching order; And If the current query point reaches or exceeds the expansion depth, perform the operation of, for each query point that is an isolation point, one by one, using the unmatched data points in the candidate set of the isolation point to replace the matched data points of the isolation point on the current branch.
8. A subgraph matching covering set device, characterized in that, The subgraph matching coverage set device includes: A memory configured to store computer-executable instructions; A processor configured to run the computer-executable instructions stored in the memory to implement the subgraph matching coverage set method according to any one of claims 1-7.
9. A storage medium, characterized in that, The storage medium stores computer-executable instructions, and when the computer-executable instructions are run by a processing unit, the subgraph matching coverage set method according to any one of claims 1-7 is implemented.
10. A computer program product, characterized in that, Including computer-executable instructions, and when the computer-executable instructions are executed by a processing unit, the subgraph matching coverage set method according to any one of claims 1-7 is implemented.