Regular path query method and system based on materialized view selection and query planning
Patent Information
- Application Number
- CN202410680530.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-29
- Publication Date
- 2026-09-29
- Estimated Expiration
- 2044-05-29
AI Technical Summary
[0019]一、正则路径查询的物化视图选择:现有技术以总内存占用而非查询负载的总查询代价为优化目标,不符合降低查询代价的现实需求;同时,不带剪枝地枚举所有可行的物化视图组合时间复杂度高,即使在很小的查询负载上也需耗费大量时间
[0062]一、正则路径查询的物化视图选择:现有技术考虑的优化目标为最小化总内存占用,而目前对处理多个正则路径查询的优化需求主要集中于降低查询执行的时间代价,内存占用存在一定的上限,但在上限之内占用多少并不是主要的性能考量,因此比起现有技术的优化目标,本发明的优化目标(在内存预算内最小化查询负载的总执行时间代价)更切合应用场景的实际需求。另外,比起现有技术不带剪枝地枚举所有可行的物化视图组合的方法,本发明提出的启发式算法更高效、可扩展性更优。
Smart Images

Figure CN118503280B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of information technology, specifically relating to a regular path query method and system based on materialized view selection and query planning. Background Technology
[0002] This invention utilizes materialized view technology in databases to solve the query optimization problem of multiple regular expression path queries. The background technologies cover areas such as materialized view selection for regular expression path queries, query planning for regular expression path queries, query optimization of multiple regular expression path queries, and query optimization of multiple relational database queries. The relevant background technologies are as follows:
[0003] I. Materialized View Selection for Regular Path Queries. The only existing technology for materialized view selection for regular path queries is [1]. [1] Considering the materialized view selection problem with the optimization objective of the total memory usage of the materialized view, the regular path query that needs to build a materialized view is first decomposed into subqueries that can be used as candidate materialized views, and then all feasible materialized view combinations are enumerated, and the one with the smallest total memory usage is selected as the result.
[0004] II. Query planning for regular path queries. The latest regular path query plan representation is Waveplan[2]. Waveplan extends finite automata in the following ways: 1) it allows forward and backward traversal on the data graph; 2) it allows the plan to consist of multiple automata; 3) it allows views (i.e., the results of other automata in the plan) as state transitions. With these extensions, Waveplan covers past techniques of regular path query planning, such as the strategy of starting the search from rare labels in the data graph[3]. [2] also proposed a cost model for estimating the efficiency of Waveplan for plan selection. Since Waveplan can treat views as state transitions, given a set of materialized views, the most efficient Waveplan based on these views can be selected using the method in [2]. Another class of techniques converts regular path queries into SQL statements or Datalog statements that support Kleene closures, and utilizes the existing query planning mechanisms of these languages to implement regular path query planning[4]. [5] A cost model for finite automata plans applicable to regular path queries is proposed, and this model is used to evaluate parallel regular path query plans. That is, the regular path query is split into a batch of subqueries and executed in parallel. The highest cost of the split subqueries is defined as the cost of this splitting method, so as to find the split with the lowest cost for parallel execution. [6] A regular path query algorithm based on a spatially efficient graph storage structure based on wavelet trees is proposed.
[0005] III. Query optimization for multiple regular path queries. Swarmguide[7] is a multiple regular path query optimization framework based on Waveplan[2]. Given a regular path query workload, Swarmguide uses the affinity propagation technique based on edge labels to group the queries according to the similarity between their minimum deterministic finite automata; then it calculates the maximum common sub-automata of each group of finite automata, selects the subquery corresponding to the sub-automata as the shared materialized view of the group, and performs query planning based on Waveplan. After calculating the results of the shared materialized views, Swarmguide regards them as the state transitions of the queries in the load and plans the queries in the load based on Waveplan. RTC[8] is a multiple query optimization framework for regular path query workloads containing a common Kleene closure. After extracting the common Kleene closure R * Or R + Then, the results are quickly calculated as follows: the paths in the label sequence that satisfy R are shrunk into edges (called edge shrinking), and each strongly connected component in the shrunk graph is shrunk into nodes (called node shrinking). The connected node pairs in the final graph are the result node pairs of the Kleene closure in the original graph. The results of the common parts and the non-common parts are connected to obtain the result of each load query.
[0006] IV. Query optimization for multiple relational database queries. Query optimization for multiple relational database queries is based on a multi-query plan representation called AND-OR Directed Acyclic Graph (AND-OR DAG) [9].
[10] proposed a materialized view selection algorithm based on AND-OR DAG. There are two types of nodes in AND-OR DAG: AND nodes represent algebraic operators, such as join operations; OR nodes represent subqueries. In AND-OR DAG, AND nodes and OR nodes are arranged in an alternating layer: the outgoing neighbor of an OR node can only be an AND node, which represents the way to calculate the result of a subquery; while the outgoing neighbor of an AND node can only be an OR node, which represents the operand of the operator. For example, Figure 1 The AND-OR directed acyclic graph in the diagram shows all possible ways to join lists A, B, and C, where OR nodes are represented by squares and AND nodes by circles. When an OR node has multiple child AND nodes (e.g., ABC), the child node with the lowest execution cost is selected according to the cost model.
[0007] The references mentioned above are as follows:
[0008] [1]S.Afonin,“The View Selection Problem for Regular Path Queries,”inLATIN 2008:Theoretical Informatics,vol.4957,E.S.Laber,C.Bornstein,L.T.Nogueira,and L.Faria,Eds.,in Lecture Notes in Computer Science,vol.4957.,Berlin,Heidelberg:Springer Berlin Heidelberg,2008,pp.121–132.doi:10.1007 / 978-3-540-78773-0_11.
[0009] [2]N.Yakovets,P.Godfrey,and J.Gryz,“Query Planning for EvaluatingSPARQL Property Paths,”in Proceedings of the 2016International Conference onManagement of Data,San Francisco California USA:ACM,Jun.2016,pp.1875–1889.doi:10.1145 / 2882903.2882944.
[0010] [3]A.Koschmieder and U.Leser,“Regular Path Queries on Large Graphs,”in Scientific and Statistical Database Management,vol.7338,A.Ailamaki andS.Bowers,Eds.,in Lecture Notes in Computer Science,vol.7338.,Berlin,Heidelberg:Springer Berlin Heidelberg,2012,pp.177–194.doi:10.1007 / 978-3-642-31235-9_12.
[0011] [4]S.Dey,V.Cuevas-Vicenttín,S. E.Gribkoff,M.Wang,and B. “On implementing provenance-aware regular path queries withrelational query engines,”in Proceedings of the Joint EDBT / ICDT 2013Workshopson-EDBT’13,Genoa,Italy:ACM Press,2013,p.214.doi:10.1145 / 2457317.2457353.
[0012] [5]V.-Q.Nguyen,Q.-T.Huynh,and K.Kim,“Estimating searching cost ofregular path queries on large graphs by exploiting unit-subqueries,”JHeuristics,vol.28,no.2,pp.149–169,Apr.2022,doi:10.1007 / s10732-018-9402-0.
[0013] [6]D.Arroyuelo,A.Hogan,G.Navarro,and J.Rojas-Ledesma,“Time-and Space-Efficient Regular Path Queries,”in 2022 IEEE 38th International Conference onData Engineering(ICDE),Kuala Lumpur,Malaysia:IEEE,May 2022,pp.3091–3105.doi:10.1109 / ICDE53745.2022.00277.
[0014] [7]Z.Abul-Basher,"Multiple-Query Optimization of Regular PathQueries,"2017 IEEE 33rd International Conference on Data Engineering(ICDE),San Diego,CA,USA,2017,pp.1426-1430,doi:10.1109 / ICDE.2017.205.
[0015] [8] I.Na, Y.-S.Moon, I.Yi, K.-Y.Whang, and SJHyun, "Regular Path QueryEvaluation Sharing a Reduced Transitive Closure Based on Graph Reduction," in 2022 IEEE 38th International Conference on Data Engineering (ICDE), Kuala Lumpur, Malaysia: IEEE, May 2022, pp.1675–1686.doi:10.1109 / ICDE53745.2022.00171.
[0016] [9]Nicholas Roussopoulos.1982.View indexing in relationaldatabases.7,2(jun 1982),258–290.https: / / doi.org / 10.1145 / 319702.319729.
[0017]
[10] Prasan Roy, S. Seshadri, S. Sudarshan, and Siddhesh Bhobe. 2000. Efficient and extensible algorithms for multi query optimization. In Proceedings of the 2000 ACM SIGMOD International Conference on Management of Data (Dallas, Texas, USA) (SIGMOD'00). Association for Computing Machinery, New York, NY, USA, 249–260. https: / / doi.org / 10.1145 / 342009.335419.
[0018] The disadvantages of existing technologies are as follows:
[0019] I. Materialized View Selection for Regular Path Queries: Existing technologies optimize based on total memory usage rather than total query cost, which does not meet the practical need to reduce query cost. At the same time, enumerating all feasible materialized view combinations without pruning has high time complexity, even for very small query loads.
[0020] II. Query planning for regular expression path queries: Existing techniques for query planning of regular expression path queries are only used to optimize a single regular expression path query and cannot be directly extended to the query optimization problem of multiple regular expression path queries.
[0021] III. Query Optimization of Multiple Regular Expression Path Queries: Existing techniques for optimizing multiple regular expression path queries make certain assumptions about the common subqueries of the query load, which may not hold true in real query loads; the speed of materialized view selection and the acceleration effect of the selected materialized view on the load query both need to be improved.
[0022] IV. Query Optimization for Multiple Relational Database Queries: Relational database queries differ from graph regular expression path queries in syntax and semantics, especially since they do not support Kleene closures, making them unsuitable for query optimization of multiple regular expression path queries. Summary of the Invention
[0023] To address the aforementioned problems, this invention provides a regular expression path query method and system based on materialized view selection and query planning.
[0024] The technical solution adopted in this invention is as follows:
[0025] A regular expression path query method based on materialized view selection and query planning includes the following steps:
[0026] Given a regular path query load and a directed graph with edge labels, use the query planner to generate a multi-query plan for the query load.
[0027] Materialized views are selected using a materialized view selector to minimize the total query cost of the query load, and redundant views are detected and removed using the multi-query plan.
[0028] During the materialized view selection process, incremental updates are performed on multiple query plans;
[0029] The executor performs load queries based on multiple query plans and with the help of materialized views.
[0030] Furthermore, the directed graph with edge labels is an AND-OR directed acyclic graph with closure; the AND-OR directed acyclic graph with closure is obtained by adding support for Kleene closure to the AND-OR directed acyclic graph, and also contains AND nodes and OR nodes, and satisfies the following conditions:
[0031] Each AND node is marked by a regular expression path query operator;
[0032] Each OR node is marked by a distinct regular path query, and the node marked by the regular path query in the load S is the root node;
[0033] Each AND node's child nodes are OR nodes, which are marked by the operands of the regular expression path query operator that marked it.
[0034] Each OR node's child nodes are AND nodes marked by the lowest priority operator in the regular expression path query associated with it.
[0035] Furthermore, the input to the query planner is: a closure-based AND-OR directed acyclic graph (aod); a regular path query load (S); a labeled directed graph (G) with edges; a cost function (cost) that maps regular path queries to execution costs; and a cardinality function (card) that maps regular path queries to cardinality. The output is: a labeled closure-based AND-OR directed acyclic graph (aod), where the label contains two parts: the cost and cardinality estimate for each node; each OR node with multiple child nodes needs to label its target child, i.e., the child AND node selected during execution.
[0036] Furthermore, the materialized view selector performs materialized view selection using the following steps:
[0037] The first step is to sort all candidate materialized views in descending order of their frequency of occurrence in the query load and initialize the total memory usage to 0.
[0038] The second step is to apply the following steps to each of the sorted candidate materialized views:
[0039] 1) If the current candidate materialized view has been used 0 times, or if the current total memory usage plus the card value of the current candidate materialized view is greater than the total memory usage limit b, skip it;
[0040] 2) Using the AND-OR directed acyclic graph (AOD) with closure and the current candidate materialized view as input, call the method to incrementally update the query plan based on the selected materialized view. Obtain the amount by which the current candidate materialized view is expected to reduce the total execution cost of the query load and the amount of change in the query plan. If the amount of reduction in the total execution cost is 0 and the card value of the current candidate materialized view is not 0, skip it.
[0041] 3) Using the changes in the current candidate materialized view and the query plan as input, call the method to apply incremental updates; add the current candidate materialized view to v; add the card value of the current candidate materialized view to the total memory usage;
[0042] 4) Check the usage count of each materialized view in v. If it is 0, delete the materialized view from v and subtract its card value from the total memory usage.
[0043] Furthermore, the incremental update of the multi-query plan includes:
[0044] Initialize the total execution cost reduction to 0 and the query plan change to empty; take the AOD, the current node, the current node's card value, the total execution cost reduction, and the query plan change as inputs, and call the single-node incremental update method;
[0045] The single-node incremental update method takes as input a closure-based AND-OR directed acyclic graph (AOD), a node in the AOD, the node's new execution cost, the total execution cost reduction, and the change in the query plan; and outputs the updated total execution cost reduction and the change in the query plan. The steps of the single-node incremental update method are as follows:
[0046] First, if the new execution cost of the input node is greater than or equal to the current cost value of the node, return;
[0047] The second step is to add the following to the current total execution cost reduction: the difference between the node's current cost value and the new execution cost of the input node multiplied by the frequency of the node in the load if the regular path query represented by the current node is in the load.
[0048] The third step is to assign the current node's cost value as the new execution cost of the input node;
[0049] The fourth step is to call the single-node incremental update method on each parent node of the current node. The input remains unchanged except for the new execution cost of the node. The new execution cost of the node varies depending on the node type.
[0050] a) If the current node is an OR node: the new execution cost of the node remains unchanged;
[0051] b) If the current node is an AND node: The new execution cost of the node is set to the cost value estimated based on the operator corresponding to the parent node, the cost of all its child nodes, and the card value.
[0052] Furthermore, the executor performs load queries based on a multi-query plan and with the aid of a materialized view, including:
[0053] The single-node query execution method is invoked for the load query q;
[0054] The input to the single-node query execution method is a node in a directed acyclic graph with closure (AND-OR), a starting node candidate set lCand, and a target node candidate set rCand; the output is the query result for that node; the steps of the single-node query execution method are as follows:
[0055] 1) If the current node is an OR node: If the query corresponding to the current node is a single edge label or has been materialized, then directly obtain its complete query result from the graph data or materialized view, and then filter it with lCand and rCand to get the final result; otherwise, call the single node query execution method on its target child, keep the starting and target node candidate sets unchanged, and return the returned result as the final result.
[0056] 2) If the current node is an AND node, the execution method is determined based on the regular expression path query operator represented by the current node.
[0057] A regular expression path query system based on materialized view selection and query planning, comprising:
[0058] A query planner that generates multiple query plans for a given regular path query load and a directed graph with edge labels.
[0059] A materialized view selector is used to select materialized views to minimize the total query cost of the query load, and to detect and remove redundant views using the multi-query plan; and to perform incremental updates to the multi-query plan during the materialized view selection process;
[0060] An executor that performs load queries based on multiple query plans using materialized views.
[0061] Compared with the prior art, the advantages and beneficial effects of the present invention are as follows:
[0062] I. Materialized View Selection for Regular Expression Path Queries: Existing technologies prioritize minimizing total memory usage. However, current optimization requirements for handling multiple regular expression path queries primarily focus on reducing query execution time. While memory usage has an upper limit, the amount used within that limit is not a primary performance consideration. Therefore, compared to existing technologies, the optimization objective of this invention (minimizing the total execution time cost of the query load within the memory budget) better suits the actual needs of the application scenario. Furthermore, compared to existing methods that enumerate all feasible materialized view combinations without pruning, the heuristic algorithm proposed in this invention is more efficient and has better scalability.
[0063] II. Query planning for regular path queries: Existing technologies for query planning of regular path queries are only used to optimize a single regular path query, while the AND-OR directed acyclic graph with closure proposed in this invention is specifically designed for the query optimization problem of multiple regular path queries, which is helpful for the joint optimization of multiple regular path queries.
[0064] III. Query Optimization for Multiple Regular Expression Path Queries: Existing techniques for optimizing multiple regular expression path queries make certain assumptions about the common subqueries of the query load. However, this invention makes no assumptions about the distribution of the query load and can handle any query load. Existing techniques have room for improvement in the speed of materialized view selection and the acceleration effect of the selected materialized view on the load query. The method proposed in this invention explicitly aims to minimize the total execution time cost of the load query and introduces an incremental plan update method to reduce the time overhead of materialized view selection itself. It has significantly improved the speed of materialized view selection and the acceleration effect of the selected materialized view on the load query compared to existing methods.
[0065] IV. Query Optimization for Multiple Relational Database Queries: Existing technologies do not support Kleene closures. This invention proposes an AND-OR directed acyclic graph with closures as a representation of multiple regular path query plans to natively support Kleene closures, and proposes corresponding cost and cardinality estimation methods. Attached Figure Description
[0066] Figure 1 This is an example of an AND-OR directed acyclic graph.
[0067] Figure 2 This is a flowchart of the method of the present invention. Detailed Implementation
[0068] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to specific embodiments and accompanying drawings.
[0069] First, let's define the problem.
[0070] First, let's introduce the definition of regular path lookup: Regular path lookup is a regular expression using the set of edge labels Σ of a directed graph G = (V, E, Σ, l) as the alphabet. It is represented as a recursive formula, like R → ∈ |a|a - |R1 / R2|R? |R * |R + Where V represents the set of nodes, E represents the set of edges, l represents the function that maps edges to their labels, R represents a regular path query, ∈ represents an empty string, a represents a single edge label, | represents an OR operation, - represents the reverse operation, R1 and R2 each represent an arbitrary regular path query, ? represents zero or one matching operation, R * This represents the Kleene star operation, which involves zero or any positive integer number of matches. R + This indicates the Kleene addition operation, which means any positive integer number of matches.
[0071] The result of a regular expression path query is a set of path endpoint pairs whose tag sequences satisfy the regular expression. Among them, [[R]] G Let represent the result of a regular path query R on graph G, where p(u,v) represents the path with starting node u and destination node v, l(p(u,v)) represents the label sequence of this path, and L(R) represents the regular expression language corresponding to the regular path query R, i.e., the set of strings that R can recognize. Specifically, [[∈]] G ={<v,v> |v∈V}, because the path length from each node to itself is 0, the label sequence is defined as an empty string. Given a directed graph G with labeled edges, any graph of the form ... <R,[[R]] G The result pairs of the regular expression path query are all regular expression path materialized views. The regular expression path query load S = {R1, ..., R...} k Let R be a finite set of regular paths, where R is a set of paths of choice. i Let R represent an arbitrary regular expression path query; to support loads with duplicate queries, let R be an arbitrary path. i The frequency in S is freq[R] i ].
[0072] The problem addressed by this invention is the materialized view selection problem for regular path queries, defined as follows: given a directed graph G with labeled edges, a regular path query load S, and a maximum total memory usage b, return a set of materialized views. The following conditions must be met:
[0073] 1) Minimize the total query cost. in For based on The time cost of executing R;
[0074] 2) Meet the total memory usage constraint.
[0075] The method proposed in this invention will be described below.
[0076] like Figure 2 As shown, the method proposed in this invention comprises two processes: 1) query planning (left side) and 2) materialized view selection (right side). Given a regularized path query load and a directed graph with edge labels, the query planner generates a multi-query plan for the load. Then, the materialized view selector selects a set of materialized views to minimize the total query cost of the load and uses the multi-query plan to detect and remove redundant views. During the view selection process, the multi-query plan is incrementally updated to reflect how query execution should utilize the selected views. Finally, the executor executes the load query using the materialized views based on the multi-query plan.
[0077] This invention proposes a novel representation of multi-regular path query plans, called the AND-OR DAG with Closure, which forms the basis for materialized view selection and view-based query planning in regular path queries. The AND-OR DAG with Closure is derived by adding support for Kleene closures to the multi-relational query plan representation of the AND-OR DAG. It also contains AND and OR nodes and satisfies the following conditions:
[0078] 1) Each AND node is marked by a regular path query operator (i.e., / , |, ?, * or +).
[0079] 2) Each OR node is marked by a distinct regular path query. The nodes marked by the regular path query in the load S are the root nodes, meaning they have no parent nodes.
[0080] 3) The child nodes of each AND node are OR nodes marked by the operands of the regular path query operator that marked it.
[0081] 4) The child nodes of each OR node (if any) are AND nodes marked by the lowest priority operator in the regular expression path query associated with it. When there are multiple lowest priority operators ( / or |), each / constitutes a child node, while all | nodes are merged into one child node, with all operands as its children, because the execution order of / affects query time, while the execution order of | does not.
[0082] Given a regular path query load S, the method to construct a closed AND-OR directed acyclic graph is as follows: for each query load, construct from top to bottom in descending order of operator precedence, sharing OR nodes where possible. The order of the load queries does not affect the final structure of the closed AND-OR directed acyclic graph.
[0083] The following describes the five methods proposed in this invention based on AND-OR directed acyclic graphs with closures: query planning method, query execution method, materialized view selection method, method for incrementally updating the query plan based on the selected materialized view, and cost and cardinality estimation method.
[0084] I. Query Planning Method:
[0085] The inputs to this method are: a closure-based AND-OR directed acyclic graph (aod); a regular path query load (S); a labeled directed graph (G) with edges; a cost function (cost) that maps regular path queries to execution costs; and a cardinality function (card) that maps regular path queries to cardinality (i.e., the number of results). Specific estimation methods for cost and cardinality will be explained in "V. Cost and Cardinality Estimation Methods".
[0086] The output of this method is: a labeled AND-OR directed acyclic graph (aod) with closure. The label consists of the following two parts:
[0087] 1) Cost and cardinality estimates for each node;
[0088] 2) Each OR node with multiple child nodes needs to mark its target child, i.e., the child AND node selected during execution.
[0089] The steps of this method are as follows:
[0090] For each query in the regular path query load S, invoke the following single-node query planning method.
[0091] The input to the single-node query planning method is a node in a closure-based AND-OR directed acyclic graph; the output is the labeled node; the steps of the single-node query planning method are as follows:
[0092] First, if the current node is already marked, exit.
[0093] The second step is to add the cost and card values (obtained directly from the graph data) to the current node's marker if the current node has no child nodes, and then exit.
[0094] The third step is to call the single-node query planning method on all child nodes of the current node.
[0095] Fourth step: If the current node is an OR node, select the child node with the smallest cost value as the target child, set the target child's cost and card values to its own cost and card values, and add the above information to its flag; otherwise, if the current node is an AND node, estimate its cost and card values based on its corresponding regular expression path query operator and the cost and card values of all child nodes. If the operator corresponding to the current node is / , it is necessary to further select the child node execution order as left to right or right to left, and add the above information to its flag.
[0096] II. Query execution method:
[0097] The input to this method is: a closure-based AND-OR directed acyclic graph (aod); a load query q∈S; and a labeled directed graph G. (Note: This method can also perform regular path queries not in the original query load, but it requires first constructing a closure-based AND-OR directed acyclic graph, merging it with the original directed acyclic graph, and then calling the single-node query planning method described above.)
[0098] The output of this method is: the result of the load query [[q]]. G
[0099] The steps of this method are as follows:
[0100] Call the following single-node query execution method on q.
[0101] The single-node query execution method takes as input a node in a closure-based AND-OR directed acyclic graph, a starting node candidate set lCand, and a target node candidate set rCand; the output is the query result for that node; the steps of the single-node query execution method are as follows:
[0102] 1) If the current node is an OR node: If the query corresponding to the current node is a single edge label or has been materialized, the complete query result can be obtained directly from the graph data or materialized view, and then filtered by lCand and rCand to obtain the final result and return it; otherwise, call the single node query execution method on its target child (selected in the query planning method), keep the starting and target node candidate sets unchanged, and return the returned result as the final result.
[0103] 2) If the current node is an AND node: determine the execution method based on the regular expression path query operator represented by the current node.
[0104] a) If the current node represents / : The execution order (left or right child node) is determined based on the chosen execution order in the query planning method. Taking the left child node as an example: First, call the single-node query execution method on the left child node, keeping the starting node candidate set unchanged and setting the target node candidate set to empty. The result is denoted as lRes. Second, call the single-node query execution method on the right child node, keeping the target node candidate set unchanged. If lRes contains [[∈]]... G If the flag is set to lRes, the candidate set of starting nodes is set to empty; otherwise, it is set to the target node set of lRes, and the result is denoted as rRes. Finally, a join operation is performed on lRes and rRes to obtain the final result and return it. The steps for the right child node are executed first, and so on.
[0105] b) If the current node represents |: Call the single-node query execution method for each child node, keeping the starting and target candidate sets unchanged. Finally, perform a union operation on the results of each child node to obtain the final result and return it.
[0106] c) If the current node represents ?: Call the single-node query execution method on the child nodes, and add [[∈]] to the result. G The flag is returned.
[0107] d) If the current node represents + or *: After calling the single-node query execution method on the child nodes, perform fixed-point iteration until no more new results are generated. Before returning the result, if the current node represents *, additionally add [[∈]]. G The marker. Returns the result.
[0108] III. Materialized View Selection Method:
[0109] (a) Sub-procedure: Method for maintaining the number of times a view is used in a query plan
[0110] Since the core of redundant view detection in the materialized view selection method lies in detecting whether the current view is selected in the query plan chosen by the load query, this invention proposes a method for maintaining the number of times a view is used in the query plan, as described below:
[0111] Input: A closure AND-OR directed acyclic graph aod; a node in aod; the change in the number of times it is used, δ.
[0112] Output: AOD updated for the number of times each node has been used.
[0113] The steps are as follows:
[0114] The first step is to increment the usage count of the current node by δ. If the current node has no child nodes, then return, i.e., end the execution of this sub-process.
[0115] The second step is to call the maintenance method of the number of times the view is used in the query plan for its target child node (selected in the query plan), and the δ value remains unchanged; otherwise, if the current node is an AND node, call the maintenance method of the number of times the view is used in the query plan for each of its child nodes, and the δ value remains unchanged.
[0116] After constructing the AND-OR directed acyclic graph with closures and calling the query planning method, it is necessary to call the maintenance method (δ=1) for the number of times the view is used in the query plan for each node corresponding to the load query to initialize the usage count of all nodes.
[0117] (II) Materialized View Selection Method
[0118] Input: A directed acyclic graph aod with closure and AND-OR; maximum total memory usage b.
[0119] Output: The selected set of materialized views, v.
[0120] The steps are as follows:
[0121] The first step is to sort all candidate materialized views in descending order of their frequency of occurrence in the query load and initialize the total memory usage to 0.
[0122] The second step is to apply the following steps to each of the sorted candidate materialized views:
[0123] 1) If the current candidate materialized view has been used 0 times, or if the current total memory usage plus the card value of the current candidate materialized view (i.e., the cardinality estimate) is greater than the total memory usage limit b, skip it.
[0124] 2) Using the AND-OR directed acyclic graph (AOD) with closure and the current candidate materialized view as input, call the method to incrementally update the query plan based on the selected materialized view, and obtain the amount by which the current candidate materialized view is expected to reduce the total execution cost of the query load and the amount of change in the query plan. If the reduction in total execution cost is 0 and the card value of the current candidate materialized view is non-zero, skip this step.
[0125] 3) Using the changes in the current candidate materialized view and the query plan as input, call the following method for incremental update: add the current candidate materialized view to v; add the card value of the current candidate materialized view to the total memory usage.
[0126] 4) Check the usage count of each materialized view in v. If it is 0, delete the materialized view from v and subtract its card value from the total memory usage.
[0127] (III) Sub-process: Applying incremental update method
[0128] Input: A closure AND-OR directed acyclic graph (AOD); a node in the AOD; the change in the query plan.
[0129] Output: AOD after applying incremental updates.
[0130] The steps are as follows:
[0131] The first step is to update the cost and card values of all relevant nodes in the AOD based on the changes in the query plan.
[0132] The second step is to call the maintenance method of the number of times the view is used in the query plan for each OR node whose target child node has changed in the query plan changes. For its old target child, call the maintenance method of the number of times the view is used in the query plan. Set δ to the negative number of times the OR node is used. For its new target child, call the maintenance method of the number of times the view is used in the query plan. Set δ to the number of times the OR node is used.
[0133] The third step is to call the maintenance method for the number of times the view is used in the query plan for the current node, with δ set to the opposite of the number of times the current node is used. After the call is completed, the number of times the current node is used is restored to the value before the call.
[0134] IV. Method for Incrementally Updating Query Plans Based on Selected Materialized Views
[0135] Input: A closure AND-OR directed acyclic graph aod; a node in aod representing the candidate materialized view currently under consideration; cost function that maps regular path queries to execution costs; card function that maps regular path queries to cardinality (i.e., the number of results).
[0136] Output: The amount by which the currently considered candidate materialized views are expected to reduce the total execution cost of the query load, and the amount of change in the query plan.
[0137] The steps are as follows:
[0138] Initialize the total execution cost reduction to 0 and the query plan change to empty; take the AOD, the current node, the current node's card value, the total execution cost reduction, and the query plan change as input, and call the following single-node incremental update method.
[0139] The single-node incremental update method takes as input a closure-based AND-OR directed acyclic graph (AOD), a node in the AOD, the node's new execution cost, the total execution cost reduction, and the change in the query plan; the output is the updated total execution cost reduction and the change in the query plan. The steps of the single-node incremental update method are as follows:
[0140] The first step is to return if the new execution cost of the input node is greater than or equal to the current cost value of the node, thus ending the execution of this process.
[0141] The second step is to add the following to the current total execution cost reduction: if the regular path query represented by the current node is in the load, then the difference between the current cost value of the node and the new execution cost of the input node is multiplied by the frequency of the node in the load.
[0142] The third step is to assign the current node's cost value as the new execution cost of the input node.
[0143] The fourth step is to call the single-node incremental update method on each parent node of the current node. The input remains unchanged except for the new execution cost of the node. The new execution cost of the node varies depending on the node type.
[0144] a) If the current node is an OR node: the new execution cost of the node remains unchanged.
[0145] b) If the current node is an AND node: The new execution cost of the node is set to the cost value estimated based on the operator corresponding to the parent node, the cost of all its child nodes, and the card value (including the updated cost value of the current node). Furthermore, after calling the single-node incremental update method, if the current node has the lowest execution cost among all the child nodes of the parent node, the current node is set as the target child node of the parent node; if the operator corresponding to the parent node is / , the execution order of the child nodes is reselected.
[0146] V. Cost and Baseline Estimation Methods
[0147] In query planning methods and methods that incrementally update query plans based on selected materialized views, cost and cardinality estimation methods are required to estimate the cost and cardinality of nodes in a closure-based AND-OR directed acyclic graph. Cardinality is the number of query results. The following cases will be discussed:
[0148] 1) If the current node is an OR node:
[0149] a) If the current node represents a single edge label: its cost and card values are the exact number of edges with this label in the graph, which can be obtained directly from the graph data.
[0150] b) Otherwise, the cost and card values of the current node are the same as the cost and card values of its target child node.
[0151] 2) If the current node is an AND node:
[0152] a) If the current node represents / :
[0153] Cost value: Where l1 is the label of the last edge of R1. Let R1 be the estimated size of the intersection of the target node set of l1 and the starting node set of R2, and let cost(R2) represent the time cost of computing the result of the regular path query R2, and cost(R1) represent the time cost of computing the result of the regular path query R1. G This represents the result of the regular expression path query R1, [[R2]]. G This represents the result of the regular expression path query R2. This invention uses Monte Carlo sampling to estimate... (The same applies below): Randomly sample several nodes from the target node set of l1, and execute R2 using a minimum deterministic finite automaton plan starting from them. If a result is found, return immediately and mark the node as having a result; otherwise, mark it as having no result. Finally, estimate the target node by multiplying the frequency of nodes with results in the sampled nodes by the total number of target nodes in l1.
[0154] Card value: Among them, [[l1]] G This represents the result of the regular expression path query l1, where t represents the target node and s represents the starting node.
[0155] The number of distinct starting nodes in the results:
[0156] The number of distinct target nodes in the results:
[0157] b) If the current node represents * or +:
[0158] Cost value: Where l1 is the label of the last edge of R. The value of D is determined as follows: when c < 1, c D-1 |[[R]] G |≥∈andc D |[[R]] G When |<∈; c≥1, take D=6.
[0159] Card value:
[0160] The number of distinct starting nodes in the result: |[[R * ]] G .s|=|[[R + ]] G .s|=|[[R]] G .s|
[0161] The number of distinct target nodes in the result: |[[R * ]] G .t|=|[[R + ]] G .t|=|[[R]] G .t|
[0162] c) If the current node represents |:
[0163] Cost value: The sum of the cost values of all child nodes;
[0164] Card value: The sum of the card values of all child nodes;
[0165] The number of distinct starting nodes in the result: the sum of the number of distinct starting nodes in all child node results;
[0166] The number of distinct target nodes in the result: the sum of the number of distinct target nodes in all child node results.
[0167] d) If the current node represents?
[0168] Cost value: The cost value of the child node;
[0169] Card value: The card value of the child node;
[0170] The number of distinct starting nodes in the result: The number of distinct starting nodes in the child node result;
[0171] The number of distinct target nodes in the result: The number of distinct target nodes in the child node result.
[0172] The key point of this invention is:
[0173] 1. This invention formally defines for the first time the regular path query materialized view selection problem that minimizes the query time cost for a given query load within a memory budget, and proposes an efficient method to solve this problem.
[0174] 2. This invention innovatively proposes a closure-based AND-OR directed acyclic graph as a joint plan representation for multiple regular path queries. It is the first joint plan representation for multiple regular path queries, which is helpful for the joint optimization of multiple regular path queries.
[0175] 3. This invention closely integrates materialized view selection with the query planning process based on materialized views, which can obtain the optimal query plan using these materialized views while selecting the optimal set of materialized views, and reduces additional overhead by using incremental query plan update technology.
[0176] 4. This invention proposes a query cost and cardinality estimation method that can cover the complete regular path syntax (including Kleene closures), based on Monte Carlo sampling technique, and has good scalability.
[0177] Other embodiments of the present invention:
[0178] 1. The present invention can also be realized if the AND-OR directed acyclic graph with closure is not explicitly represented as a directed acyclic graph, but expressed in other data structures (such as relational tables).
[0179] 2. In the query planning method, query execution method, materialized view selection method, and method of incrementally updating the query plan based on the selected materialized view, the present invention can also be achieved by adopting any other bottom-up node order (i.e., processing all child nodes of a node before processing the node itself).
[0180] 3. In the query execution method, if materialization |[[∈]] is used... G The invention can also be achieved by means of marking, rather than by using a marking method.
[0181] 4. In the cost and cardinality estimation method, if the formulas are modified or a machine learning model is used, the present invention can still be achieved as long as the input and output remain unchanged.
[0182] Regarding specific application scenarios of this invention:
[0183] In a biochemical context, nodes in a graph represent proteins, edges represent interactions between proteins, and edge labels indicate different types of interactions. Researchers in biochemistry need to query which proteins exhibit certain types of biological pathways. These biological pathways are paths in the graph that have specific tagged sequences. Therefore, queries for biological pathways can be expressed using regular expression path queries. Based on the experience of experts in biochemistry, a set of common regular expression path queries can be constructed. Applying this method to this set can select several physical views, accelerating subsequent queries from researchers.
[0184] Another embodiment of the present invention provides a regular path query system based on materialized view selection and query planning, comprising:
[0185] A query planner that generates multiple query plans for a given regular path query load and a directed graph with edge labels.
[0186] A materialized view selector is used to select materialized views to minimize the total query cost of the query load, and to detect and remove redundant views using the multi-query plan; and to perform incremental updates to the multi-query plan during the materialized view selection process;
[0187] An executor that performs load queries based on multiple query plans using materialized views.
[0188] The specific implementation process of the query planner, materialized view selector, and executor is described in the preceding description of the method of this invention.
[0189] Another embodiment of the present invention provides a computer device (computer, server, smartphone, etc.) including a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the steps of the method of the present invention.
[0190] Another embodiment of the present invention provides a computer-readable storage medium (such as ROM / RAM, disk, optical disk) storing a computer program that, when executed by a computer, implements the various steps of the method of the present invention.
[0191] The specific embodiments of the present invention disclosed above are intended to help understand the content of the present invention and to implement it accordingly. Those skilled in the art will understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention. The present invention should not be limited to the content disclosed in the embodiments of this specification; the scope of protection of the present invention is defined by the claims.
Claims
1. A regular expression path query method based on materialized view selection and query planning, characterized in that, Includes the following steps: Given a regular path query load and a directed graph with edge labels, generate a multi-query plan for the query load. Materialized views are selected to minimize the total query cost of the query load, and redundant views are detected and removed using the multi-query plan. During the materialized view selection process, incremental updates are performed on multiple query plans; Based on multiple query plans, perform load queries using materialized views; The multi-query plan for generating query load takes as input: a directed acyclic graph (AOD) with closure and AND-OR; and a regularized path query load. Directed graph with labels on the sides ; Maps regular path queries to the cost function cost; maps regular path queries to the cardinality function card; the output is: a labeled AND-OR directed acyclic graph aod with closure, where the label contains two parts: the cost and cardinality estimate of each node; each OR node with multiple child nodes needs to label its target child, i.e., the child AND node selected during execution. The selection of the materialized view includes: The first step is to sort all candidate materialized views in descending order of their frequency of occurrence in the query load and initialize the total memory usage to 0. The second step is to apply the following steps to each of the sorted candidate materialized views: 1) If the current candidate materialized view has been used 0 times, or if the current total memory usage plus the card value of the current candidate materialized view is greater than the total memory usage limit. ,jump over; 2) Using the AND-OR directed acyclic graph (AOD) with closure and the current candidate materialized view as input, call the method to incrementally update the query plan based on the selected materialized view. Obtain the amount by which the current candidate materialized view is expected to reduce the total execution cost of the query load and the amount of change in the query plan. If the amount of reduction in the total execution cost is 0 and the card value of the current candidate materialized view is non-zero, skip it. 3) Using the changes in the current candidate materialized view and the query plan as input, call the method to apply incremental updates; add the current candidate materialized view to the query plan. Add the card value of the current candidate materialized view to the total memory usage; 4) Inspection The number of times each materialized view is used; if it is 0, then from... Delete the materialized view and subtract its card value from the total memory usage.
2. The method according to claim 1, characterized in that, The labeled directed graph is a closure-based AND-OR directed acyclic graph; the closure-based AND-OR directed acyclic graph is obtained by adding support for Kleene closure to the AND-OR directed acyclic graph, and also contains AND nodes and OR nodes, and satisfies the following conditions: Each AND node is marked by a regular expression path query operator; Each OR node is marked by a distinct regular expression path query, determined by the load. The node marked by the regular expression path query in the code is the root node; Each AND node's child nodes are OR nodes, which are marked by the operands of the regular expression path query operator that marked it. Each OR node's child nodes are AND nodes marked by the lowest priority operator in the regular expression path query associated with it.
3. The method according to claim 1, characterized in that, The incremental update of multiple query plans includes: Initialize the total execution cost reduction to 0 and the query plan change to empty; take the AOD, the current node, the current node's card value, the total execution cost reduction, and the query plan change as inputs, and call the single-node incremental update method; The single-node incremental update method takes as input a closure-based AND-OR directed acyclic graph (AOD), a node in the AOD, the node's new execution cost, the total execution cost reduction, and the change in the query plan; and outputs the updated total execution cost reduction and the change in the query plan. The steps of the single-node incremental update method are as follows: First, if the new execution cost of the input node is greater than or equal to the current cost value of the node, return; The second step is to add the following to the current total execution cost reduction: the difference between the node's current cost value and the new execution cost of the input node multiplied by the frequency of the node in the load if the regular path query represented by the current node is in the load. The third step is to assign the current node's cost value as the new execution cost of the input node; The fourth step is to call the single-node incremental update method on each parent node of the current node. The input remains unchanged except for the new execution cost of the node. The new execution cost of the node varies depending on the node type. a) If the current node is an OR node: the new execution cost of the node remains unchanged; b) If the current node is an AND node: the new execution cost of the node is set to the cost value estimated based on the operator corresponding to the parent node, the cost of all its child nodes, and the card value.
4. The method according to claim 1 or 3, characterized in that, The cost and card values of nodes in a directed acyclic graph with closure are estimated using the following method: 1) If the current node is an OR node: a) If the current node represents a single edge label: its cost and card values are the exact number of edges with this label in the graph, which can be obtained directly from the graph data; b) Otherwise, the cost and card values of the current node are the same as the cost and card values of its target child node; 2) If the current node is an AND node: a) If the current node represents / : Cost value: ,in for The last side label, for target node set and An estimate of the size of the intersection of the starting node sets. This indicates the calculation of regular expression path queries. The time cost of the result This indicates the calculation of regular expression path queries. The time cost of the result Indicates a regular expression path query As a result, Indicates a regular expression path query The result; Card value: ,in Indicates a regular expression path query As a result, Indicates the target node. Indicates the starting node; The number of distinct starting nodes in the results: ; The number of distinct target nodes in the results: ; b) If the current node represents Or +: Cost value: ,in for The last side label, , The value of is determined in the following manner: hour, and ; At that time, take ; Card value: ; The number of distinct starting nodes in the results: ; The number of distinct target nodes in the results: ; c) If the current node represents |: Cost value: The sum of the cost values of all child nodes; Card value: The sum of the card values of all child nodes; The number of distinct starting nodes in the result: the sum of the number of distinct starting nodes in all child node results; The number of distinct target nodes in the result: the sum of the number of distinct target nodes in all child node results; d) If the current node represents? Cost value: The cost value of the child node; Card value: The card value of the child node; The number of distinct starting nodes in the result: The number of distinct starting nodes in the child node result; The number of distinct target nodes in the result: The number of distinct target nodes in the child node result.
5. The method according to claim 1, characterized in that, The step of performing load queries based on multiple query plans using materialized views includes: Load Query Call the single-node query execution method; The input to the single-node query execution method is a node in a directed acyclic graph with closure (AND-OR), a starting node candidate set lCand, and a target node candidate set rCand; the output is the query result for that node; the steps of the single-node query execution method are as follows: 1) If the current node is an OR node: If the query corresponding to the current node is a single edge label or has been materialized, then directly obtain its complete query result from the graph data or materialized view, and then filter it with lCand and rCand to get the final result; otherwise, call the single node query execution method on its target child, keep the starting and target node candidate sets unchanged, and return the returned result as the final result. 2) If the current node is an AND node, the execution method is determined based on the regular expression path query operator represented by the current node: a) If the current node represents / : determine whether to execute the left child node or the right child node first based on the execution order selected in the query planning method; b) If the current node represents |: call the single-node query execution method for each child node, keeping the starting and target node candidate sets unchanged, and finally perform a union operation on the results of each child node to obtain the final result and return it; c) If the current node represents ?: Call the single-node query execution method on the child nodes and add a clause containing ? to the result. The flag is returned; d) If the current node represents + or After calling the single-node query execution method on a child node, perform fixed-point iteration until no new results are generated. Before returning the result, if the current node represents... Additional inclusion The mark.
6. A regular expression path query system based on materialized view selection and query planning, characterized in that, The system comprising performing the method of any one of claims 1 to 5, wherein the system includes: A query planner that generates multiple query plans for a given regular path query load and a directed graph with edge labels. A materialized view selector is used to select materialized views to minimize the total query cost of the query load, and to detect and remove redundant views using the multi-query plan; and to perform incremental updates to the multi-query plan during the materialized view selection process; An executor that performs load queries based on multiple query plans using materialized views.
7. A computer device, characterized in that, It includes a memory and a processor, the memory storing a computer program configured to be executed by the processor, the computer program including instructions for performing the method of any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a computer, implements the method according to any one of claims 1 to 5.
Citation Information
Patent Citations
View materialization method for large-scale knowledge graph complex path query
CN106779150A
Database materialized view construction system and method and system creation method
CN111597209A