Graph query cardinal number estimation method based on path statistical information

By building a summary graph to parallelly process path query sets, decomposing and iteratively estimating user queries, the problems of cardinality estimation accuracy and latency in graph databases are solved, and the execution efficiency and system performance of graph queries are improved.

CN120763205APending Publication Date: 2025-10-10SHANGHAI JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510951559.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-10
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing graph databases have poor cardinality estimation accuracy and high statistical information collection overhead in graph query optimization, and cannot guarantee the quality of graph query execution.

Method used

Through a method based on path statistics, a summary graph (PSG) is constructed in the offline phase, a set of path queries is selected in parallel, and user queries are decomposed and the cardinality is estimated iteratively in the online phase to obtain the upper bound of the query cardinality.

Benefits of technology

It achieves high estimation accuracy and low estimation latency on complex graph queries, ensures the optimization of query execution, and improves the performance and scalability of the graph database system.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763205A_ABST
    Figure CN120763205A_ABST
Patent Text Reader

Abstract

A graph query cardinal number estimation method based on path statistical information comprises the steps that after a path query set is selected according to an input data graph in the off-line stage, a summary graph (PSG) is constructed in parallel; the method comprises the following steps of: decomposing (Dcmp) a user query received in real time according to PSG information in an online stage, and performing iterative cardinality estimation according to the decomposed query to obtain a cardinality upper bound of the user query; by means of the method, the average lowest estimation delay can be achieved on complex graph query (especially ring query), it is ensured that high estimation precision can be achieved when complex graph query is processed, and meanwhile it is ensured that an estimation value is the upper bound of a true value.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of information processing, and particularly relates to a graph query cardinality estimation method based on path statistical information. BACKGROUND

[0002] In recent years, with the growth of graph data processing demand and the maturation of graph processing related technologies, graph databases have gradually become a new force for rapid development in the field of databases. The two main functions of graph databases are to store graph data and execute graph queries. The prior art optimizes graph queries through a cost-based query optimizer (CBO) to find the optimal graph query execution mode, but the existing optimization technology has problems such as poor cardinality estimation accuracy and large statistical information collection overhead, and cannot guarantee the quality of the generated graph query execution mode. SUMMARY

[0003] The present application proposes a graph query cardinality estimation method based on path statistical information, which can achieve the lowest average estimation delay on complex graph queries (especially cyclic queries), ensure higher estimation accuracy when processing complex graph queries, and guarantee that the estimated value is an upper bound of the true value.

[0004] The present application is implemented by the following technical solutions:

[0005] The present application relates to a graph query cardinality estimation method based on path statistical information. In the offline stage, a path query set is selected according to the input data graph, and a summary graph (PSG) is constructed in parallel. In the online stage, after the user query received in real time is decomposed (Dcmp) according to the PSG information, iterative cardinality estimation is performed according to the decomposed query to obtain the upper bound of the cardinality of the user query.

[0006] The path query set is selected according to the graph schema (Graph Schema) or query requirements, and preferably all path queries with a length not exceeding a specified parameter are selected as the path query set.

[0007] In the summary graph, each vertex represents a vertex partition of the original graph obtained by a vertex partition algorithm, and the vertices are connected by undirected summary edges.

[0008] The present invention relates to a system for implementing the above-mentioned method, comprising: a parallel PSG builder, a PSG manager, and an iterative cardinality estimator based on PSG, wherein: the parallel PSG builder constructs a summary graph that meets the conditions in parallel based on a data graph G and a path query set P; the PSG manager receives and stores the generated summary graph and defines a query interface for the iterative cardinality estimator to query the summary graph; the iterative cardinality estimator decomposes the query based on a user query Q and the summary graph provided by the PSG manager, and then iteratively estimates the decomposed query to obtain the upper bound of the cardinality of Q.

[0009] Technical Effects

[0010] This invention uses a path-centric summary graph to intuitively describe information such as the cardinality and maximum degree of a path. It then decomposes the query and iteratively estimates the cardinality using a query decomposition and iterative estimation algorithm during cardinality estimation, ultimately obtaining an upper bound on the cardinality. A parallel PSG construction algorithm ensures efficient PSG construction on large-scale graph data. Compared to existing similar technologies, this invention achieves the highest average estimation accuracy and lowest average estimation latency for complex cyclic and acyclic graph queries. The CBO based on this invention can find the optimal query execution method and efficiently construct PSGs on large-scale graph data, with good scalability relative to the graph size and number of worker threads. BRIEF DESCRIPTION OF THE DRAWINGS

[0011] Figure 1 Flowchart of the present invention;

[0012] Figure 2 This is a schematic diagram of the effect of query decomposition;

[0013] In the figure: (a) represents the input graph query Q, (b) represents the new query Q' obtained after decomposition, where each edge of Q' represents a path in Q;

[0014] Figure 3 Schematic diagram of data graph, graph query, path query and PSG structure;

[0015] In the figure: (a) represents the data graph G, (b) represents the graph query Q and the path queries P0, P1, P2, P3, (c) represents the PSG constructed using the data graph G and the path query set P = {P0, P1, P2, P3}, and (d) represents the attribute value of the edge in the PSG;

[0016] Figure 4 Decompose the flow chart for the query;

[0017] Figure 5 Construct a flow chart for PSG;

[0018] Figure 6 It is an iterative estimation flow chart;

[0019] Figure 7 Schematic diagram of the effect of an embodiment of iterative estimation;

[0020] Figure 8 and Figure 9 Schematic diagram of the embodiment effect. DETAILED DESCRIPTION

[0021] like Figure 1 As shown, this embodiment relates to a graph query cardinality estimation method based on path statistical information, which includes an offline preprocessing stage and an online query stage.

[0022] The offline preprocessing stage specifically includes:

[0023] Step 1: Figure 1 As shown, all path queries whose length does not exceed the specified parameter K are selected from the input data graph G as the path query set P.

[0024] The specified parameter K refers to the maximum length of the path query.

[0025] Step 2: Use the data graph G and the path query set P to build Figure 3 (d) shows the PSG, where each vertex represents a vertex partition of the original graph obtained by the vertex partitioning algorithm, and the vertices are connected by undirected summary edges.

[0026] The vertex has label information, which is the same as the label of the corresponding vertex partition.

[0027] The summary edge refers to the connection between vertices V1 and V2, which indicates that there is a path connecting the vertex partitions V1 and V2 of the original graph that meets a certain path query p. The summary edge carries attribute information, specifically in the form of a triple (c, d1, d2), where c represents the cardinality of the path between partitions V1 and V2 that meets the path query p, and d1 and d2 represent the maximum degrees of the two ends of the edge, that is, the starting point and end point of the path, respectively.

[0028] like Figure 5 As shown, the parallel construction specifically includes:

[0029] 2.1 Input Data Graph , path query set and vertex partition map ,right Each point The initial length is Vector , specifically: , where: the path query satisfies the length not exceeding the aforementioned parameters ; Vertex partition map is The mapping of vertices in to vertex partitions in PSG, is the number of vertex partitions in PSG.

[0030] 2.2 Start with a path query of length 1 and process it according to the length Path query, until the length , with data graph The points in the middle are the granularity, parallel computing The head vector and tail vector of the path query in In the data graph The vectors corresponding to all matching start and end points in the , and the path query is calculated using the head and tail vectors The corresponding summary edge triplet.

[0031] The path query The head vector is the subquery obtained by removing an edge connected to the starting point The head vector is calculated, if The length is ,but Length is ,deal with Guaranteed The head vector of has been calculated.

[0032] The path query The tail vector of the subquery is obtained by removing an edge connected to the end point The tail vector of is calculated, if The length is ,but Length is ,deal with Guaranteed The tail vector of has been calculated.

[0033] 2.3 Using Path Query Collection The PSG is constructed by summarizing the edge triplets corresponding to all paths in .

[0034] The online query stage specifically includes:

[0035] Step 1: Figure 4 As shown, according to the PSG information in the PSG manager, the user query Q received in real time is decomposed (Dcmp) to obtain the decomposed query, which specifically includes:

[0036] 1.1 Find all pivot points in the user query Q, that is, points with degree greater than or equal to 3. If such a point does not exist in the user query Q, then take any point with the smallest degree as the pivot point.

[0037] 1.2 Perform a depth-first search (DFS) starting from each hub point obtained in step 1.1. When any point u in the user query Q is found and satisfies: point u is a hub point or u has no unvisited neighbors, DFS terminates, thereby obtaining several paths starting from each hub point and storing these paths in the candidate path set (PSet).

[0038] 1.3 Decompose each path in PSet using dynamic programming strategy , to transfer the state, specifically: ,in: Indicates the path Middle Point to The subpath consists of points, Decomposition subpath The minimum number of paths required, is the path query set in PSG.

[0039] Step 2: Figure 6 As shown, the decomposed query is iteratively estimated to obtain the upper bound of the cardinality of the original query Q, including:

[0040] 2.1 Generate an estimated order S using the decomposed query Q, where: the estimated order is the sequence of points in the query Q;

[0041] 2.2 Process the points in S in order, for each point :

[0042] 2.2.1 Extraction Central subquery ,in: Include and in All neighbor points in ;

[0043] 2.2.2 Estimation Boundary point, that is, the aforementioned The default number of neighbor points is no more than 2 corresponding summary edge triplets;

[0044] 2.2.3 Update Query :Will Replace with The edges formed by connecting boundary points;

[0045] 2.3 Repeat step 2.2 until only one point remains unprocessed in S;

[0046] 2.4 Output the final cardinality upper bound.

[0047] like Figure 8 and Figure 9The experimental results show that the application (PathCE in the figure) has obvious precision advantage over the benchmark estimators (GLogS, CEG, etc.) on multiple graph queries (the vertical axis in the figure represents the estimation error, and the lower the error, the higher the precision).

[0048] Meanwhile, compared with the prior art (CEG, GLogS, FactorJoin, etc.), the application achieves the best estimation precision and the average lowest estimation delay on complex graph queries (especially on loop queries). The test in the actual graph database system also shows that the cardinality estimation technology of the application can improve the execution performance of complex graph queries. In addition, the parallel PSG construction algorithm proposed by the application also has advantages in the performance of statistical information collection.

[0049] The above specific embodiments can be adjusted in different ways by those skilled in the art without departing from the principles and purposes of the application, the protection scope of the application is subject to the claims and is not limited by the above specific embodiments, and each implementation scheme within the scope is subject to the constraints of the application.

Claims

1. A graph query cardinality estimation method based on path statistics, characterized in that: In the offline phase, a path query set is selected based on the input data graph, and a summary graph (PSG) is constructed in parallel. In the online phase, the real-time user queries are decomposed (Dcmp) based on the PSG information. The decomposed queries are then iteratively estimated for cardinality to obtain the upper bound of the cardinality of the user queries. The path query set is selected based on the graph pattern or query requirements; In the summary graph, each vertex represents a vertex partition of the original graph obtained by a vertex partitioning algorithm, and the vertices are connected by undirected summary edges.

2. The graph query cardinality estimation method based on path statistical information according to claim 1 is characterized in that: The path query set selects all path queries whose length does not exceed the specified parameter as the path query set.

3. The graph query cardinality estimation method based on path statistical information according to claim 1 is characterized in that: The offline preprocessing stage specifically includes: Step 1: In the input data graph G, all path queries whose length does not exceed a specified parameter K, i.e., the maximum length of a path query, are selected as a path query set P; Step 2: Use the data graph G and the path query set P to construct PSG in parallel, where each vertex represents a vertex partition of the original graph obtained by the vertex partitioning algorithm, and the vertices are connected by undirected summary edges; The vertex has label information that is the same as the label of the corresponding vertex partition; The summary edge means that the edge connecting vertices V1 and V2 represents that there is a path connecting the vertex partitions V1 and V2 of the original graph that meets a certain path query p. The summary edge carries attribute information, specifically in the form of a triple (c, d1, d2), where: c represents the cardinality of the path between partitions V1 and V2 that meets the path query p, and d1 and d2 represent the maximum degrees of the two ends of the edge, that is, the starting point and end point of the path, respectively.

4. The graph query cardinality estimation method based on path statistical information according to claim 1 or 3, characterized in that: The parallel construction specifically includes: 2.1 Input Data Graph , path query set and vertex partition map ,right Each point The initial length is Vector , specifically: , where: the path query satisfies the length not exceeding the aforementioned parameters ; Vertex partition map is The mapping of vertices in to vertex partitions in PSG, is the number of vertex partitions in PSG; 2.2 Start with a path query of length 1 and process it according to the length Path query, until the length , with data graph The points in the middle are the granularity, parallel computing The head vector and tail vector of the path query in In the data graph The vectors corresponding to all matching start and end points in the , and the path query is calculated using the head and tail vectors the corresponding summary edge triples; 2.3 Using Path Query Collection The PSG is constructed by summarizing the edge triplets corresponding to all paths in .

5. The graph query cardinality estimation method based on path statistical information according to claim 4 is characterized in that: The path query The head vector is the subquery obtained by removing an edge connected to the starting point The head vector is calculated, if The length is ,but Length is ,deal with Guaranteed The head vector of has been calculated; The path query The tail vector of the subquery is obtained by removing an edge connected to the end point The tail vector of is calculated, if The length is ,but Length is ,deal with Guaranteed The tail vector of has been calculated.

6. The graph query cardinality estimation method based on path statistical information according to claim 1 is characterized in that: The online query stage specifically includes: Step 1: Decompose (Dcmp) the user query Q received in real time according to the PSG information in the PSG manager to obtain the decomposed query; Step 2: Perform iterative cardinality estimation on the decomposed query to obtain the upper bound of the cardinality of the original query Q, including: 2.1 Generate an estimated order S using the decomposed query Q, where: the estimated order is the sequence of points in the query Q; 2.2 Process the points in S in order, for each point : 2.2.1 Extraction Central subquery ,in: Include and in All neighbor points in ; 2.2.2 Estimation Boundary point, that is, the aforementioned The default number of neighbor points is no more than 2 corresponding summary edge triplets; 2.2.3 Update Query :Will Replace with The edges formed by connecting boundary points; 2.3 Repeat step 2.2 until only one point remains unprocessed in S; 2.4 Output the final cardinality upper bound.

7. The graph query cardinality estimation method based on path statistical information according to claim 1 or 6, characterized in that: The decomposition specifically includes: 1.1 Find all pivot points in the user query Q, that is, points with degree greater than or equal to 3. If there is no such point in the user query Q, then randomly select the point with the smallest degree as the pivot point; 1.2 Perform a depth-first search starting from each pivot point obtained in step 1.

1. DFS terminates when any point u in the user query Q is found and satisfies: point u is a pivot point or u has no unvisited neighbors. This results in several paths starting from each pivot point and stores these paths in the candidate path set (PSet). 1.3 Decompose each path in PSet using dynamic programming strategy , to transfer the state, specifically: ,in: Indicates the path Middle point to The subpath consists of points, Decomposition subpath The minimum number of paths required, is the path query set in PSG.

8. A graph query cardinality estimation system based on path statistics that implements the method according to any one of claims 1 to 7, characterized in that: include: A parallel PSG builder, a PSG manager, and an iterative cardinality estimator based on PSG, wherein: the parallel PSG builder constructs a qualified summary graph in parallel based on the data graph G and the path query set P; the PSG manager receives and stores the generated summary graph and defines a query interface for the iterative cardinality estimator to query the summary graph; the iterative cardinality estimator decomposes the query based on the user query Q and the summary graph provided by the PSG manager, and then iteratively estimates the decomposed query to obtain the upper bound of the cardinality of Q.