A Complex Graph Query Optimization Method Based on Vertical Partitioning and Pre-Connection of Summary Graphs
Through the method of vertical division and pre-connection of abstract graphs, RDF graph query is optimized, and the problems of excessive connection operations and error introduction in the prior art are solved, and efficient RDF graph data query is realized.
Patent Information
- Application Number
- CN202210350911.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-02
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-04-02
AI Technical Summary
The prior art has too many connection operations in RDF graph data query, and errors are introduced in the graph summary, resulting in inefficient query.
Using a method based on the vertical division and pre-joining of the summary graph, the initial RDF triple data is processed through the hash table, the hash table HTS, binary table BT and ternary table TT are generated, and the SPARQL query tuple QT is stored on HDFS, and the query results are optimized using Spark SQL query.
Reduces the number of connection operations during query, eliminates the results of not meeting connection conditions, and improves query performance, especially in large-diameter queries and large data volumes, which outperform existing methods.
Smart Images

Figure CN114706883B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of RDF graph query, and particularly relates to a complex graph query optimization method based on vertical partitioning and pre-joining of summary graphs. Background Art
[0002] With the continuous development of information technology, large RDF graph data in various fields have emerged continuously, such as Google Knowledge Graph, DBPedia, etc. The rapid growth of the data scale of large graph data has brought huge challenges to efficient graph query based on large graph data. RDF (Resource Description Framework) is a standard proposed by the W3C for semantic data modeling, which can be used to describe the characteristics of Web resources and the relationships between resources. It has a very flexible graph data model structure, so it is usually used to represent various types of highly loose and structured data. SPARQL (SPARQL Protocol and RDF Query Language) is a query language and data acquisition protocol developed for RDF. It is defined for the RDF data model developed by the W3C and can be used for querying any information resource that can be represented by RDF.
[0003] Currently, the existing systems for querying and processing RDF data are mainly divided into three categories: single-machine, centralized, and distributed. In the past few years, systems such as RDF-3X, SW-Store, Jena, Sesame, and Hexastore have been introduced. These single-machine systems all face performance bottlenecks in loading and querying large-scale RDF data. Some existing centralized RDF systems include Virtuoso, 3store, Rstar, DB2RDF, gStore, chameleon db, TripleBit, and BitMat. The main advantage of these centralized systems is that they do not generate communication overhead, but their performance is limited by the memory capacity and computing power of a single machine. As the data becomes larger and larger, the amount of data represented by RDF also becomes larger and larger, and data volumes in the billions are becoming more and more common. Therefore, efficient processing of large graph data has become increasingly important. Therefore, in order to cope with the challenges of big data, a scalable distributed method is required. Compared with centralized RDF systems, these distributed systems have a higher aggregated memory size and processing capacity.
[0004] Existing RAM-based distributed systems usually have good query performance, such as TriAD and AdPart. However, they have huge memory overhead because they store the entire RDF graph in memory. Although the query performance of storage-partition-based distributed systems is inferior to that of RAM-based systems, the requested resources are very low and they are more scalable compared to RAM-based distributed systems. However, some existing storage-partition-based systems have been optimized for highly selective star queries and can efficiently respond to star pattern queries, such as PRoST and SPT+PT, by using property tables to respond to star queries in order to achieve fast response to star queries. However, SPARQL queries are join-intensive queries. Research shows that users tend to request increasingly complex interactive queries. In the SPARQL query logs of interactive DBpedia, some queries can have up to 10 joins; analytical queries in the biomedical field even include 50 or more joins. For SPARQL queries, n query tuples need to perform n-1 join operations. The join operations will generate intermediate result information, and these intermediate results usually affect the performance of subsequent join operations. The larger the amount of generated intermediate result data, the greater not only the subsequent join overhead but also the communication overhead between cluster nodes. Therefore, how to reduce the size of the generated intermediate results, thereby minimizing the amount of data to be shuffled on the network and reducing the computational overhead of join operations to reduce CPU and memory usage has become an issue that needs to be considered in current research.
[0005] In RDF graph data, most of the data has no relation to the query results, and these data can be pruned by some means to optimize the query phase. Some systems have proposed some strategies to prune RDF data. GoFast proposed a pruning strategy to eliminate the use of fragments that do not participate in constructing the final query result upstream, and refined the selected plan by pruning invalid clusters that do not participate in the construction of the final query result. Jiang et al. proposed a general ontology-based keyword search index framework, using keyword search technology to prune semantically irrelevant subgraphs. Peng et al. utilized the interconnection topology between RDF sources to filter out irrelevant sources. Leon+ performed early pruning on RDF data through RDF summary technology and path partitioning strategy. According to analysis, the structures of most nodes in the RDF graph are similar. For example, a node of the type student usually has features such as "guided by whom" and "taking what courses". During query evaluation, when facing SPARQL queries under similar structures, they usually return as query results simultaneously, while nodes with other structures are usually irrelevant to the results. Therefore, when querying one of the results, it is hoped to quickly locate those nodes with similar structures and relationships to this result, and discard those nodes that are completely irrelevant to the query structure, which requires associating these nodes with similar structures and relationships together. Therefore, an efficient RDF graph query method based on vertical partitioning and pre-connection of the summary graph is proposed. Summary of the Invention
[0006] Aiming at the deficiencies in the prior art, the purpose of the present invention is to solve the technical problems of excessive join operations required during query in the prior art and the error existing in the introduction of graph summary.
[0007] To achieve the above purpose, the technical solution of the present invention is as follows:
[0008] A complex graph query optimization method based on vertical partitioning and pre-connection of the summary graph, comprising the following steps:
[0009] Step S1) Process the initial RDF triple data to obtain the hash table HTS, binary table BT, ternary table TT, and BT statistics and TT statistics after graph summarization. HTS, BT, and TT are all stored on HDFS;
[0010] Step S2) Optimize the SPARQL query tuple QT based on the binding number of QT and BT statistics to obtain an initial optimization sequence, and determine whether it is necessary to further optimize the obtained optimization sequence according to TT statistics based on the number of query tuples included in QT;
[0011] Step S3): Read the corresponding table from HDFS, execute the Spark SQL query based on the summary, and obtain the query result in the summary graph; perform data restoration based on the hash table HTS, and then perform the Spark SQL query to obtain the final query result.
[0012] Further, the specific steps of the said Step S1) are as follows:
[0013] Step S110): Process the initial RDF triples based on the aggregated graph summary to obtain the hash table HTS storing all summary information; where HTS represents the hash table HT summary, that is, the summary structure indexed by the hash table HT.
[0014] Step S120): Based on the vertical partitioning of the summary, partition the summary edges stored in HTS according to the predicates to obtain the binary table BT and the corresponding BT statistics.
[0015] Step S130): Calculate the pre-join of the binary table BT to obtain the TT and the corresponding TT statistics.
[0016] Step S140): HTS, BT, and TT are all stored on HDFS.
[0017] Further, the specific steps of the said Step S2) are as follows:
[0018] Step S210): Optimize the SPARQL query tuple QT based on the binding quantity of QT and the BT statistics to obtain an initial optimized sequence.
[0019] Step S220): Combine the QTs according to the possible connection methods between the QTs, calculate the corresponding priority weights of the combinations, and select an optimal combination for each tuple, that is, the tuple in the combination has not been selected and obtains the maximum weight in all combinations; if QT only contains a single tuple, no join operation is required during query execution, so there is no need to optimize QT according to the TT statistics and directly enter Step S3), otherwise enter Step S230).
[0020] Step S230): Based on the pre-computed tuple join results of TT, further optimize the obtained optimized sequence according to the TT statistics.
[0021] Further, the specific steps of the said Step S3) are as follows:
[0022] Step S310): Read BT from HDFS, execute the Spark SQL query based on BT for the optimized sequence not optimized according to the TT statistics, and obtain the query result in the summary graph.
[0023] Step S320) Read TT from HDFS, execute a SparkSQL query based on TT for the optimized sequence obtained from the TT statistics, and obtain the query result in the summary graph;
[0024] Step S330) Restore the summary data in the result queried in the summary graph to the data in the initial RDF graph based on the hash table HTS;
[0025] Step S340) Finally, execute a Spark SQL query on the data after summary restoration to obtain the final query result.
[0026] Furthermore, the specific steps of the said Step S110) are as follows:
[0027] Step S111) The initial RDF graph is used as a summary graph where all super nodes only contain one original graph node, and it is stored in a hash table, denoted as HTG;
[0028] Step S112) Calculate the edge merging error for any two super nodes;
[0029] Step S113) Only merge two super nodes when the edge merging error between them is less than the error critical value.
[0030] Furthermore, the vertical partitioning divides the triple table into different sub-tables according to the predicate p, and each sub-table only stores two columns in the triple, so as to avoid duplicate storage of predicates; each predicate in the triple table uniquely corresponds to a binary table, which has the function of an index. When querying, only connect some binary tables related to the query.
[0031] Furthermore, the specific steps of optimizing the query sequence based on the binding quantity of QT and BT statistics are as follows:
[0032] Step S211) First, obtain the binding values of all query tuples themselves and the statistics in the corresponding binary tables, and calculate the priority weights;
[0033] Step S212) Subsequently, obtain the sorted query sequence according to the weight sorting;
[0034] Step S213) Find the neighbor tuples of the query tuples and calculate the corresponding weights;
[0035] Step S214) On the basis of avoiding duplicate selection of tuples, select a neighbor tuple with the largest weight after combination for each query tuple.
[0036] Beneficial effects:
[0037] 1. The present invention can be used to implement large-scale RDF data query processing. The proposed lossless graph summarization method is used to accelerate queries without losing the original graph information. To further accelerate queries, the summary graph is vertically partitioned into binary tables and pre-joined. By pre-computing the join results under possible join methods, on the one hand, the number of join operations during query can be reduced, and on the other hand, the results that do not meet the join conditions can be eliminated. Vertical partitioning and pre-joining generate corresponding statistics for optimizing SPARQL queries to obtain a query sequence with the least time consumption. The optimized sequence is translated into the corresponding Spark SQL language and handed over to Spark for query execution.
[0038] 2. The method proposed by the present invention demonstrates excellent query performance, and its performance in all query modes is better than that of the semi-join-based distributed query processor S2RDF. Compared with most existing distributed RDF graph query methods, the method proposed by the present invention does not depend on the query diameter, shows no weaker performance in small-diameter queries, and still shows a performance one order of magnitude better than other methods in large-diameter queries. To obtain the shortest average query time, the average query time of the system under different summary sizes is tested. When the number of supernodes in the summary graph is 20% of the number of original graph nodes, the shortest average query time is obtained and the overall query performance reaches the optimal. Since the query results in the summary graph are approximately the query results in the original graph, for cases where only the existence of query results or only approximate query results are concerned, the ratio of the number of supernodes in the summary graph to the number of original graph nodes can be set smaller. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 Flowchart of a complex graph query optimization method based on vertical partitioning and pre-joining of a summary graph provided by an embodiment of the present invention;
[0040] Figure 2 RDF set and its corresponding RDF graph provided by an embodiment of the present invention;
[0041] Figure 3 Storage structure diagram of the summary graph provided by an embodiment of the present invention;
[0042] Figure 4 Summary graph corresponding to the RDF graph provided by an embodiment of the present invention;
[0043] Figure 5 Flowchart of the processing of hyperedges when merging supernodes provided by an embodiment of the present invention;
[0044] Figure 6 SPARQL and its query graph provided by an embodiment of the present invention;
[0045] Figure 7Flowchart for optimizing SPARQL queries and converting them into Spark SQL provided by embodiments of the present invention;
[0046] Figure 8 Flowchart of data changes during the execution of Spark SQL queries provided by embodiments of the present invention;
[0047] Figure 9 Comparison chart of average query times under different summary sizes provided by embodiments of the present invention;
[0048] Figure 10 Pseudocode diagram of the hash table algorithm provided by embodiments of the present invention;
[0049] Figure 11 Pseudocode diagram of the graph summary algorithm provided by embodiments of the present invention;
[0050] Figure 12 Pseudocode diagram of the vertical partitioning algorithm based on summaries provided by embodiments of the present invention;
[0051] Figure 13 Pseudocode diagram of the pre-join algorithm based on summaries provided by embodiments of the present invention;
[0052] Figure 14 Pseudocode diagram of the SPARQL query optimization algorithm provided by embodiments of the present invention;
[0053] Figure 15 Pseudocode diagram of the SPARQL2Spark SQL algorithm provided by embodiments of the present invention;
[0054] Figure 16 Pseudocode diagram of the query execution algorithm provided by embodiments of the present invention. Detailed implementation manners
[0055] The present invention will be described below with reference to specific embodiments. Those skilled in the art can understand that these embodiments are only used to illustrate the present invention and do not limit the scope of the present invention in any way.
[0056] As Figure 1 shown, a complex graph query optimization method based on vertical partitioning and pre-joining of summary graphs includes the following steps:
[0057] Step S1) Process the initial RDF triple data to obtain the hash table HTS, binary table BT, triple table TT, and BT statistics and TT statistics after graph summarization. HTS, BT, and TT are all stored on HDFS.
[0058] The initial RDF graph is a directed labeled graph: Graph G = (V G , E G , P G , φG ) where
[0059] (1)V G corresponds to the set of all subjects and objects in the RDF triples,
[0060] (2) corresponds to the set of directed edges in all RDF triples,
[0061] (3)P G is the set of all edge labels,
[0062] (4)φ G is a label mapping function, φ G : denotes the assignment of a subset of P G to the edge e ∈ E G .
[0063] As Figure 2 shown, Figure 2 a set of RDF triples enumerated in (a) in Figure 2 and (b) in
[0064] Step S110) processes the initial RDF triples based on the aggregated graph summary to obtain a hash table HTS storing all summary information.
[0065] The graph summary of the present invention aggregates the original graph nodes with similar attributes and topologies into supernodes, that is, node clusters composed of a part of the nodes in the initial RDF graph, and uses hyperedges (relationships between supernodes) to represent the full connection relationship between the nodes included in two corresponding supernodes. The graph summary forms a smaller graph by removing some detailed information in the initial graph, called the summary graph. The summary graph retains the core attributes and topological features of the initial data graph as much as possible.
[0066] As the summary graph of the graph G = (V G , E G , P G , φ G ), where
[0067] 1) is the set of supernodes and satisfies: (1): (2) (3) Let π(v) denote the supernode to which the node v ∈ V G belongs in the graph S G ;
[0068] 2) is the set of hyperedges. For any hyperedge Indicates at the super node and There is a full connection between them, that is
[0069] 3) Indicates the set of property labels on all hyper edges in graph S;
[0070] 4) φ S Indicates the set of property labels for each hyper edge and
[0071] such as Figure 3 is shown as Figure 2 a summary graph of the RDF graph in (b) of
[0072] The internal structure of the hash table is as follows:
[0073] (1) node.ID: The id of the super node,
[0074] (2) node.origNodeSet: The original graph nodes contained in the super node,
[0075] (3) node.selfConns: The edges in the original graph corresponding to the self - loop edges of the super node, that is, the edges actually existing between the internal nodes of the super node, internally stored as (s, p, o), representing an edge (s, o) with a relationship of p,
[0076] (4) node.edgeList: All hyper edges adjacent to the super node,
[0077] (5) edge.neighborID: The super node ID of the other end point of the hyper edge, which together with node.ID identifies a hyper edge,
[0078] (6) edge.interConns: The edges actually existing in the hyper edge between the super node and its neighbor, and like node.selfConn, it is also internally stored in the form of (s, p, o).
[0079] Step S111) Initialize the RDF graph G=(V G , E G , P G , φ G ) as a summary graph in which each super node contains only one original graph node, and it is stored in the hash table, denoted as HTG.
[0080] such as Figure 4As shown, the key of the hash table is the ID corresponding to the node, and the value is the relevant information of the node. The current supernode only contains one original graph node, which is stored in origNodeSet. The self-loop edges of the supernode are stored in selfConns, while the adjacent edges are stored in edgeList.
[0081] Step S112) combines two similar supernodes in a node aggregation-based manner into a supernode by ensuring that the two supernodes to be combined meet a certain similarity through error calculation.
[0082] During the summarization process, the two supernodes to be combined must have a certain similarity. Let the graph be the summary graph of graph G = (V G , E G , P G , φ G ). For any two supernodes their supernode similarity is calculated as shown in formula (1),
[0083]
[0084] where represents the set of all neighbor supernodes of supernode . Similarly,
[0085] In the summary graph , two similar supernodes are combined into a supernode In terms of node merging formally, and However, node merging will introduce edge merging errors. The edge merging errors come from two aspects: First, for the newly merged supernode ideally, if the node sets corresponding to and are fully connected, that is, Since in the real graph and the node sets corresponding to them are usually not fully connected, the edge merging error introduced thereby is called the internal merge error (IME). Second, will also introduce edge merging errors between it and its adjacent nodes, because and the node sets corresponding to them and their adjacent nodes are usually not fully connected either. This error is called the adjacent merge error (AME).
[0086] Let the figure be a summary graph of the initial RDF graph G = (V G , E G , P G , φ G ). For any two supernodes If they have been aggregated into a new supernode , then the new supernode will replace and and appear satisfying and
[0087] Internal Merging Error IME: Define the internal merging error between any two supernodes as formula (2):
[0088]
[0089] where respectively represent the full connection and the actual connection between the node sets corresponding to the supernodes and . It is set that if When there is an actual edge between only two supernodes, a false edge will be introduced to form a full connection to form a hyperedge, and the internal merging error is calculated.
[0090] Adjacent Merging Error AME: Define the adjacent merging error introduced by merging the supernode as shown in formula (3):
[0091]
[0092] Merging Error ME: Define the merging error for merging two supernodes as equation (4):
[0093]
[0094] Step S113) When the edge merging error is less than the error critical value, the supernode pair is merged.
[0095] To control the merging error of the supernode pair, preferentially merge the supernode pair with a smaller merging error, introduce a threshold called the error critical value, abbreviated as ET. Only when the merging error of two supernodes is not greater than the error critical value will they be merged.
[0096] The present invention sets the initial error threshold to 0, indicating that superpoint pairs with the same neighbors are merged first. When traversing all superpoints in the entire hash table, ET is adjusted, and the corresponding adjustment expression is shown in Formula (5):
[0097]
[0098] where |V G | and respectively represent the number of initial graph nodes and the number of superpoints in the summary graph at the current moment, and their ratio represents the number of original graph nodes contained in each superpoint on average. If ET grows too slowly, the number of superpoints meeting the merging conditions in the subsequent stage is too small; if ET grows too fast, the ability to control the merging error of superpoint pairs cannot be reflected. Therefore, the setting of the error threshold must be appropriate. In graph summarization, the introduction of merging error comes from the introduction of false edges and is related to the number of nodes inside the superpoint. To control the error threshold, the present invention sets that in the next round of loop, all nodes inside the superpoint pairs to be merged are allowed to introduce one more false edge, and the increased error is where represents taking the integer part of a real number.
[0099] For the convenience of processing, no new superpoint is created when merging superpoint pairs as the merged superpoint, but is merged into superpoint . Therefore, when merging superpoint pairs, the set of original graph nodes and adjacent edges in superpoint needs to be added to .
[0100] As Figure 10 shown in Algorithm 1 for node aggregation. First, update the edge information in all neighbor superpoints of superpoints and , as shown in lines 3 - 9 of the algorithm. In lines 3 - 6, if neighbor superpoint n is the common neighbor of and , merge the superedges and , that is, add the actual edges in superedge to superedge , where is the abbreviation of , representing the set of all actual edges in superedge , and is used to represent superedge . In lines 7 - 9, if superpoint n is the unique neighbor of , then modify the ID of the destination endpoint of superedge to At this time, the hyperedge is modified to hyperedge where represents the destination endpoint of the hyperedge . If the hypernode n is the unique neighbor of , the edge information remains unchanged, so no processing is done. After completing the edge update of the neighbor hypernodes, next update and 's adjacent edge information, as shown in lines 10 - 19 of the algorithm. In lines 10 - 16, all hyperedges in the edge set are added to the edge set. Lines 17 - 19 update the hypernode information, where is the abbreviation of , representing the set of all actual edges within the self - connecting edge on . Finally, update the hash table HTG, and delete the information of the hypernode from the hash table. The time complexity of Algorithm 1 is where represents the average degree of the graph G.
[0101] Figure 5 Describes the result after merging the hypernode into . In Figure 5 , the hyperedges between the hypernodes and and their common neighbors are merged into a larger hyperedge, and the actual edges contained in this hyperedge are the sum of the initial two hyperedges. The hyperedges between the unique neighbors of and are modified to point to , while the hyperedges between and its unique neighbors are not processed. The hyperedges between the hypernodes and are merged with the self - connecting edge on the hypernode, forming a larger self - connecting edge on .
[0102] Before merging a pair of hypernodes, it is necessary to determine which pair of hypernodes should be merged. As shown in Algorithm 2 in Figure 11 , it mainly introduces how to find the pair of hypernodes to be merged from the hash table, and then calls the nodeMerge function in Algorithm 1 to merge them. The algorithm iteratively merges hypernodes until the total number of hypernodes in the hash table is no more than k, where |HTG| represents the number of hypernodes in HTG, and k is the predefined number of hypernodes. In lines 7 - 13, the hypernode with the highest similarity to the hypernode is obtained from its two - hop neighbors, denoted as . In line 14, based on formula (4), calculate the hypernodes and The merging error. If the merging error is less than the error threshold, then call Algorithm 1 to merge the super nodes and As shown in line 15. In line 17, when the number of iterations reaches |HTG| times, then adjust the critical error value according to formula (5). When the number of super nodes in the summary graph is no greater than the preset k, then exit the loop, end the graph summarization, and return the graph summarization result HTG. The currently returned summarization result HTG corresponds to the summary graph of graph G To distinguish the hash table HTG corresponding to the initial RDF graph, hereinafter it is uniformly denoted by HTS G In Algorithm 2, the time complexity required to find a pair of the most similar super nodes and merge them is Therefore, the time complexity of the graph summarization algorithm is
[0103] Step S120) Based on the summary-based vertical partitioning, partition the summary edges stored in HTS according to the predicates to obtain the binary table BT and the corresponding BT statistics
[0104] Vertical partitioning: Partition the triple table according to the predicate p into different sub-tables, and each sub-table only stores two columns in the triple, that is, (s, o), avoiding the repeated storage of predicates. Each predicate in the triple table uniquely corresponds to a binary table and has the function of an index. When querying, only connect the binary tables related to the query
[0105] Summary graph corresponds to a triple set For graph S G According to the predicate The binary table obtained by vertical partitioning is denoted as B p , and the corresponding set expression is as shown in formula (6):
[0106]
[0107] Since graph S G is stored in HTS G , so it is necessary to obtain from HTS G , and the corresponding calculation expression is as shown in formula (7):
[0108]
[0109] Among them, is the abbreviation of , is the abbreviation of which respectively represent the hyper edges and The set of all actual edges in. Let BT denote the set of all predicate binary tables, and the calculation expression of BT is shown in formula (8):
[0110]
[0111] Abstract graph The vertical partitioning algorithm of is described in Algorithm 3 as shown in Figure 12 First, initialize a binary table for each predicate Next, in lines 4 - 13, traverse the entire hash table and add the hyperedges to the corresponding binary tables. Finally, save the binary tables and the statistics of the binary tables, as shown in lines 15 - 18 of the algorithm. In Algorithm 3, the vertical partitioning algorithm needs to traverse all the actual edges within all the hyperpoints. The time complexity of traversing all the hyperedges is And the average number of actual edges within each hyperedge is Therefore, the time complexity of the vertical partitioning algorithm is O(|E G |), where |E G | is the number of edges in the initial RDF graph.
[0112] Step S130) Calculate the pre - join of the binary table BT to obtain TT and the corresponding TT statistics.
[0113] Pre - join: The set of query tuples in SPARQL rarely contains only a single query tuple. In the vast majority of cases, it contains multiple query tuples. And when querying, n tuples require n - 1 join operations. To reduce the join operations, the pre - join based on binary tables is proposed. By pre - calculating the possible join results of query tuples, a large number of join operations during query execution are reduced, and the results that do not meet the join conditions are removed. Let the query tuple q i =(s, p i , o), q j =(s′, p j , o′) ∈ QT. If s = s′, then the query results of the query tuples q i and q j are That is, the join result at the subject of the binary tables corresponding to the predicates p i , p j . This type of join is called ubject - S ubject (SS) join, and the join result S is shown in formula (9): As shown in formula (9):
[0114]
[0115] In addition to the SS join, the possible join methods also include Subject- O bject(SO), O bject- S ubject(OS) and O bject- O bject(OO) three kinds, and the corresponding connection results are shown in the following formulas (10), (11), and (12):
[0116]
[0117]
[0118]
[0119] Lemma 2: Among all possible connection methods of the binary table, half of the connection results are duplicate results.
[0120] Proof: The predicate p j , p i ∈P G The corresponding binary table The result of the SS connection is
[0121] , it can be seen that p j , p i The SS connection result of is obtained by exchanging the order of the columns of the SS connection result of p i , p j , that is to say, the connection results of the two are the same. Similarly, when the OO connection is made, the connection result
[0122] , it can be seen that p i , p j The OO connection result of is the same as the OO connection result of p j , p i . When the predicate p j , p i ∈P G When the corresponding binary table is connected by OS,
[0123]
[0124] , so it can be known that the predicate p j , p i The SO connection of is obtained from the OS connection of the predicate p i , p j . To sum up, among the four connection methods, half of the connection operation results are obtained by other connections.
[0125] Therefore, when calculating the join result of two binary tables, only half of the SS and OO joins, and all of the OS joins need to be calculated. Let TT denote the set of triples obtained from joining all binary tables, and the corresponding set expression is shown in Equation (13):
[0126] Pre-join based on the graph summary is as shown in Algorithm 4 below. In Algorithm 4, the join results for different predicates under possible join methods are calculated. From lines 3 to 14, the join results under three join methods between predicates are obtained, and duplicate join operations are avoided. In lines 12 to 13, the join results, that is, the triples, and the statistics of the triples are saved. The reason why the statistics of the empty table are also stored in the triple statistics is that when optimizing a SPARQL query with a corresponding join, if the obtained table statistic is 0, it indicates that the result of the SPARQL query is also empty. For an RDF graph Figure 13 For each join operation, the time complexity is and it is necessary to calculate times of joins. Since Therefore, the total time complexity of Algorithm 4 is
[0127] Step S140) HTS, BT, and TT are all stored on HDFS.
[0128] Step S2) Optimize the SPARQL query tuple QT based on the binding number of QT and the BT statistics to obtain an initial optimized sequence, and determine whether it is necessary to further optimize the obtained optimized sequence according to the TT statistics based on the number of tuples included in QT.
[0129] Query graph: A SPARQL query is represented by a set of query tuples, abbreviated as QT, where the query tuple q ∈ QT is a triple in the form of (s, p, o), and s, p, o are composed of variables or constants. QT is transformed into a corresponding SPARQL query graph. As shown in Figure 1 |QT| represents the number of elements (query pattern triples) in QT, and QT is the set of query pattern triples included in a given SPARQL query.
[0130] SPARQL query graph: The query graph corresponding to SPARQL is defined as Q = (V Q , V var , E Q , P Q , φ Q ), where V Q represents the set of constant nodes included in Q, V var is the set of variable nodes of query Q, and V Q ∪ Vvar = {s, o | (s, p, o) ∈ QT}. Let Vars((s, p, o)) = {s, o} ∩ V var denote the set of variables in the query tuple q = (s, p, o). denote the set of relations of the query Q. The function φ Q : assigns a subset of P Q to each e ∈ E Q . Figure 6 is for a SPARQL query and its corresponding query graph.
[0131] Step S210) Optimize the SPARQL query tuple QT based on the number of bindings of QT and the table statistics BT to obtain an initial optimized sequence.
[0132] The table statistics of the present invention mainly target the size of the table scale, that is, the number of all records in the table. Let any table tab ∈ BT ∪ TT, and its statistics be the size of its table scale. Denote the statistical value of the table tab by SV(tab), then SV(tab) = |tab|.
[0133] Changing the join order of the tables will not change the final join result, but will obtain different intermediate results. In distributed computing, data needs to be communicated and exchanged among cluster nodes. The fewer the intermediate results generated by the join, the smaller the communication overhead. Therefore, in order to reduce the query time, the size of the generated intermediate results should be reduced as much as possible. For any two tables tab, tab′, the size of the result obtained by their join is It can be seen that generally the smaller the two joined tables are, the smaller the obtained result is. Therefore, in order to obtain smaller intermediate results, the tables with smaller table scales should be joined first as much as possible.
[0134] The number of bindings of the query tuple: Denote the number of bindings BN q of the query tuple q = (s, p, o) ∈ QT as the number of all bound items in the tuple, BN q The corresponding expression is shown in formula (14):
[0135] BN q = 2 - |Vars(q)| (14)
[0136] where Vars(q) represents the set of all variables in the tuple q, and |Vars(q)| ≤ 2. Generally speaking, the more the number of bindings of the query tuple, the more restricted the query tuple is, so the fewer the matching results are. For example, when the number of bindings of the query tuple q = (s, p, o) is 0, s and o can be matched with any value, so the number of query results is |{(s′, o′)|(s′, o′) ∈ B p}| = |Bp |, i.e., B q all elements in. And when the binding number of the query tuple q is 1, assuming that s is a constant and o is a variable at this time, the query result is Therefore, when the query is executed, the tuple with a larger binding number is queried first. Only when the binding numbers are the same, the tuple with a smaller statistical value is considered. That is to say, for any two query tuples q = (s, p, o), q' = (s', p', o') ∈ QT, the priority of q being executed is greater than the priority of q' if and only if (1) BN q > BN q′ ; (2) BN q = BN q′ ∧ SV(B p ) < SV(B p′ ).
[0137] Step S220) Combine QT according to the possible connection methods among QT, and calculate the corresponding priority weights after combination. Select an optimal combination for each tuple, that is, the tuple in the combination is not selected and obtains the maximum weight value among all combinations; if QT only contains a single tuple, no join operation is required during query execution, so there is no need to optimize QT according to TT statistics, and directly enter step S3); otherwise, enter step S230).
[0138] The present invention uses the weight W BT (q) to represent the priority of the query tuple q = (s, p, o) ∈ QT under the BT statistics, and the corresponding calculation is shown in formula (15):
[0139]
[0140] Correspondingly, for query tuples q = (s, p, o), q' = (s', p', o') ∈ QT with the connection method PreJ ∈ {SS, OS, OO}, the calculation of the priority weights in the triple table statistics is shown in (16):
[0141]
[0142] Among them, BN q and BN q′ respectively represent the binding numbers of the query tuples q and q', and SV(tab) is the statistical value of the table tab. The binary table is obtained by dividing the triples corresponding to the abstract edges according to the predicate. For any satisfying |B p | > 0. And the triple table is obtained by connecting the binary table according to the possible connection methods (SS, OS, and OO). Therefore, there exists PreJ ∈ {SS, OS, OO}, such that At this time, the specified priority weight is 6, that is, the maximum weight value (because when When the maximum priority weight is 5).
[0143] Neighbor Tuple (NT): Given a query tuple q = (s, p, o) and q' = (s', p', o') ∈ QT, if there is an intersection between the corresponding edges, that is then the tuples q and q' are called neighbor tuples of each other.
[0144] For any tuple q = (s, p, o) ∈ QT, its neighbor tuple NT q The calculation expression is shown in Equation (17):
[0145]
[0146] Since the triple table pre-computes the connections between neighbor tuples, the SPARQL query is optimized according to the triple table to reduce the number of connections. That is, during query execution, the connections between the corresponding neighbor tuples are avoided. In a SPARQL query, a query tuple may have more than one neighbor tuple. To reduce the number of connections during query evaluation, a currently unselected tuple is preferentially selected from all neighbor tuples, thus avoiding the connections of the selected tuples.
[0147] The implementation of SPARQL query optimization is described in Figure 14 Algorithm 5 as shown. First, the query tuple set (QT) is sorted according to the priority weight. In lines 1 - 9, the entire QT is traversed, and the priority weight of each query tuple is calculated based on the BT statistics and the binding number of the query tuple. Each time, a query tuple with the maximum priority weight is taken out and saved in orderedQT. The obtained orderedQT is an ordered sequence of query tuples. Next, orderedQT is further optimized based on the pre-joined triple table statistics. In line 10, since pre-joining requires a connection between two tuples, when |QT| = 1, two tuples cannot be obtained, so the result is directly returned. In lines 11 - 27, for each currently unselected query tuple q in orderedQT, a query tuple q' is selected from its neighbor tuple NT q as its best combination. If there are still unselected query tuples in NT q , then select from these unselected query tuples. If not, select from NT q . When sorting QT, a query tuple can be taken out each time QT is traversed, so it is necessary to traverse times. When optimizing orderedQT based on pre-joining, in the worst case, it needs to loop times. Therefore, the time complexity of Algorithm 5 is O(|QT| 2 ).
[0148] Figure 7 describes the entire process of optimizing the query sequence according to the number of bindings and statistical data when querying the RDF graph in Figure 6 using the SPARQL query in Figure 2 (the corresponding summary graph is Figure 4 ). First, obtain the binding values of all query tuples themselves and the statistics in the corresponding binary table, and calculate the priority weights using formula (15), such as Figure 7 (a)-(b) in Figure 7 ). Subsequently, obtain the sorted query sequence according to the weights, such as Figure 7 (b)-(c) in Figure 7 . In
[0149] (c)-(d), use formula (17) to find the neighbor tuples of the query tuple, and calculate the corresponding weights using formula (16). Finally, on the basis of avoiding duplicate selection of tuples, select a neighbor tuple that obtains the maximum weight when combined with each query tuple, as shown in
[0150] Step S230) Further optimize the obtained optimized sequence according to the TT statistics based on the tuple join result pre-calculated by TT.
[0151] The present invention does not directly evaluate the SPARQL query sequence, but converts it into an equivalent Spark SQL statement and executes the query by the Spark system. The variables in all query tuples in the query sequence are placed after the "SELECT" keyword as query targets, the binding values are placed after the "WHERE" keyword as query conditions, and the data tables based on which the query is made are placed after the "FROM" keyword. Algorithm 6 uses the relational algebra notation to describe the mapping from query tuples to Spark SQL statements. All variables of the query tuples are added to the projection list (denoted as projections), and the constants are added to the condition list (denoted as conditions). As Figure 15 shown, the time complexity of Algorithm 6 is O(|qtSet|). Figure 7 (c)-(d) and (f)-(g) in
[0152] describe examples of converting the query sequence into the corresponding Spark SQL statements.
[0153] Step S310): Read BT from HDFS, perform a Spark SQL query based on BT on the optimized sequence that has not been statistically paired according to TT, and obtain the query result in the summary graph.
[0154] Step S320): Read TT from HDFS, perform a SparkSQL query based on TT on the optimized sequence obtained according to the TT statistical pair, and obtain the query result in the summary graph.
[0155] Matching pattern: Let the initial RDF graph be G = (V G , E G , P G , φ G ), and the graph Q = (V Q , V var , E Q , P Q , φ Q ) be the query pattern graph corresponding to the query tuple Q T . The pattern matching of graph Q on graph G is to find a binary relation that satisfies the following conditions (1) For any u ∈ V Q ∪ V var , there exists a node u' ∈ V G such that (2) For any (v, v) ∈ E Q , there exists an edge (u', v') ∈ E that satisfies G , such that
[0156] The present invention optimizes the RDF graph based on the summary graph to accelerate graph queries. Similarly, for pattern matching based on the summary graph, as follows:
[0157] Pattern matching based on the summary: Let the graph be the summary graph of the initial RDF graph G = (V G , E G , P G , φ G ), and the graph Q = (V Q , V var , E Q , P Q , φ Q ) be the graph pattern corresponding to the query tuple QT. The pattern matching of the pattern matching based on the summary (on graph Q on graph S G ) is to find a binary relation that satisfies: (1) For any u ∈ V Q ∪ V var , there exists a supernode such that (2) For any (u, v) ∈ E Q , there exists a hyperedge such that
[0158] Query result: The query result of query graph Q = (V Q , V var , E Q , P Q , φ Q ) on the RDF graph G = (V G , E G , P G , φ G ) is represented as The query result of graph Q on the summary graph is represented as
[0159] Irrelevant queries: Let QT′ and QT″ be two sets of SPARQL query tuples. QT′ and QT″ are irrelevant if and only if
[0160] Let QT be a set of SPARQL query tuples. The subqueries and satisfy QT′ ∪ QT″ = QT and If QT′ and QT″ are two irrelevant queries, it means that: (1) Since the subjects and objects of the irrelevant query tuples have no intersection, no SQL operations based on table joins need to be performed between any two query tuples from QT′ and QT″ respectively. (2) Database table join operations can only occur between the tuples within QT′ or QT″. Let the query pattern graphs of QT, QT′, and QT″ be denoted as Q, Q′, and Q″ respectively. Then, according to the definition of the query result, there is the following lemma.
[0161] Lemma 3. Let QT′ ∪ QT″ = QT and If QT′ and QT″ are not relevant in query, then
[0162] Therefore, for a given query QT, the first step of optimization is to detect whether there are irrelevant query tuples, so as to reduce the join operations of irrelevant tables during the query process, thereby accelerating the query process. In the following of the present invention, it is assumed that all tuples in the query tuple QT are relevant.
[0163] Step S330) Based on the hash table HTS, restore the summary data in the query result in the summary graph to the data in the initial RDF graph.
[0164] Since the lossless graph summarization process of the present invention adds some spurious relationships to satisfy the fully connected relationship in the node merging process, the Spark SQL query based on the summary graph will inevitably obtain some incorrect query results due to these spurious relationships. Accurate graph queries need to filter out these incorrect results.
[0165] Use to represent all the actual edges with predicate p in the hyperedge , and the corresponding calculation expression is shown in formula (18):
[0166]
[0167] Since the time complexity of calculating according to (19) is too high, and the summary edges and the actual edges they contain are stored in HTS G , in the actual implementation, calculate G from HTS , and the corresponding calculation formula is shown in formula (19):
[0168]
[0169] where and respectively represent the actual edges in the hyperedges and .
[0170] Let Q be the query pattern graph of QT, and use sumRecov((s, p, o)) to represent the set of results corresponding to the query tuple (s, p, o) ∈ QT in that are restored to the initial RDF data, and the corresponding calculation is shown in formula (20):
[0171]
[0172] where represents obtaining the attribute column with column name s from , and similarly obtain and According to formula (20), obtain the data set after the summary of the query results in the summary graph is restored, and
[0173] Step S340) Finally, execute the Spark SQL query on the data after summary restoration to obtain the final query result.
[0174] The process of data change during the execution of the Spark SQL query is as Figure 8As shown, the form '(x)-{y}' is used to represent hyperpoint information, where x represents the hyperpoint and y represents the original graph nodes contained in hyperpoint x. First, the result of the query in the summary graph is obtained based on the SparkSQL statement queried in the summary graph. Since the result of the query in the summary graph contains introduced false edges, summary recovery is required to restore the summary edges to actual edges, where the crossed-out ones are the introduced false edges. Subsequently, all binary tables are joined using the Spark SQL statement queried in the restored data to obtain the final query result.
[0175] As Figure 16 shown, Algorithm 7 described below depicts the entire process of query execution. In lines 1 - 13, first, the Spark SQL statement queried in the summary graph is obtained, and the tables (BT or TT) required for the query are cached in memory and the query is executed. In lines 1 - 13, if the number of query tuples QT is equal to 1, it indicates that pre - join optimization is not based on, so it is directly converted into the corresponding SparkSQL statement. In lines 7 - 12. In line 14, Spark executes the query of the Spark SQL statement to obtain the corresponding query result, denoted as the query intermediate result. Then, the query result in the summary graph is restored to the data in the initial graph, as shown in lines 15 - 21. Finally, the ordered query tuples are converted into the corresponding Spark SQL statements and handed over to Spark to execute the final query to obtain the final query result of the SPARQL query, as shown in lines 22 - 27 of the algorithm.
[0176] Example 2: Experimental Evaluation and Result Analysis
[0177] 2.1 Experimental Setup
[0178] The entire experiment was run on a Spark cluster with 3 worker nodes, each set up on 3 machines. Each machine is equipped with a CPU of model Intel i5 - 9300h (with 4 cores and a frequency of 3.98 GHZ), 8GB of memory, and a hard disk capacity of 1TB. The operating system used is Ubuntu 18.04, and the versions of Hadoop, Spark, and Scala are 2.6.7, 2.4.4, and 2.11.12 respectively. All experiments were repeated 3 times, and then the average value was taken as the experimental result.
[0179] 2.2 Datasets
[0180] The present invention defines the diameter of the query graph as the maximum value of the shortest path lengths between any two points when ignoring the edge directions.
[0181] For example, query graph Q=(V Q , E Q , P Q , φQ ) has a query diameter of diameter = max((minLen(v i , v j ) | v i , v j ∈ V Q}), where minLen(v i , v j ) represents the length of the shortest path between nodes v i , v j . The SPARQL Basic Graph Pattern (BGP) has four basic forms, namely linear, star, snowflake, and composite. In the linear pattern, the corresponding query graph is a straight line, and the diameter is the number of edges, i.e., |E Q |. In the star pattern, all edges meet at a single point, and the diameter is 2. The snowflake pattern is formed by connecting several star queries with short paths, and the composite pattern is composed of a combination of the other three basic patterns. Therefore, their diameters are determined by the actual graph.
[0182] The present invention selects two general datasets, namely Lehigh University Benchmark (LUBM) (http: / / swat.cse.lehigh.edu / projects / lubm / ) and Waterloo SPARQL Diversity TestSuite (WatDiv) (http: / / dsg.uwaterloo.ca / watdiv / ). LUBM is an RDF dataset describing universities, which includes information such as students, courses, professors, and schools. The number of universities in LUBM is a parameter that can control the size of the dataset. LUBM provides 14 basic query use cases, which not only include the four basic forms in BGP but also include SPARQL queries with only one query tuple (Q6 and Q14 respectively), called unit tuple patterns. The present invention generates datasets with the number of universities being 1 and 5 respectively, denoted as LUBM1 and LUBM5.
[0183] WatDiv provides rich query use cases, which also cover all forms in BGP and can comprehensively reflect the performance of the system under different query modes. WatDiv uses the parameter SF to control the size of the dataset. The present invention selects the dataset with SF = 1 (denoted as WatDiv1) for experiments. Table 1 lists the characteristics of all datasets.
[0184] Table 1 RDF Datasets
[0185] Dataset Number of Triples Number of Predicates LUBM1 100543 17 LUBM5 624530 17 WatDiv1 111804 84
[0186] 2.3 Experimental Contents and Result Analysis
[0187] The present invention compares GraSP with two current advanced and performance-excellent RDF query systems implemented on the Spark platform, namely S2RDF and S3QLRDF. S2RDF uses vertical partitioning to partition the dataset and introduces a semi-join-based RDF data relationship partitioning mode ExtVP to minimize the size of the query input, and can exhibit excellent performance under SPARQL queries of different modes and diameters. S3QLRDF proposes a new relationship splitting architecture - Property Table Partitioning (PTP), which minimizes the input data size and join operations by partitioning the existing property table into multiple sub-tables based on different properties, and has very superior performance in queries with a small number of different subjects (usually with a very small diameter, such as the star schema). Both S2RDF and S3QLRDF save corresponding statistics according to the storage characteristics of their own data for optimizing SPARQL queries, and convert the optimized query sequence into the corresponding Spark SQL statement, and use the relational interface of Spark SQL to execute the final query.
[0188] To fully evaluate the system performance, the experiments are carried out from two aspects: namely, the basic query test and the incremental linear query test (Incremental Linear Testing, IL) with the increasing number of tuples. The purpose of the basic query test is to test the performance of the system under different query modes. Since the diameter of the query cases in the basic query test is small, the incremental linear query test with the increasing number of tuples is designed to test the performance of the system under query cases with a large diameter.
[0189] 2.3.1 Basic Query Test
[0190] The LUBM dataset provides 14 basic test cases, including four modes of BGP and the unit group mode. Since there are isomorphisms in the query graphs corresponding to some query cases, only some query cases are selected for the experiment. Table 2 lists the response times of the system on the LUBM basic queries.
[0191] Table 2 Response Times (ms) of the System on LUBM Basic Queries
[0192]
[0193]
[0194] Benefiting from the aggregation strategy of the graph summary, GraSP outperforms semi-join-based S2RDF in all modes of queries, and the query time is less than half of that of S2RDF. In most query cases, GraSP outperforms S3QLRDF, and in 55% of the queries, the query time of GraSP is less than half of that of S3QLRDF. In Q4, the query time of GraSP is twice that of S3QLRDF. The reason is that Q4 has 5 query tuples in total, and all query tuples have the same subject. Therefore, S3QLRDF can respond to the query without any joins, so it shows better query performance in Q4. In Q8, the query time of GraSP is 20% more than that of S3QLRDF because Q8 is composed of two star patterns with the same subject, and S3QLRDF has optimizations for such queries. In addition, the query result set of Q8 is large. Therefore, when GraSP executes the query, the result set returned by the query in the graph summary is also large. That is to say, less irrelevant data is removed in the query of the graph summary. Therefore, the role of the graph summary in accelerating the query cannot be well demonstrated. In the snowflake query Q12 with a small query result set, when GraSP executes the query, a large amount of data irrelevant to the query result is removed in the summary graph, so it shows a performance twice as good as that of S3QLRDF. In the queries of the linear mode, since each query tuple only contains two query tuples and has a common subject, the query time of GraSP is only 21% less than that of S3QLRDF.
[0195] Table 3 lists the response times of each system for the basic queries in WatDiv. Similar to its performance in the basic queries of LUBM, GraSP outperforms S2RDF in all basic queries of WatDiv, and for 75% of the queries, the query time of GraSP is no more than half of that of S2RDF. At the same time, GraSP outperforms S3QLRDF in 70% of the basic queries of WatDiv. In the star schema, all query tuples in S2, S3, and S5 have a common subject, so S3QLRDF performs better than GraSP, while GraSP performs better than S3QLRDF in the remaining 4 query cases. In all snowflake and linear queries, GraSP outperforms S3QLRDF. Since C3 contains 6 query tuples and the 6 tuples have a common subject, the query time of GraSP is 90% longer than that of S3QLRDF. In all snowflake and linear queries, GraSP outperforms S3QLRDF. Queries C2 and F2 contain predicates that do not exist in the RDF dataset and cannot find corresponding predicates during the query, so the S2RDF and S3QLRDF systems prompt errors (indicated by "-" in Table 3), while when GraSP converts SPARQL to Spark SQL (introduced in Section 5.2), when it detects that the table corresponding to the predicate does not exist in the BT or TT sets, it will return an empty query result, and obviously this processing method is more reasonable. Since such queries are given the highest priority value during query optimization, when converting SPARQL to Spark SQL statements during the query, it is possible to determine whether the table corresponding to the predicate exists, so the query response time is very short.
[0196] Table 3 Response times (ms) of the systems for the basic queries in WatDiv
[0197]
[0198] It can be clearly seen from the basic queries that GraSP based on graph summarization and pre-join outperforms semi-join-based S2RDF in all query modes. For S3QLRDF based on property table partitioning, due to the characteristics of the property table storage mode, its performance is better than GraSP when the number of different subjects of the query tuples is small. However, in other queries, GraSP outperforms S3QLRDF in query performance, even by two orders of magnitude.
[0199] 2.3.2 Linear queries with increasing number of tuples
[0200] The basic query use cases provided by LUBM and WatDiv are suitable for testing the performance of the system under different query shapes, but the diameters of most query use cases are small. In fact, the diameters of only two query use cases are greater than 3, namely C1 and C2. Most current RDF query systems are optimized for small-diameter queries. These systems usually place triples with the same subject in the RDF dataset on the same cluster nodes close to each other or make a copy of them and place them on the same machine node based on hash or graph partitioning algorithms. Therefore, not much data exchange or long-distance data transmission is required when evaluating small-diameter queries.
[0201] To test the performance of the system under large-diameter query use cases, the present invention designs an additional experiment on WatDiv, called the linear query test with increasing number of tuples, which focuses on linear queries with an increasing number of query tuples (i.e., increasing the diameter of the query). The linear query test with increasing number of tuples selects the provided test cases ( https: / / github.com / nosrepus / Stream-WatDiv / tree / master / testsuite / linear_incremental ) as the extended experimental test cases on the WatDiv dataset. The test cases of the increasing linear query test are composed of 3 types (IL-1, IL-2, IL-3), which represent user binding, retailer binding, and no binding respectively, and the diameter size of each type increases from 5 to 10.
[0202] Table 4 Linear Queries with Increasing Number of Tuples in WatDiv (ms)
[0203]
[0204] Table 4 records the query response times of all systems in different linear query use cases. For large-diameter queries, GraSP shows performance superior to S3QLRDF by nearly two orders of magnitude. This is because S3QLRDF is optimized for queries with fewer different subjects and performs well in star queries (usually subject-subject connected), but it cannot leverage its advantages in linear queries (usually object-subject connected). In addition, S3QLRDF is based on property table partitioning. During the join operation, since there are columns in the input data table that are irrelevant to the query, it will increase its memory overhead and the burden of data transmission during communication. Therefore, S3QLRDF has the longest query time in all test cases. S2RDF shows very good performance for large-diameter queries because it adopts a semi-join-based data storage strategy and is not optimized for a specific pattern, so it can perform well in queries of different shapes and diameters. As in the basic query test results, GraSP benefits from graph summarization to accelerate queries, so it outperforms S2RDF in all query cases by more than one order of magnitude.
[0205] The linear query test with an increasing number of tuples shows that the performance of some distributed RDF stores significantly degrades for such workloads because their data models are optimized to answer small-diameter queries. As in the basic test cases, GraSP also demonstrates excellent performance in large-diameter queries. Overall, GraSP has excellent performance and scalability.
[0206] 2.3.3 Impact of Summary Size on Query Speed
[0207] The summary graph of a data graph is its abstract representation, much smaller in scale than the data graph but retaining the structure and information of the data graph. Therefore, querying in the summary graph only takes a small amount of time, yet it can obtain approximately the results of querying the query graph in the data graph (including the correct query results and introduced false edges). The present invention represents the summary size as the ratio of the total number of supernodes in the graph summary to the total number of nodes in the data graph, that is If the summary size is larger, querying in the summary graph will take more time, and the acceleration effect of the graph summary for querying becomes less obvious. If the summary graph size is smaller, more false edges are introduced in the summary graph, and more time will be consumed in summary restoration and querying in the restored data. Therefore, a reasonable summary size should be selected. To obtain the value of the summary size when the average query time is the shortest, the present invention tested the impact on query time under different summary sizes.
[0208] Figure 9Under 14 basic test cases, the impact of different summary sizes on the average query time of the system under different scales of the LUBM dataset. When the summary size is not less than 40%, as the summary ratio decreases, the average query time also gradually decreases, indicating that the graph summary has a positive effect on accelerating graph queries. When the summary size is between 20% and 40%, the average query time is roughly the same. When the proportion is less than 20%, as the summary size gradually decreases, the average query time gradually increases. When the summary size is 20%, that is, when the total number of supernodes in the summary graph is 20% of the total number of nodes in the original graph, the shortest average query response time is obtained.
Claims
1. A complex graph query optimization method based on vertical partitioning and pre-connection of summary graphs, characterized in that It includes the following steps: Step S1) Process the initial RDF triple data to obtain the hash table HTS, binary table BT, ternary table TT, and BT statistics and TT statistics after graph summarization. HTS, BT, and TT are all stored on HDFS; Step S2) Optimize the SPARQL query tuple QT based on the binding number of QT and BT statistics to obtain an initial optimized sequence. Determine whether it is necessary to further optimize the obtained optimized sequence according to TT statistics based on the number of query tuples included in QT; Step S3) Read the corresponding tables from HDFS, execute the Spark SQL query based on the summary to obtain the query result in the summary graph; perform data restoration based on the hash table HTS, and then perform the Spark SQL query to obtain the final query result.
2. The complex graph query optimization method based on vertical partitioning and pre-connection of summary graphs according to claim 1, characterized in that The specific steps of step S1) are as follows: Step S110) Process the initial RDF triples based on aggregated graph summarization to obtain the hash table HTS storing all summary information; Step S120) Vertically partition the summary edges stored in HTS according to the predicates to obtain the binary table BT and the corresponding BT statistics; Step S130) Calculate the pre-join of the binary table BT to obtain TT and the corresponding TT statistics; Step S140) HTS, BT, and TT are all stored on HDFS.
3. The complex graph query optimization method based on vertical partitioning and pre-connection of summary graphs according to claim 1, characterized in that, The specific steps of step S2) are as follows: Step S210) Optimize the SPARQL query tuple QT based on the binding number of QT and BT statistics to obtain an initial optimized sequence; Step S220) Combine QTs according to the possible connection methods between QTs, and calculate the corresponding priority weights after combination. Select a best combination for each tuple, that is, the tuple in the combination is not selected and obtains the maximum weight in all combinations; if QT only contains a single tuple, no join operation is required during query execution, so there is no need to optimize QT according to TT statistics and directly enter step S3), otherwise enter step S230); Step S230) Further optimize the obtained optimized sequence according to TT statistics based on the pre-calculated tuple connection results of TT.
4. The complex graph query optimization method based on vertical partitioning and pre-connection of summary graphs according to claim 1, characterized in that The specific steps of step S3) are as follows: Step S310) Read BT from HDFS and execute the SparkSQL query based on BT for the optimized sequence that has not been optimized according to TT statistics to obtain the query result in the summary graph; Step S320) Read TT from HDFS and execute the Spark SQL query based on TT for the optimized sequence optimized according to TT statistics to obtain the query result in the summary graph; Step S330) Restore the summary data in the query result in the summary graph to the data in the initial RDF graph based on the hash table HTS; Step S340) Finally, execute the Spark SQL query in the data after summary restoration to obtain the final query result.
5. The complex graph query optimization method based on vertical partitioning and pre-connection of summary graphs according to claim 2, characterized in that The specific steps of step S110) are as follows: Step S111) The initial RDF graph is a summary graph in which all super nodes contain only one original graph node, and it is stored in a hash table, denoted as HTG; the definition of the initial RDF graph as a directed labeled graph is: Graph G = (V G , E G , P G , φ G ), where (1) V G corresponds to the set of all subjects and objects in the RDF triples, (2) corresponds to the set of directed edges in all RDF triples, (3) P G is the set of labels on all edges, (4) φ G is a label mapping function φ G : denotes the assignment of a subset of P G to the edge e ∈ E G ; Step S112) Aggregate two similar supernodes in a node-aggregation-based manner to merge them into a supernode and ensure that the two merged supernodes meet a certain similarity through error calculation; where, let the graph be the summary graph of graph G=(V G , E G , P G , φ C ), where 1) is a set of super points and satisfies: (1): Let π(v) denote the node v ∈ V G in the graph S G the belonging super node; 2) is the hyperedge set. For any hyperedge denotes that there is a full connection between the super nodes and i.e., ; 3) denotes the set of attribute labels on all hyperedges in the graph S; 4) φ S denotes the set of attribute labels of each hyperedge and for any two super nodes, the similarity between the super nodes is calculated as shown in formula (1). Among them, Denote the superpoint the set of all neighbor superpoints; similarly, obtain Step S113) When the edge merging error is less than the error critical value, the superpoint pair is merged; among them, for any two superpoints If they have been aggregated into a new superpoint then the new superpoint will replace and appear in satisfy and 6. The complex graph query optimization method based on vertical partitioning and pre-connection of summary graphs according to claim 5, characterized in that, The calculation of the edge merging error ME is realized through the internal merging error IME and the adjacency merging error AME: 1) Internal Merging Error IME: The internal merging error of any two super points is defined as Formula (2): The internal merging error Among them, , respectively represent super nodes and the full connection and the actual connection between the corresponding node sets; 2) Adjacent Merging Error AME: When the superpoints are merged into the superpoint the introduced adjacent merging error is defined as shown in formula (3): 3) Merging Error ME: Merging two super points The merging error is defined as Equation (4): Set if That is, only when there is an actual edge between two super nodes, a false edge will be introduced to form a fully connected super edge, and the internal merging error will be calculated.
7. The complex graph query optimization method based on vertical partitioning and pre-connection of summary graphs according to claim 2, characterized in that, The vertical partitioning divides the triple table into different sub-tables according to the predicate p, and each sub-table only stores two columns in the triple, namely (s, o), to avoid duplicate storage of predicates: Abstract Drawing corresponding to a set of triples Convert Figure S G According to the predicate The binary table obtained by vertical partitioning is denoted as B P , and the corresponding set expression is shown in formula (6): Since Figure S G is stored in the HTS G , so it is necessary to obtain from the HTS G , and the corresponding calculation expression is shown in Formula (7): Among them, is the abbreviation of , is the abbreviation of , respectively representing the set of all actual edges in the hyperedge and ; Let BT represent the set of all predicate binary tables, and the calculation expression of BT is shown in formula (8): In summary, each predicate in the triple table uniquely corresponds to a binary table and functions as an index. During querying, only the binary tables related to the query need to be joined.
8. The complex graph query optimization method based on vertical partitioning and pre-connection of summary graphs according to claim 2, characterized in that The pre-joining of the binary table BT is to reduce a large number of join operations during query execution and remove the results that do not meet the join conditions by pre-computing the possible join results of the query tuples: Let the query tuple be q i =(s, p i , o), q j =(s′, p j , o′) ∈ QT. If s = s′, then the query results of the query tuple q i and q j are That is, the connection result at the subject of the binary table corresponding to the predicates p i , p j . Such a connection is called subject- S subject (SS) connection, and the connection result S is as shown in formula (9): As shown in formula (9): In addition to S ubject- S ubject (SS) connection, the existing connection methods also include S ubject- O bject (SO), O bject- S ubject (OS) and O bject- O bject (OO), and the corresponding connection results are shown in the following formulas (10), (11), and (12): It is verified that among the four connection methods, the results of half of the connection operations are obtained by other connections; therefore, when calculating the connection result of two binary tables, only half of the S ubject- S ubject(SS), O bject- O bject(OO) connections, and all of the O bject- S ubject(OS) connections. Let TT represent the set of all ternary tables obtained by connecting binary tables. The corresponding set expression is shown in Equation (13):
9. The complex graph query optimization method based on vertical partitioning and pre-connection of summary graphs according to claim 8, characterized in that, The specific steps for optimizing the query sequence based on the binding number of QT and the statistics of BT are as follows: Step S211) First, obtain the self-bound values of all query tuples and their statistics in the corresponding binary table, and calculate the priority weight using formula (15); use the weight W BT (q) represents the priority of the query tuple q = (s, p, o) ∈ QT under the BT statistics, and the corresponding calculation is as shown in the formula As shown in (15): In step S212), the sorted query sequence is obtained according to the weight sorting; In step S213), the neighbor tuples NT of the query tuples are obtained using formula (17), and the corresponding weights are calculated using formula (16); Let the query tuple q = (s, p, o), q' = (s', p', o') ∈ QT. If there are intersections between the corresponding edges, that is then the tuples q and q' are called neighbor tuples of each other; for any tuple q = (s, p, o) ∈ QT, its neighbor tuple NT q The calculation expression is shown in Equation (17): For connection types PreJ ∈ { S ubject- S ubject(SS), O bject- S ubject(OS), O bject- O bject(OO)}, the calculation of the priority weights of query tuples q = (s, p, o), q' = (s', p', o') ∈ QT in the triple table statistics is shown in (16): Among them, BN q and BN q′ respectively represent the number of bindings of the query tuples q and q′, and SV(tab) is the statistical value of the table tab; the binary table is obtained by dividing the triples corresponding to the summary edges according to the predicates. For any satisfies |B p | > 0; the ternary table is obtained by connecting the binary table according to the possible connection methods S ubject- S ubject(SS), O bject- S ubject(OS) and O bject- O bject(OO). So there exists PreJ ∈ { S ubject- S ubject(SS), O bject- S ubject(OS), O bject- O bject(OO)}, such that If the specified priority weight is M, that is, the maximum weight, because when , the maximum priority weight is M - 1; In step S214), on the basis of avoiding duplicate selection of tuples, a neighbor tuple with the maximum weight after combination is selected for each query tuple.
10. The complex graph query optimization method based on vertical partitioning and pre-connection of summary graphs according to claim 4, characterized in that In the step S330), the summary data in the query results retrieved from the summary graph is restored to the data in the initial RDF graph based on the hash table HTS, that is, the incorrect query results obtained by querying some spurious relationships added due to satisfying the full connection relationship in the node merging process are filtered out: Use to represent the hyperedge All actual edges within with the predicate p, and the corresponding calculation expression is shown in formula (18): Since the time complexity of calculating is too high, and the summary edges and the actual edges they contain are stored in HTS G in the actual implementation, calculate from HTS G The corresponding calculation formula is as shown in formula (19): The corresponding calculation formula is as shown in formula (19): Among them, and respectively represent the actual edges and in the hyperedge; Let Q be the query pattern graph of QT, and use sumRecov((s, p, o)) to represent that the result corresponding to the query tuple (s, p, o) ∈ QT in is restored to the set of initial RDF data, and the corresponding calculation is shown in formula (20): Among them indicates obtaining the attribute column with the column name s from similarly, obtain and According to formula (20), obtain the data set after the result summary in the abstract graph query is restored, and
Citation Information
Patent Citations
Spatial RDF data keyword query method based on abstract graph
CN110222240A
Data storage and query method and device based on distributed RDF
CN110825738A