A Method and Device for Distributed Storage of Graph Data
By performing Jena analysis and relational model construction on RDF ontology data, semantic similarity between vertices and graph division, the problems of low quality of graph data division and insufficient semantic consideration in the prior art are solved, and efficient graph data query processing is achieved.
Patent Information
- Application Number
- CN202310476159.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-04-27
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2043-04-27
AI Technical Summary
The existing graph data division algorithm has low division quality on large graph data and lacks consideration of semantic dimensions, resulting in poor query processing time.
By Jena analysis of RDF ontology data, a relationship tree is generated and a relationship model is constructed, the semantic similarity between vertices is calculated, and the graph division algorithm is used to maximize the semantic coherence to obtain the division result.
A partitioning scheme that maximizes semantic coherence in the static division stage is realized, which improves the query processing efficiency of graph data and adapts to changes in query workloads.
Smart Images

Figure CN116501739B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of software engineering. Specifically, it relates to a method and device for distributed storage of graph data. Background Art
[0002] Methods for storing and querying RDF (Resource Description Framework) graph data in a distributed environment have become a popular research direction in the academic community. Corresponding large-scale graph data distributed storage systems are divided into three categories. The first category is to use a cloud platform to store and manage graph data, use a distributed file system to store RDF graph data, and process SPARQL queries through the MapReduce programming model. The main difference between different storage systems is the way of converting RDF graphs into underlying data structures. This type of system has good scalability and fault tolerance, and most are oriented to offline analysis of data. The second category is to distribute RDF graph data to different nodes based on data. Each single-machine graph data management system manages RDF subgraphs, and divides SPARQL queries into each node to calculate partial solutions, and aggregates the partial solutions to obtain the final solution. The main difference between different storage systems is mainly the difference in graph data partitioning strategies. This type of system has a small communication cost, but is overly dependent on the data partitioning method and has certain limitations. The third category is a federated RDF graph data management system. Each data owner is independent and autonomous as a data source, independently processes SPARQL sub-query calculations for partial solutions, and the graph data management system obtains the final solution by connecting the partial solutions.
[0003] However, the above systems also have two disadvantages. On the one hand, the existing graph data partitioning algorithms have low partitioning quality for large graph data, need to simultaneously meet the minimization of cut edges and load balancing, and have a high time complexity problem. On the other hand, RDF graph data contains rich semantic information, and previous graph partitioning metrics lack consideration of the semantic dimension. A good partitioning algorithm needs to adapt to changes in query workloads and maintain good query processing times. Compared with static graph data partitioning algorithms, dynamic graph data partitioning algorithms often have better performance in large-scale graph data query processing. Graph data query and storage are closely related. An efficient query engine is a very important part of a graph data storage system. SPARQL is the standard graph data query language, and the essence of a SPARQL query is a subgraph matching problem. Therefore, it is necessary to design an efficient subgraph matching algorithm to quickly select a result set in a large-scale data graph. Summary of the Invention
[0004] The purpose of this application is to provide a method and device for distributed storage of graph data to overcome the existing technical defects. In the static partitioning stage, by calculating the similarity of classes in the ontology structure, a partitioning scheme that maximizes semantic coherence is sought.
[0005] The purpose of this application is achieved through the following technical solutions:
[0006] In the first aspect, this application proposes a method for distributed storage of graph data, including:
[0007] Perform Jena parsing on the RDF ontology data, generate a relationship tree based on the relationships between classes, and construct a relationship model;
[0008] Define the distance similarity, structural similarity, and attribute similarity between vertices according to the hierarchical structure of the relationship tree in the relationship model;
[0009] Calculate the semantic similarity between vertices, and obtain the corresponding semantic coherence through the semantic similarity. The semantic similarity is a weighted value of the distance similarity, the structural similarity, and the attribute similarity;
[0010] Use a graph partitioning algorithm to maximize the partitioning of the semantic coherence to obtain a partitioning result.
[0011] In a possible implementation manner, the RDF ontology data includes a plurality of triple items. The step of performing Jena parsing on the RDF ontology data includes:
[0012] Group the subjects and objects of the plurality of triple items by class to obtain the classes to which the plurality of triples belong;
[0013] Assign a class code to each class to which a triple belongs as an identifier, and use the hash value of the URI prefix of each triple item as a prefix code;
[0014] Obtain the suffix code of the triple item according to the class to which the triple item belongs and the sequential number of the prefix code;
[0015] Obtain the triple item code according to the class code, the prefix code, and the suffix code.
[0016] In a possible implementation manner, the calculation formula of the class code CC(t,i) is: CC(t,i) = Flag & Num & f(t - 1, m) & f(t, i), where Flag is the class flag bit, Num is the number of direct parent classes, f(t - 1, m) represents the sequential code of the parent class node Y of node X, and f(t, i) represents the sequential code of the i-th node X in the t-th layer, which is expressed as: g(t - 1) is the sequential code of the class node in the (t - 1)-th layer.
[0017] In a possible implementation, the distance similarity Sim D (x, y) is as follows: X represents the class to which vertex x belongs, Y represents the class to which vertex y belongs, D(X) and D(Y) respectively represent the path lengths from X and Y to the least common superclass, D(LCC(X, Y)) represents the path length from LCC(X, Y) to the root node, and LCC(X, Y) represents the least common superclass of X and Y;
[0018] The structure similarity Sim S (x, y) is as follows: I(X) is the semantic feature information amount of class X, I(Y) is the semantic feature information amount of class Y, and I(LCC(X, Y)) is the semantic feature information amount of LCC(X, Y);
[0019] The attribute similarity Sim A (x, y) is as follows:
[0020]
[0021] where p D is a function regarding the influence magnitudes of class X and class Y on the attribute similarity, and μ(X) and μ(Y) respectively correspond to the number of attributes of class X and class Y.
[0022] In a possible implementation, the steps after maximizing the partitioning of the semantic coherence by using the graph partitioning algorithm to obtain the partitioning result include:
[0023] Optimizing the partitioning result by determining the load transfer amount and the vertex movement target partition;
[0024] Calculating the signature of each vertex, where the signature is composed of vertex information and the structural information adjacent to the vertex;
[0025] Partitioning the data graph G of the RDF ontology data into star-shaped substructures through the signature, and pruning the graph data matching process according to the graph structure information provided by the vertex signature.
[0026] In a possible implementation, the steps of pruning the graph data matching process include:
[0027] Obtaining all types Class(Q) of vertices in the query graph Q;
[0028] Comparing each type in Class(Q) with each vertex in the data graph G, and removing the vertices that do not belong to Class(Q) and all their adjacent edges in the data graph G.
[0029] In a possible implementation, the step of pruning the graph data matching process further includes:
[0030] Traverse all vertices in the query graph Q, and denote the minimum degree of the query graph Q as degree min (Q);
[0031] Traverse all vertices in the data graph G, and delete all vertices and adjacent edges that satisfy the vertex degree degree(v) < degree min (Q).
[0032] In a second aspect, the present application proposes a graph data distributed storage device, and the device includes:
[0033] A construction module for performing Jena parsing on RDF ontology data, generating a relationship tree according to the relationships between classes, and constructing a relationship model;
[0034] A similarity generation module for defining the distance similarity, structural similarity, and attribute similarity between vertices according to the hierarchical structure of the relationship tree in the relationship model;
[0035] A semantic calculation module for calculating the semantic similarity between vertices, and obtaining the corresponding semantic coherence through the semantic similarity, where the semantic similarity is a weighted value of the distance similarity, the structural similarity, and the attribute similarity;
[0036] A generation module for maximizing the division of the semantic coherence by using a graph partitioning algorithm to obtain a division result.
[0037] In a third aspect, the present application further proposes a computer device, the computer device includes a processor and a memory, and a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the graph data distributed storage method as described in any item of the first aspect.
[0038] In a fourth aspect, the present application further proposes a computer-readable storage medium, and a computer program is stored in the storage medium, and the computer program is loaded and executed by a processor to implement the graph data distributed storage method as described in any item of the first aspect.
[0039] The main solution of the present application and its various further alternative solutions can be freely combined to form multiple solutions, all of which are solutions that can be adopted and claimed by the present application; and in the present application, (each non-conflicting option) options can be freely combined with each other and with other options. Those skilled in the art can understand that there are various combinations according to the prior art and common general knowledge after understanding the solution of the present application, and all of them are the technical solutions to be protected by the present application, and will not be enumerated here.
[0040] This application discloses a method for distributed storage of graph data. First, the RDF ontology data is parsed by Jena, a relationship tree is generated according to the relationships between classes, and a relationship model is constructed. The distance similarity, structural similarity, and attribute similarity between vertices are defined according to the hierarchical structure of the relationship tree in the relationship model, the semantic similarity between vertices is calculated, and the corresponding semantic coherence is obtained through the semantic similarity. Finally, a graph partitioning algorithm is used to maximize the partitioning of the semantic coherence to obtain the partitioning result. In the static partitioning stage, the relationships between ontology concepts are represented by a tree structure, a relationship model of classes and attributes is constructed, so that the RDF data encoding embeds the semantic information of the ontology, and by calculating the similarity of classes in the ontology structure, a partitioning scheme that maximizes semantic coherence is sought. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] Figure 1 FIG. shows a schematic flowchart of a method for distributed storage of graph data proposed in an embodiment of this application.
[0042] Figure 2 FIG. shows a comparison chart of the number of cut edges of different partitioning algorithms proposed in an embodiment of this application.
[0043] Figure 3 FIG. shows a comparison chart of semantic consistency of different partitioning algorithms proposed in an embodiment of this application.
[0044] Figure 4 FIG. shows a schematic diagram of the query time of different partitioning algorithms proposed in an embodiment of this application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0045] The following uses specific specific examples to illustrate the implementation manners of this application. Those skilled in the art can easily understand the other advantages and effects of this application from the content disclosed in this specification. This application can also be implemented or applied through other different specific implementation manners, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of this application. It should be noted that, without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0046] Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without making creative efforts belong to the scope protected by this application.
[0047] From the solutions disclosed in the prior art, the RDF distributed storage method based on a multi-layer partitioning framework stores different data subsets on different nodes by adopting a multi-layer partitioning framework. Although the storage efficiency can be improved, the storage capacity of each node is limited. When the scale of the stored data exceeds the capacity of the node, it may be necessary to expand or add nodes, increasing the storage cost.
[0048] The large-scale knowledge graph storage solution based on a distributed key-value store uses a key-value store. Each query requires looking up through the key-value, which will increase the complexity and time of the query and affect the query performance. At the same time, since data synchronization and communication are required between multiple nodes, this will also affect the read and write performance.
[0049] A graph data storage method and a distributed graph data calculation method use distributed computing and allocate computing tasks to different computing nodes. However, since different nodes may be responsible for different subgraphs, if the subgraph sizes are unbalanced, it may lead to unbalanced loads on the computing nodes, with some nodes overloaded and other nodes underloaded.
[0050] Due to the above-mentioned drawbacks in the prior art, it does not conform to the design of an efficient subgraph matching algorithm and cannot achieve quickly selecting a result set in a large-scale data graph. Therefore, the embodiments of this application propose a graph data distributed storage method and device, a large-scale graph data distributed storage solution based on Hbase, to improve the efficiency of data storage and query. In the static partitioning stage, by calculating the similarity of classes in the ontology structure, a partitioning scheme that maximizes semantic coherence is sought. The following is a detailed description of it.
[0051] Please refer to Figure 1 , Figure 1 FIG. shows a schematic flowchart of a graph data distributed storage method proposed by the embodiments of this application. The large-scale graph data distributed storage method has broad application prospects. It can help users process massive graph data, improve the reliability and fault tolerance of data, achieve rapid data sharing and collaboration, and achieve rapid data processing and analysis. This method can be applied to multiple technical fields, such as: social network analysis, road network analysis, bioinformatics, financial risk analysis, and e-commerce platforms, etc. The effects produced by it are described as follows:
[0052] Social network analysis: A social network is a typical graph structure, where nodes represent people or organizations and edges represent the relationships between people. By analyzing the social network, information such as interpersonal relationships, social influence, and hot topics can be mined. Using the distributed storage method can process massive social network data and improve the efficiency of social network analysis.
[0053] Road network analysis: A road network is a complex graph structure, where nodes represent intersections or road junctions and edges represent the relationships between roads. By analyzing the road network, information such as traffic conditions, vehicle flow, and congestion can be mined. Using the distributed storage method can process massive road network data and improve the efficiency of road network analysis.
[0054] Bioinformatics: Bioinformatics is a field involving a large number of gene and protein sequences, where the sequences can be regarded as a graph structure. By analyzing gene and protein sequences, their roles and significance in biology can be explored. Using a distributed storage method can process massive sequence data and improve the efficiency of bioinformatics analysis.
[0055] Financial Risk Analysis: The financial market is a complex system that contains a large number of financial products and trading relationships. By analyzing the financial market, information such as market trends and risk distributions can be mined. Using a distributed storage method can process massive financial data and improve the efficiency of financial risk analysis.
[0056] E-commerce platforms usually have a large amount of data such as users, products, and orders, which can be regarded as a graph structure. By analyzing the data of e-commerce platforms, sales, precision marketing, and user satisfaction can be improved. Using a distributed storage method can process massive e-commerce data and improve the efficiency of e-commerce platform analysis.
[0057] A graph data distributed storage method proposed in an embodiment of this application includes the following steps:
[0058] Step S100: Perform Jena parsing on RDF ontology data, generate a relationship tree according to the relationships between classes, and construct a relationship model.
[0059] When generating the relationship tree, first set the root node of the ontology tree as R, and construct the relationship tree through construction rules. The construction rules are: and where I * (X,Z) means that X and Z are sibling nodes, I(X,Y) means that X is a subclass of Y, and so on.
[0060] The RDF ontology data includes multiple triple items. The steps of performing Jena parsing on the RDF ontology data include:
[0061] Group the subjects and objects of multiple triple items by class to obtain the classes to which multiple triples belong;
[0062] Assign class codes with each class to which a triple belongs as an identifier, and use the hash value of the URI prefix of each triple item as the prefix code;
[0063] Obtain the suffix code of the triple item according to the order number of the class to which the triple item belongs and the prefix code;
[0064] Obtain the triple item code according to the class code, prefix code, and suffix code.
[0065] Before calculating the class encoding, it is necessary to first calculate the number of digits of the node encoding (NodeDigit, ND). If the total number of nodes is T, the number of digits of the node encoding ND(T) under this total number of nodes T is: floor() represents the upper bound. Then, calculate the class encoding (ClassCode, CC). The calculation formula for the class encoding CC(t,i) is: CC(t,i) = Flag & Num & f(t - 1, m) & f(t,i), where Flag is the class flag bit, Num is the number of direct parent classes, f(t - 1, m) represents the sequential encoding of the parent node Y of node X, and f(t,i) represents the sequential encoding of the i-th node X in the t-th layer, which is expressed as: g(t - 1) is the sequential encoding of the class node in the (t - 1)-th layer. When the number of direct parent classes Num is greater than 1, the sequential encoding of the parent node is composed of the combined sequential encodings of all direct parent nodes. The number of digits of the parent node encoding and the sequential encoding of this node are both ND(T), and CC(t,i) represents the encoding of the i-th node in the t-th layer.
[0066] In the embodiment of the present application, the triple is encoded. First, the subject and object in the triple item are grouped by class, the class to which each triple belongs is used as an identifier to assign the class encoding, the hash value of the URI prefix of each triple item is used as the prefix encoding of the triple item, and the suffix encoding of the triple item is obtained according to the class to which the triple item belongs and the sequential numbering of the prefix encoding. Combining the class encoding, prefix encoding, and suffix encoding of the triple, the final encoding of the triple item is obtained. The encoding format of each triple item is: class encoding + prefix encoding + suffix encoding. By establishing a dictionary table from the triple item to its encoding, the encoding corresponding to the triple item can be quickly obtained, and the triple encoding can also be quickly restored to the original data.
[0067] In addition, the property encoding (PropertyCode, PC) is composed of the property bit flag Flag, the class encoding to which it belongs, the sequential encoding of the parent node, and the sequential encoding of this node. On the premise that the total number of nodes is T, the number of digits of the sequential encoding of the parent node and the sequential encoding of this node is ND(T), and PC(t,i) represents the property node encoding of the i-th node C in the t-th layer. The formula is PC(t,i) = Flag & NC(p,r) & f(t - 1, m) & f(t,i), where The class to which C belongs is set as L, the class node encoding of L is represented as NC(p,r), f(t,i) represents the sequential encoding of the i-th node C in the t-th layer. g(t) represents the sequential encoding of the property node in the t-th layer, and f(t - 1, m) represents the sequential encoding of the parent node D of node C.
[0068] Step S200: Define the distance similarity, structural similarity, and attribute similarity between vertices according to the hierarchical structure of the relationship tree in the relationship model.
[0069] To measure the similarity of vertices, based on the hierarchical structure of the ontology tree, the definitions of distance similarity, depth similarity, and attribute similarity are proposed, and the similarity of vertices is judged by synthesizing the similarities of three different dimensions.
[0070] The distance similarity Sim D (x, y) between vertex x and y is:
[0071]
[0072] Let X represent the class to which vertex x belongs, Y represent the class to which vertex y belongs, D(X) and D(Y) respectively represent the path lengths of X and Y to the least common superclass, D(LCC(X, Y)) represents the path length of LCC(X, Y) to the root node, and LCC(X, Y) represents the least common superclass of X and Y. Since the semantic similarity between class X and class Y can use distance as an evaluation index, when the depth of their least common superclass is greater, the distance similarity between the two classes is higher; when the respective depths of class X and class Y are greater, the distance similarity between the two classes is lower. That is, the distance similarity is positively correlated with the depth of the least common superclass and negatively correlated with their respective depths.
[0073] The structural similarity Sim S (x, y) between vertex x and vertex y is: Let I(X) be the semantic feature information amount of class X, I(Y) be the semantic feature information amount of class Y, and I(LCC(X, Y)) be the semantic feature information amount of LCC(X, Y). The calculation formula for the semantic feature information amount I(X) of class X is: I(X) = -log(r(X)), where r(X) represents the ratio of the number of subclasses of X to the total number of classes. Similarly, the calculation formula for the semantic feature information amount I(Y) of class Y is: I(Y) = -log(r(Y)), where r(Y) represents the ratio of the number of subclasses of Y to the total number of classes.
[0074] Except for the root node, each class in the ontology has a superclass, corresponding to the inheritance chain from the superclass to the subclass in the ontology tree. The amount of semantic features contained in a class increases with the increase in the length of the inheritance chain. The structural similarity between two classes is related to the semantic information amount. The structural similarity between class X and class Y is proportional to the semantic information amount contained in their least common superclass and inversely proportional to the semantic information amount contained in themselves.
[0075] The attribute similarity Sim A (x, y) is:
[0076]
[0077] Among them, μ(Y) and μ(Y) respectively correspond to the number of attributes of class X and class Y, pD (X, Y) is a function that measures the impact of classes X and Y on attribute similarity, and its calculation formula is: where H(X) is the depth of class X and H(Y) is the depth of class Y.
[0078] Since the semantic similarity between classes X and Y can be measured by the number of common attributes, the more common attributes there are, the higher the attribute similarity. The measurement of the attribute similarity between classes X and Y is affected by the depth difference between them. The attribute similarity is determined by the number of common attributes and the difference in the number of attributes, being directly proportional to the number of common attributes and inversely proportional to the integrated difference in the number of attributes.
[0079] Step S300: Calculate the semantic similarity between vertices and obtain the corresponding semantic coherence through the semantic similarity.
[0080] The semantic similarity is the weighted value of distance similarity, structural similarity, and attribute similarity. The calculation method of the semantic similarity Sim(x, y) is: Sim(x, y) = αSum D (x, y) + βSum S (x, y) + γSum A (x, y), where α (0 ≤ α ≤ 1), β (0 ≤ β ≤ 1), and γ (0 ≤ γ ≤ 1) are three coefficients respectively, and the semantic similarity represents the weights of the three similarities each accounting for a certain proportion.
[0081] The partitioning scheme based on semantic similarity measures the partitioning criterion by using the semantic coherence of the partitions. The semantic coherence represents the correlation between the vertices at each partition boundary and each partition. Assigning vertices to partitions with higher semantic coherence can significantly reduce the possibility of cross-region queries, and the objective function of partitioning is to maximize the semantic coherence of the graph data.
[0082] Set the RDF ontology data as graph G = (V, E), and partition graph G = (V, E) to obtain P = {P1, …, P k}, vertex v i ∈ V, the corresponding partition is p i , the set of adjacent vertices of v i is denoted as v i ∈ NBR(v i ). At this time, the semantic coherence of vertex v i is the ratio of the sum of the semantic similarities from v i to its adjacent vertices NBR(v i ) in partition p i to the sum of the similarities from v i to all its adjacent vertices: where Sim(v i , v j ) represents the similarity between vertex vi and vertex v j semantic similarity.
[0083] Step S400: Use the graph partitioning algorithm to maximize the partitioning of semantic coherence to obtain the partitioning result.
[0084] The graph partitioning algorithm is a re - partitioning algorithm based on semantic information. To maximize semantic coherence, re - partition the vertices that meet the conditions on the basis of the initial partition. It is divided into a partition calculation stage and a re - partitioning stage as a whole. The input is the RDF graph G, the initial partition P = {P1,…,P k}, the number of iterations δ of the partitioning algorithm, and the balance factor θ. The output is the partitioning result that satisfies the maximization of semantic coherence, as follows:[[]]
[0085]
[0086] The goal of semantic - based graph data partitioning is to maximize the sum of the semantic coherence of vertices v i ∈V, and its objective function is:[[]] where SC(v i , P i ) represents the semantic coherence corresponding to vertex v i in partition p i .
[0087] A distributed storage scheme for large - scale graph data based on Hbase to improve the efficiency of data storage and query. In the static partitioning stage, by calculating the similarity of classes in the ontology structure, a partitioning scheme that maximizes semantic coherence is sought.
[0088] To maintain the quality of graph data partitioning, this application performs dynamic partitioning of graph data based on the query load, adjusts the vertex partition by analyzing the load information of historical queries, so that while the load within the nodes is balanced, the communication overhead between nodes can be reduced. After the static partitioning of the graph, dynamic partitioning of the graph is also required. The graph dynamic partitioning scheme is: optimize the partitioning result by determining the load transfer amount and the target partition for vertex movement, calculate the signature of each vertex, and the signature is composed of vertex information and the structural information adjacent to the vertex. Divide the data graph G of the RDF ontology data into star - shaped sub - structures through the signature, and prune the graph data matching process according to the graph structure information provided by the vertex signature.
[0089] Among them, the process of determining the load transfer amount is: assume that the current graph query task is q. If node p i is overloaded in query q, that is, there are many active vertices participating in the query in node p i , which causes a sharp increase in the number of edges between vertices within the node. Set the load transfer amount function It is shown that to achieve load balancing for the current node, the amount of load to be transferred before the next query task (q + 1) is where is the difference in workload between node p i and node p j in the previous query task (q - 1), and k is a constant. At the same time, to prevent a large amount of load movement in a short period of time, the threshold α is calculated as: to limit the amount of workload movement in the query task. |V| + |E| is the sum of the number of vertices and edges, and k is a constant.
[0090] The process of determining the target partition for vertex movement is as follows: Set vertex v to belong to partition p i , and the number of active incoming edges of active vertex v spanning partition p j in query q is The number of active outgoing edges spanning partition p j is The communication score is the sum of the number of active incoming edges and the number of active outgoing edges. When the communication score is the largest after calculating that the vertex moves to node p j , the benefit of moving the vertex is the highest. Define the vertex movement benefit B(v): Calculate the communication scores of the vertex to be moved with all partitions, and select the maximum value as the candidate target partition p j for the vertex to move. If it is found during the process of solving the target partition for vertex v that the number of external edges is much smaller than the number of internal edges, that is the value of is close to |p j |, it indicates that vertex v is not suitable to move to p j , and it should continue to be retained in p i . In addition, through the communication overhead of the communication overhead between nodes and the load balancing within nodes can be comprehensively considered.
[0091] Calculate the signature of each vertex. The data graph G of the RDF ontology data is divided into star-shaped substructures through the signature. According to the graph structure information provided by the vertex signature, pruning the graph data matching process is the index construction process because the search space in the subgraph matching process is C(ν i ) is the candidate set corresponding to vertex v i in the data graph for the query graph. Then the worst-case space complexity is Θ(m n), where n is the space to be matched for vertices in the query graph, and m is the space already matched for vertices in the data graph. To reduce the space to be matched for vertices in the query graph, many feature-based algorithms build graph index structures by decomposing the data graph into structures such as subgraphs and subtrees. However, this method will split a large-scale data graph into a large number of subgraph structures.
[0092] To solve this problem, first represent the vertex signature as sign(v) = {vertex(v), pre(v), suc(v)}, where vertex(v) is the identification code of the vertex vertex(v) = {ID, din, dout}, which consists of the vertex code, in-degree, and out-degree. ID represents the code corresponding to vertex v, din represents the in-degree of the vertex, and dout represents the out-degree of the vertex. The in-edge set The out-edge set p represents the code of the adjacent edge, and ID represents the code of the adjacent vertex.
[0093] In a possible implementation, to improve the efficiency of finding the index in the matching phase, store the graph index through a B+ tree. A B+ tree is a multi-way tree. The leaf nodes store the actual values of the data and pointers to the data. The leaf nodes are stored on the disk in order for efficient sequential access. Build the index tree in the ascending order of the vertex codes. Since the data in the B+ tree nodes is arranged and connected in order, finding the data only requires comparing layer by layer downward until the matching leaf node is found.
[0094] In the process of subgraph matching, since the time complexity of subgraph query is exponential and the search space involved in the traversal process is extremely large, effective pruning rules must be designed to reduce the number of matching structures in the matching set through pruning and avoid unnecessary search traversals. In the embodiments of the present application, by introducing features in multiple dimensions such as vertex type and vertex degree into the pruning judgment rules, the search space is reduced from multiple aspects, the verification times in the matching process are reduced, and the efficiency of subgraph matching is improved. According to the timing of the pruning operation, divide the pruning rules into preliminary pruning and fine-grained pruning, and implement pruning before and during the search process respectively.
[0095] The first rule of preliminary pruning is pruning based on vertex type. The type of a vertex, as an important indicator for subgraph matching, can filter out the set of vertices with matching types before the matching starts. First, obtain all the types of vertices in the query graph Q, denoted as Class(Q). Then, compare each type in Class(Q) with each vertex in the data graph G, and remove the vertices in graph G that do not belong to Class(Q) and all their adjacent edges. The second rule is pruning based on vertex degree. The degree of a vertex, as a vertex feature of the query graph, can filter out vertices whose degrees cannot be matched. First, traverse all vertices in the query graph Q, and denote the minimum degree of the query graph Q as degree min (Q). Then, traverse all vertices in the data graph G, and delete all vertices that satisfy degree(v) < degree min (Q) and their adjacent edges.
[0096] Convert the query graph into a weighted graph, assign corresponding weights to each edge. The magnitude of the weight represents the frequency of this edge appearing in the data graph. Edges with high weights appear more frequently in the data graph. If these edges are used to filter the candidate set during subgraph matching, the search space will be very large. While edges with low weights appear less frequently in the data graph, a smaller-scale candidate set can be generated. Therefore, when selecting the matching order of edges, a strategy of preferentially matching low-weight edges is needed to quickly reduce the search space.
[0097] The preliminary pruning of vertices reduces the scale of the search space. It is necessary to build an index based on vertex signatures for the pruned data graph and query graph, and further prune during the search process according to the index. The subgraph matching scheme is based on the strategy of "pruning + verification". Therefore, during the vertex matching stage, it is necessary to verify the subgraphs that may be matched in the candidate set. The vertex matching rule is the basis for judging whether vertices in two graphs can be matched. By introducing fine-grained pruning into the vertex matching rule, the final matching and mapping of vertices are completed.
[0098] The first rule is the matching rule based on vertex type, which requires that the types of vertices are the same. Vertices in two graphs can be matched only when their types are exactly the same. The second is the matching rule based on vertex degree, which is the most common vertex matching rule. Degree refers to the number of edges connected to a vertex, and its magnitude can measure the importance and influence of this vertex in the entire graph. The third is the matching rule based on vertex adjacency structure, which requires that when vertices in two graphs are matched, the structures around them must be similar.
[0099] This application constructs an index based on vertex signatures, prunes the matching space from both the vertex's own information and neighborhood structure. The algorithm optimizes the generation order of the candidate set for the subgraph matching algorithm, preferentially screening vertices with strong filtering capabilities, and improving the query efficiency of graph data.
[0100] To prove the semantic partitioning of the embodiments of this application, Hash with the lowest algorithmic complexity and METIS with the highest maturity are selected for comparison. Through experiments conducted on three scales of datasets in LUBM, analyze the variation of the number of edge cuts of different algorithms under datasets of different scales. Please refer to Figure 2 , Figure 2 which shows the comparison chart of the number of cut edges of different partitioning algorithms proposed in the embodiments of this application. Among the three algorithms, METIS has the lowest number of edge cuts, Hash partitioning generates the most number of edge cuts, and the semantic partitioning proposed in this application is between the two. The semantic partitioning does not strictly partition the graph data in the way of minimizing the number of edge cuts. Based on Figure 2 Please refer to Figure 3 , Figure 3 which shows the comparison chart of the semantic consistency of different partitioning algorithms proposed in the embodiments of this application. It can be seen that the semantic partitioning algorithm has the highest semantic coherence. Therefore, the semantic partitioning algorithm sacrifices part of the edge cut performance and improves the semantic coherence.
[0101] In addition, to verify the efficiency of the graph storage method proposed in this application, this application conducts query experiments on the dataset LUBM1000. The query statements come from 14 test statements provided by LUBM. In the experiment, Q1, Q8, and Q9 are selected as test statements, and query tests are respectively conducted on the graph partitioning results of Hash, METIS, and the semantic partitioning algorithm. The average value of multiple queries is taken as the query time. The test results are as Figure 4 shown. Figure 4 which shows the schematic diagram of the query time of different partitioning algorithms proposed in the embodiments of this application. The experiment shows that for simple queries, the query times of the three partitioning algorithms are almost equal. As the complexity of the query statement increases, the graph storage method proposed in this application is slightly better than METIS and has the best query efficiency.
[0102] Therefore, since this application takes into account the efficiency of the algorithm and the partitioning effect in a balanced way, although the method of this application does not perform outstandingly in terms of the edge cut rate compared with the Metis method, the load balancing of the partitioning algorithm has been significantly improved, and for RDF data, considering the semantic information of the data during partitioning can provide effective support for subsequent upper-layer semantic queries, and still has great advantages.
[0103] Compared with the prior art, the embodiments of this application have the following beneficial effects:
[0104] First, the graph data compression and encoding method based on ontology represents the relationships between ontology concepts with a tree structure, constructs a relationship model of classes and attributes, and embeds the semantic information of RDF data into the ontology.
[0105] Second, a graph data partitioning method based on semantic information and query load, which initially partitions by maximizing semantic coherence and dynamically adjusts the graph data partitions based on the query load.
[0106] Third, an index construction method based on vertex signatures, which calculates the signatures of each vertex in the data graph with the vertex as the core, and the signature is jointly composed of the information of the vertex itself and the structural information adjacent to the vertex.
[0107] Fourth, a subgraph matching method based on index pruning, which adopts the idea of backtracking pruning, uses vertex signatures to prune candidate vertices, and recursively solves the subgraph matching problem.
[0108] The following gives a possible implementation of a graph data distributed storage device, which is used to execute each execution step and corresponding technical effect of the graph data distributed storage method shown in the above embodiments and possible implementations. The graph data distributed storage device includes:
[0109] A construction module, which is used to perform Jena parsing on RDF ontology data, generate a relationship tree according to the relationships between classes, and construct a relationship model;
[0110] A similarity generation module, which is used to define the distance similarity, structural similarity, and attribute similarity between vertices according to the hierarchical structure of the relationship tree in the relationship model;
[0111] A semantic calculation module, which is used to calculate the semantic similarity between vertices and obtain the corresponding semantic coherence through the semantic similarity. The semantic similarity is a weighted value of the distance similarity, structural similarity, and attribute similarity;
[0112] A generation module, which is used to maximize the division of the semantic coherence by using a graph partitioning algorithm to obtain a division result.
[0113] This preferred embodiment provides a computer device, which can implement the steps in any embodiment of the graph data distributed storage method provided by the embodiments of the present application. Therefore, the beneficial effects of the graph data distributed storage method provided by the embodiments of the present application can be achieved. For details, see the previous embodiments and will not be repeated here.
[0114] Those of ordinary skill in the art can understand that all or part of the steps in the above various methods can be completed by instructions or by controlling relevant hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. For this purpose, the embodiments of the present application provide a storage medium, in which multiple instructions are stored, and the instructions can be loaded by a processor to execute the steps in any embodiment of the graph data distributed storage method provided by the embodiments of the present application.
[0115] Among them, the storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0116] Since the instructions stored in the storage medium can execute the steps in any of the embodiments of the graph data distributed storage method provided in the embodiments of the present application, the beneficial effects achievable by any of the graph data distributed storage methods provided in the embodiments of the present application can be realized. For details, refer to the previous embodiments and will not be elaborated here.
[0117] The above are only the preferred embodiments of the present application and are not intended to limit the present application. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present application shall be included in the protection scope of the present application.
Claims
1. A method for distributed storage of graph data, characterized in that, including: Performing Jena parsing on RDF ontology data, generating a relationship tree according to the relationships between classes, and constructing a relationship model; Defining distance similarity, structural similarity, and attribute similarity between vertices according to the hierarchical structure of the relationship tree in the relationship model; Calculating the semantic similarity between vertices, and obtaining the corresponding semantic coherence through the semantic similarity, where the semantic similarity is a weighted value of the distance similarity, the structural similarity, and the attribute similarity; Using a graph partitioning algorithm to perform a maximization partition on the semantic coherence to obtain a partition result; Optimizing the partition result by determining the load transfer amount and the vertex movement target partition; Calculating the signature of each vertex, where the signature is composed of vertex information and the structural information adjacent to the vertex; Partition the data graph of the RDF ontology data through the signature into star-shaped substructures, and prune the graph data matching process according to the graph structure information provided by the vertex signature.
2. The method for distributed storage of graph data according to claim 1, characterized in that, The RDF ontology data includes a plurality of triple items, and the step of performing Jena parsing on the RDF ontology data includes: Grouping the subjects and objects of the plurality of triple items by class to obtain the classes to which the plurality of triples belong; Assigning class codes with each class to which a triple belongs as an identifier, and using the hash value of the URI prefix of each triple item as a prefix code; Obtaining a suffix code for each triple item according to the class to which the triple item belongs and the sequential number of the prefix code; Obtaining a triple item code according to the class code, the prefix code, and the suffix code; 3. The method for distributed storage of graph data according to claim 2, characterized in that, The class encoding has the following calculation formula: , where is the class marker bit, is the number of direct parent classes, represents the parent class node of node , and is the node sequence encoding of node represents the th layer and the th node The sequence encoding of is expressed as: , where is the sequence encoding of the class node in the th layer.
4. The method for distributed storage of graph data according to claim 1, characterized in that, The distance similarity is defined as: , represents the class to which vertex belongs, represents the class to which vertex belongs, and respectively represent the path lengths from to the least common superclass, represents the path length from to the root node, and represents the least common superclass of The structural similarity is as follows: , is the semantic feature information quantity of class , is the semantic feature information quantity of class , is the semantic feature information quantity; The described attribute similarity is as follows: ; Among them, is a function regarding the influence magnitude of and on the attribute similarity, and correspond to the number of attributes of and respectively.
5. The method for distributed storage of graph data according to claim 1, characterized in that, The step of pruning the graph data matching process includes: Obtain query graph All types of vertices in ; Compare each type in with each vertex in the data graph . Remove from the data graph the vertices that do not belong to and all the edges adjacent to them.
6. The method for distributed storage of graph data according to claim 5, characterized in that, The step of pruning the graph data matching process further includes: Traverse the query graph for all vertices, and denote the minimum degree of the query graph as ; Traverse the data graph for all vertices, and delete all vertices that satisfy the vertex degree and the adjacent edges.
7. A device for distributed storage of graph data, characterized in that, The device includes: A construction module for performing Jena parsing on RDF ontology data, generating a relationship tree according to the relationships between classes, and constructing a relationship model; A similarity generation module for defining distance similarity, structural similarity, and attribute similarity between vertices according to the hierarchical structure of the relationship tree in the relationship model; A semantic calculation module for calculating the semantic similarity between vertices, and obtaining the corresponding semantic coherence through the semantic similarity, where the semantic similarity is a weighted value of the distance similarity, the structural similarity, and the attribute similarity; A generation module for using a graph partitioning algorithm to perform a maximization partition on the semantic coherence to obtain a partition result; An optimization module for optimizing the partition result by determining the load transfer amount and the vertex movement target partition; A calculation module for calculating the signature of each vertex, where the signature is composed of vertex information and the structural information adjacent to the vertex; The splitting and pruning module is used to split the data graph of the RDF ontology data into star-shaped sub-structures through the signature, and prune the graph data matching process according to the graph structure information provided by the vertex signature. Split into star-shaped sub-structures, and prune the graph data matching process according to the graph structure information provided by the vertex signature.
8. A computer device, characterized in that, The computer device includes a processor and a memory, and a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the graph data distributed storage method according to any one of claims 1-6; 9. A computer-readable storage medium, characterized in that, A computer program is stored in the storage medium, and the computer program is loaded and executed by a processor to implement the graph data distributed storage method according to any one of claims 1-6.
Citation Information
Patent Citations
Storage optimization-based distributed graph processing method
CN107122248A
Semantic search implementation method and system, computer equipment and storage medium
CN114490928A