A Large-Scale Semantic Graph Approximate Summarization Method and System Based on a Partial Order Lattice
By classifying the relationship type of large semantic graphs and calculating the approximate abstract using partial order grid method, the problems of poor readability and high complexity of semantic graph abstracts in the prior art are solved, and an efficient and concise semantic graph abstract is achieved.
Patent Information
- Application Number
- CN202210049687.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-01-17
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2042-01-17
AI Technical Summary
The existing semantic graph summary method cannot effectively process the features of different semantic graphs, resulting in poor readability and simplicity of the summary, and high complexity, making it difficult to apply to large semantic graphs.
By classifying the relationship types of large semantic graphs, the approximate summary is calculated using a partially ordered grid-based method, and using algorithm 1 and algorithm 2 to process rich and simple relationship-type semantic graphs respectively, and the information degree of the abstract is calculated to evaluate the coverage and filter ratio of the abstract.
It improves the efficiency and readability of semantic graph abstracts, can effectively support users to understand and browse large semantic graphs, and reduces computational complexity.
Smart Images

Figure CN114385807B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of computers, and relates to a method and system for approximately summarizing large semantic graphs based on a partially ordered lattice. Background Art
[0002] A semantic graph, that is, a graph structure formed by semantic data, is applied in many fields, including medical, education, e-commerce, and agriculture, etc. Currently, semantic data has grown explosively. For example, semantic data from different fields such as geography, bioscience, lexicostatistics, linguistics, and sociology. The Linkgeodata dataset in the geography field alone contains more than 20 billion triples and 3 billion node data; the Semantic Web Linked Open Data Cloud (LOD) has more than 6.3 million different large datasets, and the linked datasets include AGROVOC, DBpedia, and wikidata, etc. Due to the continuous growth of semantic data, it has become extremely difficult to understand and use large semantic graphs.
[0003] Semantic graph summarization aims to reduce the scale of the semantic graph by extracting or compressing the data in the original semantic graph, so as to solve the above-mentioned semantic graph application problems. The existing semantic graph summarization mainly focuses on: (1) statistical methods, that is, extracting important nodes of the semantic graph through various methods of calculating central nodes to form a summary; (2) pattern mining, that is, mining frequent subgraphs of the semantic graph, and using the set of subgraphs as the summary of the original semantic graph; (3) graph structure equivalence relations, and using the equivalence relations between nodes to form a quotient graph as the summary of the original semantic graph.
[0004] The main defects of the above methods are in three aspects: First, the characteristics of different semantic graphs are not considered, and a single strategy is used to summarize the semantic graph without discrimination; second, the readability and conciseness of the summaries generated for large semantic graphs are poor, and they cannot support users to understand and browse; third, the complexity of some summary methods is too high to be applied to large semantic graphs. Summary of the Invention
[0005] In view of this, the purpose of the present invention is to provide a method and system for approximately summarizing large semantic graphs based on a partially ordered lattice, so as to solve the problems that the existing semantic graphs are too large for users to effectively browse, understand, and use, etc.
[0006] To achieve the above purpose, the present invention provides the following technical solutions:
[0007] A method for approximately summarizing large semantic graphs based on a partially ordered lattice, the method comprising the following steps:
[0008] S1: Classify large semantic graphs according to the richness of relationship types into: Type I, i.e., rich-relationship semantic graphs, and Type II, i.e., simple-relationship semantic graphs;
[0009] S2: For Type I semantic graphs, use Algorithm 1 to calculate the approximate summary based on the poset lattice according to its characteristics, and then use Algorithm 3 to calculate the information degree of the summary, that is: the ratio of covering the original semantic graph;
[0010] S3: For Type II semantic graphs, use Algorithm 2 to calculate the approximate summary based on the poset lattice according to its characteristics, and then use Algorithm 4 to calculate the information degree of the summary, that is: the filtering ratio of the entities in the original semantic graph;
[0011] S4: Generate the poset lattice summary result of the semantic graph.
[0012] Optionally, the specific content of S1 is as follows:
[0013] A semantic graph is composed of semantic data RDF triples, and the semantic graph is defined as where V is the set of entities, R is the set of relationships between entities, is the relationship type, that is, the set of object properties, is the property, that is, the set of data type properties, is the mapping from the relationship to the relationship type, is the mapping from the entity to the property set; the properties of the entities in the semantic graph are regarded as the properties that only associate with the entity, rather than the relationship between the entity and the property value;
[0014] Define the relationship type index δ:
[0015]
[0016] to measure the richness of the relationships in the semantic graph; where the larger δ is, the richer the relationship types of the semantic graph are; conversely, the simpler the relationship types are;
[0017] The steps of semantic graph classification are as follows:
[0018] S11: First, extract the number of entities |V| and the number of relationship types of the large semantic graph by parsing the RDF file of the semantic graph or importing the semantic graph into the corresponding database, including the graph database and the semantic database, and obtaining it using the database query language;
[0019] S12: Second, calculate the relationship index δ according to formula (1);
[0020] S13: Compare the size relationship between the relationship index δ and the set index threshold δ T ; According to the situation of existing large semantic graphs, set the default value of δ T to 10-4 ; The user sets according to the specific situation of the processed semantic graph;
[0021] S14: Finally, according to the magnitudes of δ and δ T , obtain the semantic graph type: when δ < δ T , the semantic graph is a type-I semantic graph; when δ ≥ δ T , the semantic graph is a type-II semantic graph.
[0022] Optionally, the specific content of S2 is as follows:
[0023] Define 1 entity pattern: Given a semantic graph G, let be the set of all subjects s in all triples (s, p, o) in the entity; for any be the feature set of entity s; an entity pattern (EP) is defined as c = (S, T, A), where: (i) (ii) CS(s) = T; (iii) A = ∪ s∈S L A (s);
[0024] Let C be the set of all entity patterns, then forms a poset; if two special entity patterns and are set, then forms a poset lattice;
[0025] Define 2 key relationship types: Given a semantic graph G, if the subset of relationship types: is the top σ% of the relationship types that the semantic graph is retrieved most frequently, where then R t * is called the set of key relationship types, and the elements in R t * are key relationship types;
[0026] Set the value of σ to 20;
[0027] Define 3 approximate summary of type-I semantic graph based on poset lattice: Given a semantic graph G and the set of key relationship types R t *, the approximate summary of type-I semantic graph based on poset lattice is defined as the lattice σL formed by the poset (σC, ≤), where σC is the set of entity patterns and each entity pattern contains at least one key relationship type, that is:
[0028] Algorithm 1 gives the steps to calculate the approximate summary ELSRR of type-I semantic graph based on poset lattice; the input of this algorithm is semantic graph G, the set of key types R t*, parameter σ and semantic graph type, the output is the type-I semantic graph approximate summary σL based on the poset lattice;
[0029] S21: Initialize the entity pattern set;
[0030] S22: For each entity s in each semantic graph, if it is associated with a key relationship type, add the entity s and all its associated relationship types to σC;
[0031] S23: Merge entities with the same feature set CS, and layer the entity patterns EP according to the cardinality of the feature set CS; CS_T k Store the entity patterns EP of the k-th layer, that is: all entity patterns EP in the k-th layer satisfy: the cardinality |T| of the feature set of all entities = k; m represents the maximum value of all feature sets CS;
[0032] S24: Generate the poset lattice σL according to the entity patterns CS_T of each layer;
[0033] S25: Return the poset lattice σL.
[0034] Optionally, the specific content of S3 is as follows:
[0035] Definition 4 Type-II semantic graph approximate summary based on the poset lattice: Given a semantic graph G and a set of key relationship types The type-II semantic graph approximate summary based on the poset lattice is defined as the lattice μL formed by the poset (μC, ≤), where: There is |E p* | ≥ μ(p*), E p* Is the edge set with the relationship type p*, and μ(p*) is the threshold of p*;
[0036] Set μ(p*) = 2 to filter at least 50% of the entities related to p*; μ(p*) is set by the user himself, and different thresholds are set for different key relationship types p* to achieve filtering of the specified entities;
[0037] Algorithm 2 gives the steps to calculate the type-II semantic graph approximate summary ELSSR based on the poset lattice; the input of this algorithm is the semantic graph G, the set of key types R t *, the threshold μ(p*) of p*, and the semantic graph type, and the output is the type-II semantic graph approximate summary μL based on the poset lattice;
[0038] S31: Initialize the entity pattern set;
[0039] S32: For each entity s in each semantic graph, if the entity s is associated with the key relationship type p*, then check its associated corresponding edge set |E p*|Relationship with the set threshold μ(p*). If |E p* |≥μ(p*), then add the entity s and all its associated relationship types to μC;
[0040] S33: Merge entities with the same feature set CS and layer the entity patterns EP according to the cardinality of the feature set CS; CS_T k Store the entity patterns EP of the k-th layer, i.e., all entity patterns EP in the k-th layer satisfy that the cardinality |T| of the feature set of all entities = k; m represents the maximum value of all feature sets CS;
[0041] S34: Generate the poset μL according to the entity patterns CS_T of each layer;
[0042] S35: Return the poset μL.
[0043] Optionally, the specific content of S4 is as follows:
[0044] Define the base graph of 5ELSRR: Given the semantic graph The set of key relationship types R t *, and the ELSRR summary σL = (σC, ≤) of this semantic graph, the base graph of G is defined as: Is a subgraph of the semantic graph G that satisfies:
[0045] (1) V b = V σ ∪V N , where V N Contains the adjacent nodes of all nodes in V σ ;
[0046] (2) R b = {(u, v)|u ∈ V σ or v ∈ V σ};
[0047] (3)
[0048] (4)
[0049] (5) Is a mapping that maps the relationships in R b to the relationship types in the semantic graph;
[0050] (6) Is a mapping that maps the entities in V b to the attribute set in the semantic graph;
[0051] The base graph of ELSRR is the subgraph of the original semantic graph covered by the summary;
[0052] Define the information degree of 6ELSRR: Given a semantic graph The set R of key relationship types t *, and the ELSRR summary σL = (σC, ≤) of this semantic graph, the information degree of ELSRR is defined as:
[0053]
[0054] where V b and R b are the entity set and relationship set of the base graph, and V and R are the entity set and relationship set of the semantic graph;
[0055] Algorithm 3 is the calculation method of the ELSRR information degree; the specific steps are as follows:
[0056] S41: Initialize the corresponding variable I σ , V b , V σ , V N , R b ;
[0057] S42: Calculate the base graph G of σL b ;
[0058] S43: Calculate the information degree I according to formula (2) σ ;
[0059] S44: Return the information degree I σ ;
[0060] Define the information degree of 7ELSSR: Given a semantic graph The set R of key relationship types t * and its threshold μ(R t *), the ELSRR summary μL = (μC, ≤) of this semantic graph, the information degree of ELSRR is defined as:
[0061]
[0062] where V μ is the set of all entities included in μC;
[0063] Algorithm 4 is the calculation method of the ELSRR information degree; the specific steps are as follows:
[0064] S51: Initialize the corresponding variable I μ , V b , V σ , V N , R b ;
[0065] S52: Calculate the number of entities of μL
[0066] S53: Calculate the information degree I according to formula (3) μ ;
[0067] S54: Return the information degree I μ 。
[0068] A large-scale semantic graph approximate summary system based on a partial order lattice, including a memory, a processor, and a computer program stored on the memory and capable of running on the processor. When the processor executes the computer program, the method described above is implemented.
[0069] A computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a processor, the method described above is implemented.
[0070] The beneficial effects of the present invention are as follows:
[0071] (1) Classify the semantic graph according to the relationship type index, and adopt different summary strategies for different semantic graph types;
[0072] (2) Through the key relationship type, use the partial order lattice to perform approximate summarization on the semantic graph, which greatly improves the efficiency of the summarization method, and the form of its summary is not only concise but also highly readable, facilitating users to understand and browse large-scale semantic graphs.
[0073] Other advantages, objectives, and features of the present invention will be described to some extent in the subsequent description, and to some extent, they will be obvious to those skilled in the art based on the study of the following text, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be achieved and obtained through the following description. Description of the Drawings
[0074] In order to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be described in detail preferably with reference to the accompanying drawings, where:
[0075] Figure 1 is the flowchart of the method of the present invention;
[0076] Figure 2 is the flowchart of semantic graph classification;
[0077] Figure 3 is an example type I (rich relationship type) semantic graph;
[0078] Figure 4 is the approximate summary of type I semantic graph based on the partial order lattice;
[0079] Figure 5 is the partial order lattice approximate summary of the semantic graph YAGO;
[0080] Figure 6 For Figure 3 the base graph of the ELSRR summary in Detailed implementation manners
[0081] The following uses specific specific examples to illustrate the implementation manners of the present invention. Those skilled in the art can easily understand other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific implementation manners. Various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the diagrams provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner. Without conflict, the following embodiments and the features in the embodiments can be combined with each other.
[0082] Among them, the attached drawings are only for illustrative purposes, showing only schematic diagrams, not physical diagrams, and cannot be understood as a limitation to the present invention; in order to better illustrate the embodiments of the present invention, some components in the attached drawings will be omitted, enlarged or reduced, and do not represent the dimensions of the actual product; for those skilled in the art, it is understandable that some well-known structures and their descriptions in the attached drawings may be omitted.
[0083] In the attached drawings of the embodiments of the present invention, the same or similar reference numerals correspond to the same or similar components; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "rear", etc. indicating the orientation or positional relationship, it is based on the orientation or positional relationship shown in the attached drawings. It is only for the convenience of describing the present invention and simplifying the description, rather than indicating or implying that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation. Therefore, the terms describing the positional relationship in the attached drawings are only for illustrative purposes and cannot be understood as a limitation to the present invention. For those of ordinary skill in the art, the specific meanings of the above terms can be understood according to specific circumstances.
[0084] Figure 1 is the process included in the large-scale semantic graph approximate summary method based on the poset lattice. First, the large-scale semantic graph is classified according to the richness of the relationship types, and is divided into: type I (rich relationship type) semantic graph and type II (simple relationship type) semantic graph. For the type I (rich relationship type) semantic graph, use algorithm 1 to calculate the approximate summary based on the poset lattice according to its characteristics, and then use algorithm 3 to calculate the information degree of the summary, that is: the ratio of covering the original semantic graph. For the type II (simple relationship type) semantic graph, use algorithm 2 to calculate the approximate summary based on the poset lattice according to its characteristics, and then use algorithm 4 to calculate the information degree of the summary, that is: the filtering ratio of the entities in the original semantic graph. Finally, generate the poset lattice summary result of the semantic graph.
[0085] (1) Semantic graph classification
[0086] The semantic graph is composed of semantic data RDF triples. The present invention defines the semantic graph as where V is a set of entities, R is a set of relationships between entities, is a set of relationship types (i.e., object properties), is a set of properties (i.e., datatype properties), is a mapping from relationships to relationship types, is a mapping from entities to sets of properties. Different from the definition of common semantic graphs, the present invention regards the properties of entities in the semantic graph as the properties only associated with the entity, rather than the relationship between the entity and the property value. This way of defining the semantic graph greatly reduces the number of edges in the semantic graph and can effectively reduce the computational complexity of large semantic graphs.
[0087] Table 1 lists 3 semantic graphs and related information. Among them, the number of entities and relationships in the YAGO and DBpedia semantic graphs are in the millions or tens of millions, which are large semantic graphs, but the number of relationship types is relatively small. While the relationship types of ODU are relatively rich.
[0088] Table 1 Example semantic graphs and related metrics
[0089]
[0090] The present invention defines the relationship type index δ:
[0091]
[0092] to measure the richness of relationships in the semantic graph. Among them, the larger δ is, the richer the relationship types in the semantic graph; on the contrary, the relationship types are simpler. For YAGO, δ = 6 / 4,103,888 = 1.46e-6. This indicates that the relationships in this semantic graph are simple. A similar conclusion can be drawn for DBpedia. However, the δ index value of ODUS is 1.66e-3, showing relatively rich relationship types.
[0093] Figure 2 Shows the classification steps of the semantic graph, specifically as follows:
[0094] Step 1: First, extract the number of entities |V| and the number of relationship types of the large semantic graph This process can be completed by parsing the RDF file of the semantic graph, or the semantic graph can be imported into the corresponding database (graph database or semantic database), and obtained using the database query language.
[0095] Step 2: Secondly, calculate the relationship index δ according to formula (1).
[0096] Step 3: Compare the size relationship between the relationship index δ and the set index threshold δ T According to the situation of the existing large semantic graph, the present invention sets the default value of δ T to 10 -4 . Users can set it according to the specific situation of the semantic graph being processed.
[0097] Step 4: Finally, based on the size of δ and δ T , obtain the semantic graph type: when δ < δ T , the semantic graph is a type I (rich relationship type) semantic graph; when δ ≥ δ T , the semantic graph is a type II (simple relationship type) semantic graph.
[0098] As can be seen from the above classification method, among the 3 semantic graphs shown in Table 1, YAGO and DBpedia are type II (simple relationship type) semantic graphs, while ODUS is a type I (rich relationship type) semantic graph.
[0099] (2) Approximate summary of type I (rich relationship type) semantic graph based on poset lattice
[0100] Definition 1 Entity Pattern (EP) Given a semantic graph G, let be the set of all subjects s in all triples (s, p, o) in the entity. For any is the feature set of entity s. An entity pattern (EP) is defined as c = (S, T, A), where: (i) (ii) CS(s) = T; (iii) A = ∪ s∈ S L A (s).
[0101] Let C be the set of all entity patterns, then forms a poset. If two special entity patterns and are set, then forms a poset lattice.
[0102] Definition 2 Critical relation type Given a semantic graph G, if the subset of relation types: is the top σ% of the relation types that are most frequently retrieved in this semantic graph, where then R t * is called the critical relation type set, and the elements in R t * are critical relation types.
[0103] According to the report in the literature (Peng P, Zou L, Chen L and Zhao D. Adaptive Distributed RDF Graph Fragmentation and Allocation based on Query Workload. in IEEE Transactions on Knowledge and Data Engineering, 31(4)(2019):670 - 685.), in semantic graphs such as DBpedia, 90% of the queries contain only 20% of the relation types. Therefore, the present invention sets the default σ value to 20.
[0104] Definition 3: Essential Lattice based Summary for RDFG with Rich Relations (ELSRR) Given a semantic graph G and a set of key relation types R t *, the essential lattice based summary of type I for the semantic graph is defined as the lattice σL formed by the poset (σC, ≤), where σC is the set of entity schemas and each entity schema contains at least one key relation type, that is:
[0105] Algorithm 1 gives the steps to calculate the essential lattice based summary of type I for the semantic graph ELSRR. The input of this algorithm is the semantic graph G, the set of key types R t *, the parameter σ and the semantic graph type, and the output is the essential lattice based summary of type I for the semantic graph σL.
[0106] Step 1: In line 1), initialize the set of entity schemas.
[0107] Step 2: In lines 2) - 5), for each entity s in the semantic graph, if it is associated with a key relation type, then add the entity s and all its associated relation types to σC.
[0108] Step 3: In line 6), merge the entities with the same feature set CS, and layer the entity patterns EP according to the cardinality of the feature set CS. CS_T k Stores the entity patterns EP of the k - th layer, that is: all entity patterns EP in the k - th layer satisfy: the cardinality of the feature set of all entities |T| = k. m represents the maximum value of all feature sets CS.
[0109] Step 4: In line 7), generate the poset lattice σL according to the entity patterns CS_T of each layer.
[0110] Step 5: In line 8), return the poset lattice σL.
[0111]
[0112] Example 1 Figure 3 shows a semantic graph, where ( is the key relationship type). Table 2 lists the entity patterns: c 0 , c 1 , c 2 , c 3 , c 4 , c 5 , c 6 , c 7 . Figure 4 is the ELSRR summary of this semantic graph.
[0113] Table 2 Entity Patterns of Example 1
[0114]
[0115] (3) Type-II (Simple Relationship Type) Semantic Graph Approximate Summary Based on Poset Lattice
[0116] Definition 4 Type-II Semantic Graph Approximate Summary Based on Poset Lattice (Essential Lattice based Summary for RDFG with Simple Relations, ELSSR) Given a semantic graph G and a set of key relationship types The type-II semantic graph approximate summary based on poset lattice is defined as the lattice μL formed by the poset (μC, ≤), where: There is |E p* | ≥ μ(p*), where E p* is the set of edges with the relationship type p*, and μ(p*) is the threshold of p*.
[0117] According to the report of the literature (Kumar R, Raghavan P, Rajagopalan S, Tomkins A. Trawling the Web for emerging cyber-communities, Computer Networks, 31(1999):1481-1493.), in real large graph structures, the node degrees generally follow a power-law distribution, that is: the probability that a node has at least k out-degrees is Therefore, 50% of the nodes have at least 2 out-degrees. In the present invention, μ(p*) = 2 is set, then at least 50% of the entities related to p* can be filtered. μ(p*) can also be set by the user himself, and different thresholds can be set for different key relationship types p* to achieve the filtering of specified entities.
[0118]
[0119] Algorithm 2 gives the steps for calculating the type-II semantic graph approximate summary ELSSR of a partial order lattice. The input of this algorithm is the semantic graph G, the set of key types R t *, the threshold μ(p*) of p*, and the semantic graph type, and the output is the type-II semantic graph approximate summary μL based on the partial order lattice.
[0120] Step 1, line 1) initializes the entity pattern set.
[0121] Step 2, lines 2)-13) for each entity s in each semantic graph, if the entity s is associated with the key relationship type p*, then check the relationship between the corresponding edge set |E p* | it is associated with and the set threshold μ(p*). If |E p* | ≥ μ(p*), then add the entity s and all the relationship types it is associated with to μC.
[0122] Step 3, line 14) merges entities with the same feature set CS and hierarchically arranges the entity patterns EP according to the cardinality of the feature set CS. CS_T k stores the entity patterns EP of the k-th layer, that is: all entity patterns EP in the k-th layer satisfy that the cardinality |T| of the feature set of all entities is = k. m represents the maximum value of all feature sets CS.
[0123] Step 4, line 15) generates the partial order lattice μL according to the entity patterns CS_T of each layer.
[0124] Step 5, line 16) returns the partial order lattice μL.
[0125] Example 2 Figure 5 is the approximate summary based on the partial order lattice of the YAGO semantic graph. Table 3 lists the information of the corresponding 20 entity patterns. In this example, set R t * = {isCitizenOf} and μ(isCitizenOf) = 2, that is: for the relationship type "isCitizenOf", this summary extracts entities with dual nationality or more. Compared with the original semantic graph, the summary proposed by the present invention reduces the number of entities from 3,098,907 to 1,421,131, increasing the readability of the summary.
[0126] Note: (i) i-j represents the entity pattern in the i-th layer and the j-th column; (ii) "1" and "0" represent and
[0127] Table 3 Entity patterns of the YAGO semantic graph and the approximate summary based on the partial order lattice (ELSSR)
[0128]
[0129] (4) ELSRR Abstract Information Degree Calculation Method
[0130] The following gives the information degree calculation method of the type-I (rich in relationships) semantic graph approximate abstract (ELSRR) based on the poset lattice.
[0131] Definition 5: The base graph of ELSRR Given a semantic graph The set R of key relationship types t *, and the ELSRR abstract σL = (σC, ≤) of this semantic graph, the base graph of G is defined as: is a subgraph of the semantic graph G that satisfies: (1) where V N contains the adjacent nodes of all nodes in V σ ; (2) R b = {(u, v) | u ∈ V σ or v ∈ V σ}; (3) (4) (5) is a mapping that maps the relationships in R b to the relationship types in the semantic graph; (6) is a mapping that maps the entities in V b to the set of attributes in the semantic graph.
[0132] From the above definition, it can be seen that the base graph of ELSRR is the subgraph of the original semantic graph covered by this abstract.
[0133] Definition 6: The information degree of ELSRR Given a semantic graph The set R of key relationship types t *, and the ELSRR abstract σL = (σC, ≤) of this semantic graph, the information degree of ELSRR is defined as:
[0134]
[0135] where, V b and R b are the entity set and relationship set of the base graph, and V and R are the entity set and relationship set of the semantic graph.
[0136] Example 3 Figure 6 is the base graph of the ELSRR abstract of the semantic graph shown in Example 1, where the black nodes are the node set V Figure 3 of the base graph, and the gray nodes are the adjacent node set V σ of V σ . According to formula (2), N . Figure 4Information Degree of the Shown ELSRR Summary
[0137] It can be seen from this that Figure 4 the shown ELSRR summary covers 77% of the entities and relationships in the original semantic graph, including the main and most of the information in the original semantic graph.
[0138] Algorithm 3 is the calculation method of the ELSRR information degree. The specific steps are as follows:
[0139] Step 1: Initialize the corresponding variables I σ , V b , V σ , V N , R b ;
[0140] Step 2: Calculate the base graph G of σL in lines 2)-9) b ;
[0141] Step 3: Calculate the information degree I according to formula (2) in line 10 σ ;
[0142] Step 4: Return the information degree I in line 11 σ .
[0143]
[0144] (5) Calculation Method of the ELSSR Summary Information Degree
[0145] The following gives the calculation method of the information degree of the type-II (relationship-simple type) semantic graph approximate summary (ELSSR) based on the poset lattice. Since the number of relationship types in the type-II semantic graph is small, but the number of relationship instances (edges with a certain relationship type) is huge, the information degree of this summary ELSSR is used to measure the ratio of the "representative" entities it includes to the entities in the original semantic graph. The so-called "representative" entities refer to the entities that meet the threshold μ(p*) requirements of the key type p*.
[0146] Definition 7 Information Degree of ELSSR Given a semantic graph Set of key relationship types R t * and its threshold μ(R t *), for the ELSSR summary μL = (μC, ≤) of this semantic graph, the information degree of ELSSR is defined as:
[0147]
[0148] where V μ is the set of all entities included in μC.
[0149] Example 4 Table 3 shows the information of 21 entity patterns of YAGO, where R t * = {isCitizenOf} and μ(isCitizenOf) = 2. This ELSSR summary contains 1,421,131 entities. The number of entities in the original YAGA semantic graph is 3,098,907, so the information degree Therefore, by setting the key attribute threshold, the ELSSR summary filters 54% of the entities, retains the representative entities, and increases the readability of the summary.
[0150] Algorithm 4 is the ELSSR information degree calculation method. The specific steps are as follows:
[0151] Step 1: Initialize the corresponding variables I μ , V b , V σ , V N , R b ;
[0152] Step 2: Calculate the number of entities of μL in lines 2) - 4);
[0153] Step 3: Calculate the information degree I according to formula (3) in line 5 μ ;
[0154] Step 4: Return the information degree I in line 6 μ .
[0155]
[0156] It should be recognized that the embodiments of the present invention can be implemented or carried out by computer hardware, a combination of hardware and software, or by computer instructions stored in a non - transitory computer - readable memory. The method can be implemented in a computer program using standard programming techniques - including a non - transitory computer - readable storage medium configured with a computer program, where the storage medium so configured causes the computer to operate in a specific and predefined manner - according to the methods and drawings described in the specific embodiments. Each program can be implemented in a high - level procedural or object - oriented programming language to communicate with the computer system. However, if necessary, the program can be implemented in assembly or machine language. In any case, the language can be a compiled or interpreted language. In addition, for this purpose, the program can run on a dedicated integrated circuit programmed for this purpose.
[0157] In addition, the operations of the processes described herein can be performed in any suitable order, unless otherwise indicated herein or otherwise clearly contradicted by the context. The processes described herein (or variations and / or combinations thereof) can be executed under the control of one or more computer systems configured with executable instructions and can be implemented as code (e.g., executable instructions, one or more computer programs, or one or more applications) executed jointly on one or more processors, by hardware, or by a combination thereof. The computer program includes a plurality of instructions executable by one or more processors.
[0158] Further, the method can be implemented in any type of computing platform operably connected to a suitable one, including but not limited to personal computers, minicomputers, mainframes, workstations, network or distributed computing environments, separate or integrated computer platforms, or communicating with charged particle tools or other imaging devices, etc. Aspects of the present invention can be implemented in machine-readable code stored on a non-transitory storage medium or device, whether removable or integrated into the computing platform, such as a hard disk, optical read and / or write storage medium, RAM, ROM, etc., such that it can be read by a programmable computer and, when the storage medium or device is read by the computer, can be used to configure and operate the computer to execute the processes described herein. In addition, the machine-readable code, or portions thereof, can be transmitted via a wired or wireless network. When such media include instructions or programs that implement the above-described steps in combination with a microprocessor or other data processor, the inventions described herein include these and other different types of non-transitory computer-readable storage media. When programmed according to the large semantic graph approximate summarization method and technology based on a partially ordered lattice of the present invention, the present invention also includes the computer itself.
[0159] The computer program can be applied to the input data to perform the functions described herein, thereby transforming the input data to generate output data stored in non-volatile memory. The output information can also be applied to one or more output devices such as a display. In a preferred embodiment of the present invention, the transformed data represents physical and tangible objects, including a specific visual depiction of the physical and tangible objects generated on the display.
[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the present technical solution, and they should all be covered by the scope of the claims of the present invention.
Claims
1. A method for approximate summarization of large semantic graphs based on poset lattices, characterized in that: The method comprises the following steps: S1: Classify the large semantic graph according to the richness of relationship types into: Type I, i.e., rich relationship type semantic graph and Type II, i.e., simple relationship type semantic graph; The semantic graph classification step is specifically as follows: S11: First, extract the number of entities |V| and the number of relationship types in the large semantic graph Specifically: Parse the RDF file of the semantic graph, import the semantic graph into the corresponding databases, including the graph database and the semantic database, and obtain it using the database query language; S12: Secondly, calculate the relationship index δ according to formula (1); The semantic graph is composed of semantic data RDF triples, and the semantic graph is defined as where V is the set of entities, R is the set of relationships between entities, is the relationship type, that is, the set of object properties, is the property, that is, the set of data type properties, is the mapping from relationships to relationship types, is the mapping from entities to property sets; the properties of entities in the semantic graph are regarded as the properties that only associate with the entity, rather than the relationship between the entity and the property value; Define the relationship type index δ: to measure the richness of relationships in the semantic graph; where the larger δ is, the richer the relationship types in the semantic graph; conversely, the simpler the relationship types. S13: Compare the size relationship between the relationship index δ and the set index threshold δ T ; According to the situation of the existing large-scale semantic graph, set the default value of δ T to 10 -4 ; The user sets it according to the specific situation of the processed semantic graph S14: Finally, based on the magnitudes of δ and δ T , the semantic graph type is obtained: when δ < δ T , the semantic graph is a type-I semantic graph; when δ ≥ δ T , the semantic graph is a type-II semantic graph; S2: For Type I semantic graphs, use Algorithm 1 to calculate the approximate summary based on poset lattices according to their characteristics, and then use Algorithm 3 to calculate the information degree of the summary, i.e., the ratio of covering the original semantic graph; S3: For Type II semantic graphs, use Algorithm 2 to calculate the approximate summary based on poset lattices according to their characteristics, and then use Algorithm 4 to calculate the information degree of the summary, i.e., the filtering ratio of the entities in the original semantic graph; S4: Generate the poset lattice summary result of the semantic graph; Algorithm 1 gives the steps for calculating the type-I semantic graph approximate summary ELSRR of a partial order lattice; the input of this algorithm is the semantic graph G, the key type set R t *, the parameter σ and the semantic graph type, and the output is the type-I semantic graph approximate summary σL of the partial order lattice; S21: Initialize the entity pattern set; S22: For each entity s in the semantic graph, if it is associated with a key relationship type, add the entity s and all the relationship types it is associated with to σC; S23: Merge entities with the same feature set CS, and hierarchically classify entity patterns EP according to the cardinality of the feature set CS; CS_T k Store the entity pattern EP of the k-th layer, that is: all entity patterns EP in the k-th layer satisfy: the cardinality |T| of the feature set of all entities = k; m represents the maximum value of all feature sets CS; S24: Generate the poset lattice σL according to the entity patterns CS_T of each layer; S25: Return the poset lattice σL; Algorithm 2 gives the steps for calculating the approximate summary ELSSR of type II semantic graphs on a poset lattice; the inputs of this algorithm are the semantic graph G, the set R of key types t *, the threshold μ(p*) of p*, and the semantic graph type, and the output is the approximate summary μL of type II semantic graphs based on the poset lattice; S31: Initialize the entity pattern set; S32: For each entity s in each semantic graph, if the entity s is associated with a key relationship type p*, then check the relationship between the corresponding edge set |E p* | it is associated with and the set threshold μ(p*). If |E p* | ≥ μ(p*), then add the entity s and all the relationship types it is associated with to μC; S33: Merge entities with the same feature set CS, and hierarchically classify entity patterns EP according to the cardinality of the feature set CS; CS_T k Store the entity patterns EP of the k-th layer, that is: all entity patterns EP in the k-th layer satisfy that the cardinality |T| of the feature sets of all entities is equal to k; m represents the maximum value of all feature sets CS; S34: Generate the poset lattice μL according to the entity patterns CS_T of each layer; S35: Return the poset lattice μL; Algorithm 3 is the ELSRR information degree calculation method; the specific steps are as follows: S41: Initialize the corresponding variable I σ ,V b ,V σ ,V N ,R b ; S42: Calculate the base graph G of σL b ; S43: Calculate the information degree I according to formula (2) σ ; Define the information degree of 6ELSRR: Given a semantic graph a set R of key relationship types t *, and the ELSRR summary σL = (σC, ≤) of this semantic graph, the information degree of ELSRR is defined as: Among them, V b and R b are the entity set and relationship set of the base graph, and V and R are the entity set and relationship set of the semantic graph; S44: Return information degree I σ ; Algorithm 4 is the ELSSR information degree calculation method; the specific steps are as follows: S51: Initialize the corresponding variable I μ ,V b ,V σ ,V N ,R b ; S52: Calculate the number of entities in μL; S53: Calculate the information degree I according to formula (3) μ ; Define the information degree of 7ELSSR: Given a semantic graph a set R of key relationship types t * and its threshold μ(R t *), the ELSSR summary μL = (μC, ≤) of this semantic graph, and the information degree of 7ELSSR is defined as: where V μ is the set of all entities included in μC; S54: Return information degree I μ .
2. A method for approximate summarization of large semantic graphs based on poset lattices according to claim 1, characterized in that: The specific content of S2 is: Definition 1 Entity Pattern: Given a semantic graph G, let be the set of all subjects s in the triples (s, p, o) of the entity; for any be the set of features of the entity s; an entity pattern EP is defined as c = (S, T, A), where: (i) (ii) CS(s) = T; (iii) A = ∪ s∈S L A (s); Let \(C\) be the set of all entity patterns, then forms a poset; if two special entity patterns are set and then forms a poset lattice; Define two types of key relationships: Given a semantic graph G, if a subset of relationship types: is among the top σ% of the relationship types that the semantic graph is most frequently retrieved for, where then R t * is called the set of key relationship types, and the elements in R t * are key relationship types; Set the value of σ to 20; Definition 3: I-Type Semantic Graph Approximate Summary Based on Poset Lattice: Given a semantic graph G and a set R of key relationship types t *, the I-type semantic graph approximate summary based on the poset lattice is defined as the lattice σL formed by the poset (σC, ≤), where σC is a set of entity patterns and each entity pattern contains at least one key relationship type, that is:
3. A method for approximate summarization of large semantic graphs based on poset lattices according to claim 2, characterized in that: The specific content of S3 is: Definition 4: Type-II Semantic Graph Approximate Summary Based on Poset Lattice: Given a semantic graph G and a set of key relationship types The type-II semantic graph approximate summary based on poset lattice is defined as the lattice μL formed by the poset (μC, ≤), where: There is |E p* | ≥ μ(p*), E p* is the set of edges with the relationship type p*, and μ(p*) is the threshold of p*; Set μ(p*) = 2, and filter at least 50% of the entities related to p*; μ(p*) is set by the user himself, and different thresholds are set for different key relationship types p* to achieve the filtering of specified entities.
4. A method for approximate summarization of large semantic graphs based on poset lattices according to claim 3, characterized in that: The specific content of S4 is: Define the base graph of 5ELSRR: Given a semantic graph a set R of key relationship types t *, and the ELSRR summary σL = (σC, ≤) of this semantic graph, the base graph of G is defined as: is a subgraph of the semantic graph G that satisfies: (1)V b = V σ ∪ V N , where V N includes the adjacent nodes of all nodes in V σ ; (2)R b ={(u, v) | u ∈ V σ or v ∈ V σ}; (3) (4) (5) is a mapping that maps the relationships in R b to the relationship types in the semantic graph; (6) is a mapping that maps entities in V b to a set of properties in the semantic graph; The base graph of ELSRR is the subgraph of the original semantic graph covered by the summary.
5. A system for approximate summarization of large semantic graphs based on poset lattices, including a memory, a processor, and a computer program stored on the memory and capable of running on the processor, characterized in that: When the processor executes the computer program, it implements the method according to any one of claims 1 to 4.
6. A computer-readable storage medium, on which a computer program is stored, characterized in that: When the computer program is executed by the processor, it implements the method according to any one of claims 1 to 4.
Citation Information
Patent Citations
A semantic approximate query method for RDF knowledge map
CN108959613A
Fraud behavior mining method for multi-dimensional sparse sales data warehouse
CN111275480A