An intelligent fusion method for heterogeneous knowledge resources
By transforming heterogeneous knowledge resources into directed graphs and performing semantic similarity calculation and graph clustering, the problems of entity merger and semantic merging of heterogeneous knowledge resources are solved, and a unified knowledge system is constructed, realizing the integration and organization of knowledge content.
Patent Information
- Application Number
- CN202210864447.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-21
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-07-21
AI Technical Summary
The existing technology is difficult to effectively integrate heterogeneous knowledge resources, especially in entity mergers and semantic mergers, which leads to the dispersion of knowledge content and is unable to build a complete knowledge system.
Transform heterogeneous knowledge resources into directed graphs, calculate the similarity between nodes through semantic embedding vectors, establish connections or merge nodes, and use graph clustering and OWL ontology transformation to form a unified knowledge representation.
The entity merger and semantic merger of heterogeneous knowledge resources are realized, a complete knowledge system is constructed, and the unity and richness of knowledge organization is improved.
Smart Images

Figure CN115391550B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of computer software, and relates to a method for fusing heterogeneous knowledge resources, and particularly to a method for using a directed graph as a unified knowledge representation of heterogeneous resources. Background Art
[0002] Existing knowledge resources have various formats, including unstructured natural language texts, semi-structured HTML documents, structured XML documents, relational databases, etc. Although there are significant differences in the forms of these knowledge resources, the knowledge content involved may have a high degree of relevance, which is a description of the same real-world entity or related to a specific problem. In order to obtain a complete knowledge description or a solution to a problem, it is necessary to fuse heterogeneous knowledge resources, extract the relevant knowledge content, and organize it in an orderly manner to form a unified description structure and build a knowledge system.
[0003] The present invention mainly focuses on the intelligent fusion method between heterogeneous structured data. The problems to be solved in this regard are: (1) The entity merging problem of heterogeneous knowledge resources. The same knowledge object may have different expressions in different knowledge resources, and there may be different description formats in heterogeneous resources. The fusion of heterogeneous knowledge resources needs to be able to identify the same entity in different original resources and merge different description formats; (2) The semantic merging problem of heterogeneous knowledge resources. Different entities may have different forms in different resources, and at the same time, there are usually semantic information expressions for these entities in the original resources. When fusing the original resources, it is necessary to extract this knowledge content and perform knowledge fusion. Summary of the Invention
[0004] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide an intelligent fusion method for heterogeneous knowledge resources.
[0005] The technical solution of the present invention is as follows:
[0006] An intelligent fusion method for heterogeneous knowledge resources, the steps of which include:
[0007] 1) Respectively convert each knowledge resource to be fused into a corresponding directed graph;
[0008] 2) Generate a semantic embedding vector for each node in the directed graph, and calculate the semantic similarity between nodes according to the semantic embedding vectors of the nodes; if the semantic similarity between two nodes is greater than a set connection threshold, an undirected edge is established to connect the corresponding two nodes; if the semantic similarity between two nodes is greater than a set merging threshold, the corresponding two nodes are merged;
[0009] 3) Perform graph clustering on each of the directed graphs processed in step 2) to obtain multiple clusters;
[0010] 4) Generate semantic topics for the corresponding clusters according to the nodes included in each of the clusters and construct a semantic directed graph;
[0011] 5) Convert the semantic directed graph into an OWL ontology to obtain the fused knowledge resources.
[0012] Furthermore, the knowledge resources include structured data with a nested hierarchical structure and structured data without a nested hierarchical structure; the structured data with a nested hierarchical structure includes XML - formatted data and non - XML - formatted data; among them,
[0013] A) The method for converting XML - formatted knowledge resources into a directed graph is as follows:
[0014] 11) Convert the elements in the XSD document used to describe entities into entity nodes Ve; convert the elements in the XSD document to be processed that describe entity attributes into attribute nodes Vp;
[0015] 12) For the nested relationship N(a,b) in the XSD document to be processed, where a is the parent element and b is the child element; generate a directed edge from the node corresponding to element a to the node corresponding to element b according to N(a,b), and name this directed edge "has" + b; if element b satisfies any one of the conditions (1) - (3), the edge between the node corresponding to element a and the node corresponding to element b is called a class edge; where the conditions (1) - (3) are: (1) the node corresponding to element b is a node under Ve; (2) element b has specific constraint conditions for restriction in the XSD document to be processed; (3) element b is a named node in the XSD document to be processed, that is, element b is an actual business object;
[0016] B) The method for converting non - XML - formatted knowledge resources with a nested hierarchical relationship into a directed graph is as follows:
[0017] 21) Convert the elements in the knowledge resource document used to describe entities into entity nodes Ve; use the attributes of the elements describing entities as attribute nodes Vp corresponding to the entity nodes Ve;
[0018] 22) Generate a directed edge <Ve,Vp> according to the corresponding relationship between the entity node Ve and the attribute node Vp;
[0019] C) The method for converting knowledge resources without a nested hierarchical structure into a directed graph is: take the description unit of each type of entity in the knowledge resources as an entity node Ve; take each attribute unit included in the description unit of the entity as an attribute node Vp; generate a directed edge <Ve,Vp> according to the corresponding relationship between the entity node Ve and the attribute node Vp.
[0020] Further, the knowledge resource without a nested hierarchical structure is a relational database, the description unit is a table in the relational database, and the attribute unit is a field in the relational database; or the knowledge resource without a nested hierarchical structure is a spreadsheet, the description unit is several columns in the spreadsheet, and the attribute unit is a column in the spreadsheet.
[0021] Further, the method for merging corresponding two nodes is as follows: retain the node with more attributes among the two nodes, and add the attributes of the node with fewer attributes to the attributes of the retained node.
[0022] Further, the method for merging corresponding two nodes is as follows: manually determine the node to be retained among the two nodes, and rename, add attributes or update the retained node.
[0023] Further, use the HDBSCAN algorithm to perform graph clustering on each directed graph processed in step 2).
[0024] Further, the method for generating the semantic theme corresponding to the cluster and constructing the semantic directed graph is as follows:
[0025] 41) For each entity node Ve in the cluster, splice the text description of the entity node Ve in the knowledge resource with the texts corresponding to each attribute node Vp connected to the entity node Ve as the description information of the entity node Ve;
[0026] 42) Generate the semantic vector vs corresponding to the entity node Ve according to the description information of the entity node Ve;
[0027] 43) Extract the themes from the description information of the entity node Ve, perform semantic embedding on the top K theme words obtained for each theme using the word2vec algorithm, and splice the obtained semantic embedding representation with each theme category number to obtain the theme vector vt of the entity node Ve;
[0028] 44) Generate the complete vector vc of the entity node Ve according to the semantic vector vs and the theme vector vt;
[0029] 45) Use the clustering algorithm to cluster the obtained complete vectors, generate the theme words of each cluster according to the clustering results; use the set of theme words of each cluster as the theme of the cluster, create a new node Vec as the core node of the cluster, and establish a directed edge <Vec,Ve> between the other entity nodes Ve in the cluster and the node Vec.
[0030] Further, the method for converting the semantic directed graph into an OWL ontology is as follows: for each node Vec and entity node Ve of a clustering group, directly convert them into classes in the OWL language; for the directed edges <Vec, Ve> and <Ve, Ve>, convert them into object properties in the OWL language, and convert the source node in the directed edge into the domain of the object property, the target node into the range of the object property, and the name of the directed edge into the naming of the object property; for the edges <Ve, Vp> and the attribute vertex Vp, convert the name of Vp into the naming of the data property in the OWL language, convert the Ve connected to Vp through the edge <Ve, Vp> into the domain of the data property, and convert the data type of the element corresponding to Vp in the knowledge resource into the range of the data property.
[0031] A server, comprising a memory and a processor, wherein the memory stores a computer program configured to be executed by the processor, and the computer program includes instructions for executing the steps in the above method.
[0032] A computer-readable storage medium, on which a computer program is stored, characterized in that the steps of the above method are implemented when the computer program is executed by a processor.
[0033] Compared with the prior art, the positive effects of the present invention are as follows:
[0034] 1. Using a directed graph as a unified representation form for heterogeneous knowledge resources. By using the definitions of nodes and edges in the graph, the knowledge content in the form of triples in the original knowledge resources is described without considering the original existence form of this knowledge content, and focusing on the knowledge content itself.
[0035] 2. Realize the entity merging of heterogeneous knowledge resources. For the same or similar knowledge entities in different knowledge resources, extract the semantic information of the original expression methods, and make connections or merges according to the semantic similarity. At the same time, further classify the related entities to improve the degree of organization.
[0036] 3. Realize the knowledge merging of heterogeneous knowledge resources. Using the rich semantic information contained in the text descriptions of knowledge entities in the original knowledge resources, integrate and organize the knowledge content scattered in heterogeneous knowledge resources, and combine it with the original entity relationships and attribute information, which is more conducive to constructing a complete knowledge system. BRIEF DESCRIPTION OF THE DRAWINGS
[0037] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0038] The present invention will be further described in detail below with reference to the accompanying drawings. The examples are only used to explain the present invention and are not intended to limit the scope of the present invention.
[0039] Considering that the knowledge content semantics in different knowledge resources can all be expressed in the form of triples <S, P, O>, where S is the subject, that is, the described entity; O is the object, which can be another entity or an attribute possessed by the entity; P is the predicate, which can express both the association relationship between entities and the correspondence relationship between entities and attributes. The OWL language provides a way of ontology expression, in which classes (owl:Class) are defined as descriptions of entities, data properties (owl:DataProperty) are defined as descriptions of the attributes possessed by entities, and object properties (owl:ObjectProperty) are defined as descriptions of the relationships between entities. Using the OWL language can reveal and make explicit the semantic information existing in the original knowledge resources, and also provides a unified expression form, which can be used as the goal of heterogeneous knowledge resource fusion.
[0040] However, if different knowledge resources are directly converted into OWL language descriptions, the problem of the association between different OWLs still exists, and the knowledge content in different original resources has not been integrated and still exists in a scattered form in different OWL ontologies. To address this problem, this paper uses a directed graph structure as the unified expression form for heterogeneous knowledge resource fusion. A directed graph is defined as G=(V, E), where V is the set of vertices, including entities (Ve) and attributes (Vp); E is the set of edges, and each edge contains two vertices, pointing from one vertex to another vertex, including two forms: <Ve, Ve> and <Ve, Vp>.
[0041] The specific process of the intelligent fusion method for heterogeneous knowledge resources is as follows:
[0042] Step1: Convert each knowledge resource to be fused into a corresponding directed graph respectively; the original structured knowledge resources mainly have two types. The first type is XML and data formats similar to XML with nested hierarchical relationships, and the second type is structured data without nested hierarchical structures. For the first type of resource, the conversion method from XML resources to directed graphs is Steps 11 to 12. For data formats similar to XML with nested hierarchical relationships, such as JSON, etc., only the processes of Steps 11 to 12 are required to extract the nested relationships between elements at each level of nesting. These elements themselves will be used as entity nodes Ve in the vertex set, and the attributes possessed by the elements will be used as attribute nodes Vp in the vertex set, and the correspondence relationship between entities and attributes will be used as the directed edge <Ve, Vp>.
[0043] Step11: Determination of directed graph nodes. To determine a directed graph, it is first necessary to determine that the XSD document serves as a specification for the logical structure of the XML document. The complex elements defined in it usually can contain other elements or have many attributes, and are mostly used to describe a class of entities in real life. Therefore, all complex elements in the XSD document are converted into nodes. Secondly, the XSD document also defines some custom element types used in the corresponding XML document. These element types are usually defined to describe specific real entities and contain the knowledge content required to solve the corresponding problems. Therefore, they are also converted into nodes. However, for the sake of distinction, the nodes obtained by converting complex elements are used as entity nodes Ve, corresponding to classes (owl:Class) in OWL, and the nodes obtained by converting other elements are used as attribute nodes Vp, corresponding to data properties (owl:DataProperty) in OWL.
[0044] Step12: Determination of directed graph edges. After obtaining the nodes, it is necessary to determine the link relationship between the nodes, that is, the edges. In the XSD document, there is a nested hierarchical relationship between elements. Therefore, it needs to be converted into a directed graph, and the original hierarchical information is retained in the form of directed edges. The generation of this directed graph is different from the traversal of the XSD structure tree. Instead, it focuses on the nested relationship itself, reorganizes the original tree structure according to the nested relationship, which is a further processing of the XSD structure tree, and there will be loops in the finally obtained directed graph. In the specific conversion, for the nested relationship N(a,b), where a is the parent element and b is the child element, it is converted into a directed edge pointing from a to b, named "has" + the name of b, corresponding to the object property (owl:ObjectProperty) in OWL. Similar to the nodes, in order to distinguish the knowledge content and knowledge functions of different XML elements, the above edges are also distinguished. The edge between a and b that meets the following conditions is called a class edge <Ve,Ve>. The specific conditions that b needs to meet are: (1) It is a node under Ve; (2) There are specific constraint conditions in the XSD for restriction, that is, there exists <restriction>Restrictions are imposed on the elements, which indicates that these elements have certain practical meanings, and these constraints are the formal expressions of these practical meanings; (3) is a named node. In an XML document, node elements may not have names. These nodes may only serve to provide a nested hierarchical structure, but have no practical meaning and contain little knowledge content. The naming of named child nodes usually reflects the actual business object corresponding to this element and needs to be distinguished from other nodes.
[0045] Step2: For structured data without a nested hierarchical structure, such as relational databases, spreadsheets, etc., the description unit for each type of entity, such as a table in a relational database, a form in a spreadsheet, or the whole of several manually specified columns, is used as an entity node Ve in the node set, and a name is given according to the actual meaning. For each field in a relational database and each column in a spreadsheet, they are used as attribute nodes Vp in the vertex set. The corresponding relationship between the entity and the attribute, that is, the inclusion relationship between the whole manually specified as Ve and the fields or columns therein, is also used as a directed edge <Ve, Vp>.
[0046] Step3: Entity and knowledge merging. After the above two operations, heterogeneous knowledge resources can all be expressed using the unified structure of a directed graph. At this time, entity merging and knowledge merging are required. First, semantic embedding needs to be performed on the nodes. Algorithms such as GraphSAGE can be used for learning to obtain the vector representation of the nodes. After that, the semantic similarity between the vertices can be calculated, and connection thresholds and merging thresholds are set manually, where the connection threshold is less than the merging threshold. An undirected edge connection is established between two vertices with a similarity greater than the connection threshold, and the vertices themselves and the attributes of the two vertices with a similarity greater than the merging threshold are merged. Here, either the machine can automatically select the vertex with more attributes and add the attributes of the other vertex to this vertex, or the user can decide which vertex to retain and rename the vertex, add or update attributes, etc. Then, the HDBSCAN algorithm is used for graph clustering on the obtained extended directed graph, and the resulting clusters contain vertices from different original resources.
[0047] Step4: To more clearly obtain the semantic topics of each cluster, further processing needs to be done within each cluster:
[0048] 1) For the text description of each vertex Ve in the original knowledge resource, the text description of vertex Ve is concatenated with the text of each attribute Vp connected to vertex Ve;
[0049] 2) Use the BERT algorithm to learn the semantic embedding information of the concatenated result to obtain the semantic vector vs of each vertex;
[0050] 3) Use the LDA algorithm to extract the themes of the text descriptions of each vertex obtained by splicing in 1), perform semantic embedding on the top 10 theme words of each obtained theme using the word2vec algorithm, and splice them with the respective theme category numbers to obtain the theme vector vt of each vertex;
[0051] 4) Splice vs and vt, and manually set the parameter λ to adjust the weights of the semantic vector and the theme vector to obtain the complete vector vc of each vertex;
[0052] 5) Use the clustering algorithm to cluster the obtained complete vectors, and manually assign theme words according to the clustering results or automatically generate the theme words of the clustering clusters based on the semantic descriptions of the vertices. The overall set of theme words of each subclass cluster is used as the theme of the cluster, and a new vertex Vec is created as the core vertex of the cluster, and a directed edge <Vec,Ve> is established between Vec and other vertices Ve in the cluster.
[0053] After this step is completed, it is equivalent to completing the fusion of heterogeneous knowledge resources and further knowledge content extraction. After that, in order to more clearly express the semantic content in the fusion result, the graph structure is described using the OWL language to obtain the knowledge fusion ontology.
[0054] Step5: According to the mapping rules defined in Table 1, convert the semantic directed graph obtained in Step4 into an OWL ontology.
[0055] Table 1 is the mapping rules from the directed graph to OWL
[0056] Directed graph element OWL element Vec, Ve owl:Class <Vec,Ve>, <Ve,Ve> owl:ObjectProperty domain: Vec or the first Ve range: the second Ve <Ve,Vp>,Vp owl:DataProperty domain: Ve range: the data type of Vp
[0057] For the core vertex Vec and the entity vertex Ve of each clustering cluster, they are directly converted into classes (owl:Class) in the OWL language. For the edges <Vec,Ve> and <Ve,Ve>, they are converted into object properties (owl:ObjectProperty) in the OWL language. Since the class edges are all directed edges, the source node (source) of the directed edge is converted into the domain (domain) of the object property, the target node (target) is converted into the range (target), and the name of the directed edge is converted into the naming of the object property. For the edge <Ve,Vp> and the attribute vertex Vp, the name of Vp is converted into the naming of the data property (owl:DataProperty) in the OWL language, the Ve connected to Vp through the edge <Ve,Vp> is converted into the domain (domain) of the data property, and the data type of the element in the original document corresponding to the Vp vertex is converted into the range (range).
[0058] Although specific embodiments of the present invention are disclosed for illustrative purposes, which are intended to help understand the content of the present invention and implement it accordingly, those skilled in the art can understand that various substitutions, changes, and modifications are possible without departing from the spirit and scope of the present invention and the appended claims. Therefore, the present invention should not be limited to the content disclosed in the preferred embodiments, and the scope of protection claimed by the present invention shall be defined by the scope defined in the claims.< / restriction>
Claims
1. An intelligent fusion method for heterogeneous knowledge resources, the steps of which include: 1) Respectively transform each knowledge resource to be fused into a corresponding directed graph; 2) Generate a semantic embedding vector for each node in the directed graph, and calculate the semantic similarity between nodes according to the semantic embedding vectors of the nodes; if the semantic similarity between two nodes is greater than the set connection threshold, establish an undirected edge to connect between the corresponding two nodes; if the semantic similarity between two nodes is greater than the set merging threshold, merge the corresponding two nodes; 3) Perform graph clustering on the directed graphs processed in step 2) to obtain multiple cliques; 4) Generate a semantic theme corresponding to each clique and construct a semantic directed graph according to the nodes included in each clique, and the method is: 41) For each entity node Ve in the clique, splice the text description of the entity node Ve in the knowledge resource with the texts corresponding to each attribute node Vp connected to the entity node Ve as the description information of the entity node Ve; 42) Generate a semantic vector vs corresponding to the entity node Ve according to the description information of the entity node Ve; 43) Extract the theme from the description information of the entity node Ve, semantically embed the first K topic words obtained for each theme using the word2vec algorithm, and splice the obtained semantic embedding representation with each theme category number to obtain the theme vector vt of the entity node Ve; 44) Generate a complete vector vc of the entity node Ve according to the semantic vector vs and the theme vector vt; 45) Use a clustering algorithm to cluster the obtained complete vectors, generate topic words for each clustering group according to the clustering results; use the set of topic words of each clustering group as the theme of the clique, and create a new node Vec as the core node of the clique, and establish a directed edge <Vec, Ve> between other entity nodes Ve in the clique and this node Vec; 5) Transform the semantic directed graph into an OWL ontology to obtain the fused knowledge resources.
2. The method according to claim 1, characterized in that, The knowledge resources include structured data with a nested hierarchical structure and structured data without a nested hierarchical structure; the structured data with a nested hierarchical structure includes XML - format data and non - XML - format data; among them, A) The method for transforming XML - format knowledge resources into a directed graph is: 11) Transform the elements used to describe entities in the XSD document to be processed into entity nodes Ve; transform the elements used to describe entity attributes in the XSD document to be processed into attribute nodes Vp; 12) For the nesting relationship N(a, b) in the XSD document to be processed, a is the parent element and b is the child element; a directed edge pointing from the node corresponding to element a to the node corresponding to element b is generated according to N(a, b), and this directed edge is named "has" + b; if element b satisfies any one of conditions (1) to (3), the edge between the node corresponding to element a and the node corresponding to element b is called a quasi-edge; where conditions (1) to (3) are: (1) the node corresponding to element b is a node under Ve; (2) element b has specific constraint conditions for restriction in this XSD to be processed; (3) element b is a named node in this XSD to be processed, that is, element b is an actual business object. B) The method for converting a non-XML format knowledge resource with a nested hierarchical relationship into a directed graph is as follows: 21) Convert the elements used to describe entities in the knowledge resource document into entity nodes Ve; use the attributes of the elements describing the entities as the attribute nodes Vp corresponding to the entity nodes Ve. 22) Generate a directed edge <Ve, Vp> according to the corresponding relationship between the entity node Ve and the attribute node Vp. C) The method for converting a knowledge resource without a nested hierarchical structure into a directed graph is as follows: use the description unit of each type of entity in the knowledge resource as an entity node Ve; use each attribute unit included in the description unit of the entity as an attribute node Vp; generate a directed edge <Ve, Vp> according to the corresponding relationship between the entity node Ve and the attribute node Vp.
3. The method according to claim 2, wherein The knowledge resource without a nested hierarchical structure is a relational database, the description unit is a table in the relational database, and the attribute unit is a field in the relational database; or the knowledge resource without a nested hierarchical structure is a spreadsheet, the description unit is several columns in the spreadsheet, and the attribute unit is a column in the spreadsheet.
4. The method according to claim 1 or 2 or 3, characterized in that, The method for merging corresponding two nodes is as follows: retain the node with more attributes among the two nodes, and add the attributes of the node with fewer attributes to the attributes of the retained node.
5. The method according to claim 1 or 2 or 3, characterized in that The method for merging corresponding two nodes is as follows: manually determine the node to be retained among the two nodes, and rename, add attributes or update the retained node.
6. The method according to claim 1 or 2 or 3, characterized in that Use the HDBSCAN algorithm to perform graph clustering on each directed graph processed in step 2).
7. The method according to claim 1, characterized in that, The method for converting the semantic directed graph into an OWL ontology is as follows: for the nodes Vec and entity nodes Ve of each cluster, directly convert them into classes in the OWL language; for the directed edges <Vec, Ve> and <Ve, Ve>, convert them into object properties in the OWL language, and convert the source node in the directed edge into the domain of the object property, the target node into the range of the object property, and the name of the directed edge into the naming of the object property; for the edges <Ve, Vp> and the attribute vertex Vp, convert the name of Vp into the naming of the data property in the OWL language, convert the Ve connected to Vp through the edge <Ve, Vp> into the domain of the data property, and convert the data type of the element corresponding to Vp in the knowledge resource into the range of the data property.
8. A server, characterized in that, It includes a memory and a processor. The memory stores a computer program which is configured to be executed by the processor. The computer program includes instructions for executing each step in any one of the methods recited in claims 1 to 7.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of any one of the methods recited in claims 1 to 7.
Citation Information
Patent Citations
Archive management model construction method and system based on knowledge graph
CN111737471A