An Ontology Intelligent Generation Method

Through the XML to OWL transformation method based on directed graphs, the resource organization and knowledge content disclosure problems when XML knowledge resources are converted into OWL ontology are solved, and more efficient knowledge resource utilization and richer knowledge content expression are achieved.

CN115292512BActive Publication Date: 2025-06-20PEKING UNIV +1
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202210864480.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-07-21
Publication Date
2025-06-20
Estimated Expiration
2042-07-21

AI Technical Summary

Technical Problem

When the existing technology converts XML knowledge resources into OWL ontology, it is impossible to effectively organize multi-source heterogeneous knowledge resources, resulting in complex ontology structure, low knowledge utilization efficiency, and inability to fully reveal the deep relationship of knowledge content.

Method used

Using the XML to OWL conversion method based on directed graphs, the directed graph is generated by converting elements in the XSD document into class nodes and data attribute nodes, and combining nodes through semantic embedding and similarity to form cluster nodes, and finally generating the resource knowledge content ontology described by the OWL language.

Benefits of technology

It has improved the degree of organization of original resources, revealed more knowledge content, formed a more effective knowledge system, and improved the efficiency of knowledge resource utilization.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115292512B_ABST
    Figure CN115292512B_ABST
Patent Text Reader

Abstract

The present invention discloses a method for intelligently generating an ontology, and its steps include: 1) converting the elements for describing entities in the XSD document to be processed into class nodes; converting the elements for describing entity attributes in the XSD document to be processed into data attribute nodes; 2) determining the edges between the nodes corresponding to each element according to the nested hierarchical relationship between the elements in the XSD document to be processed, and generating a directed graph corresponding to the XSD document to be processed; 3) generating semantic embedding vectors for each node in the directed graph, calculating the semantic similarity between nodes according to the semantic embedding vectors of the nodes; merging the nodes with semantic similarity greater than a set threshold into cluster nodes; 4) obtaining an ontology of resource knowledge content described in OWL language according to the directed graph processed in step 3). The present invention can reveal more knowledge content in the original XML resources and improve the description and revelation ability of the ontology for the original knowledge content.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a method for intelligent generation of an ontology, and particularly to an ontology intelligent generation method for extracting knowledge content from original knowledge resources. Background Art

[0002] An ontology has strong semantic description capabilities and can describe entities in the real world and reveal the associations between knowledge. The currently widely used data exchange format is XML, but XML can only express the hierarchical nesting relationships between different elements and cannot well reveal the rich semantic content in an XML document. While the OWL ontology has rich expressiveness and can describe the interrelationships between knowledge contents in the original knowledge resources and express them in a systematic and formal way. Therefore, in order to better mine the knowledge content in XML knowledge resources, a conversion method from XML to OWL is needed. Existing conversion methods mostly establish mappings directly, or directly perform conversions according to the element types defined by XSD, or use the tree structure of XSD itself for conversion. The OWL ontologies obtained by these methods can only express the semantic information in the original XML document's hierarchical nesting structure. When extracting knowledge content from XML knowledge resources, the following problems exist: (1) Traditional methods cannot organize multi-source heterogeneous knowledge resources better. The tags involved in the original knowledge resources (such as XML resource files) are complex and diverse, and simple mapping relationships alone cannot well organize and sort the tags, resulting in an extremely complex ontology structure as the resource scale increases, without forming an effective knowledge system and extremely low utilization efficiency of knowledge resources. (2) Traditional methods cannot well reveal the rich knowledge content contained in the original knowledge resources. Existing methods mainly perform conversions on the hierarchical nesting structure of XML, but information such as different nesting levels and nesting positions has not been fully utilized, lacking the upper and lower position relationships for forming a knowledge system. In fact, only obtaining the XSD structure tree and performing conversions does not deeply analyze at the semantic level, and the deeper knowledge content contained in XML knowledge resources has not been further described and revealed. Summary of the Invention

[0003] Aiming at the problems existing in the prior art, the purpose of the present invention is to provide an ontology intelligent generation method. The present invention is based on the conversion from XML to OWL of a directed graph to obtain the ontology in XML.

[0004] The technical solution of the present invention is as follows:

[0005] An ontology intelligent generation method, the steps of which include:

[0006] 1) Convert the elements for describing entities in the to-be-processed XSD document into class nodes; convert the elements for describing entity attributes in the to-be-processed XSD document into data attribute nodes;

[0007] 2) Determine the edges between the nodes corresponding to each element according to the nesting level relationship between the elements in the XSD document to be processed, and generate a directed graph corresponding to the XSD document to be processed;

[0008] 3) Generate the semantic embedding vector of each node in the directed graph, and calculate the semantic similarity between nodes according to the semantic embedding vectors of the nodes; Merge the nodes with semantic similarity greater than the set threshold into cluster nodes;

[0009] 4) Obtain the ontology of the resource knowledge content described in OWL language according to the directed graph processed in step 3).

[0010] Furthermore, the method for generating the directed graph is as follows: For the nesting relationship N(a, b) in the XSD document to be processed, a is the parent element and b is the child element; Generate a directed edge from the node corresponding to element a to the node corresponding to element b according to N(a, b), and name the directed edge "has" + b; If element b satisfies any one of the conditions (1) to (3), the edge between the node corresponding to element a and the node corresponding to element b is called a class edge; Where the conditions (1) to (3) are: (1) The node corresponding to element b is a node under the class node; (2) Element b has specific constraint conditions in the XSD document to be processed for restriction; (3) Element b is a named node in the XSD document to be processed, that is, element b is an actual business object.

[0011] Furthermore, the method for merging the nodes with semantic similarity greater than the set threshold into cluster nodes is as follows: 1) Generate the XML structure tree of the processed XSD document; Put the nodes with semantic similarity greater than the set threshold into the same node clique, perform clustering on each node clique, and select a node from each clustering cluster I as the cluster node, where the node corresponding to the element with the shortest distance from the root node of the XML structure tree among the elements corresponding to the nodes in the clustering cluster I is selected as the cluster node of the clustering cluster I; 2) Establish a directed edge from the cluster node in the clustering cluster I to other nodes in the clustering cluster I, and name it "hasMember" + node name.

[0012] Furthermore, the method for obtaining the ontology of the resource knowledge content described in OWL language according to the directed graph processed in step 3) is as follows: Convert the class nodes or cluster nodes in the directed graph into classes in OWL language; Convert the directed edges between class nodes into object properties in OWL language, convert the source node of the directed edge into the domain of the object property, the target node into the range, and convert the name of the directed edge into the name of the object property; Convert the names of non-class nodes in the directed graph into the names of data properties in OWL language, use the class nodes connected by non-class nodes as the domain of the data property, and convert the data type of the element corresponding to the non-class node into the range of the data property.

[0013] Furthermore, the GraphSAGE algorithm is used to generate the semantic embedding vectors of each node in the directed graph.

[0014] A server, comprising a memory and a processor, wherein the memory stores a computer program configured to be executed by the processor, and the computer program includes instructions for performing the steps in the above method.

[0015] A computer-readable storage medium, on which a computer program is stored, characterized in that the steps of the above method are implemented when the computer program is executed by a processor.

[0016] The advantages of the present invention are as follows:

[0017] 1. The organization level of the original resources is higher. When generating the ontology, it does not directly map only based on the nested hierarchical relationship of elements. Instead, a directed graph is first used as a knowledge representation form of the original knowledge resources in the XML document rather than a resource structure representation, and on this basis, it is extended to improve the organization level of the original resources.

[0018] 2. The extraction of knowledge content in the original knowledge resources is richer. In addition to using a directed graph as a knowledge representation form, the semantic embedding of the nodes in the graph is learned, and new nodes are created according to the semantic similarity as an explicit expression of the original similar semantic information, revealing more knowledge content in the original XML resources.

[0019] 3. The transformation method can be automated, and manual participation can be introduced to improve efficiency. Although this method uses a directed graph as a transformation intermediary, the entire process can still be implemented in an automated manner, and expert knowledge can be introduced during the transformation process to further improve the description and revelation ability of the ontology for the original knowledge content. Description of the Drawings

[0020] Figure 1 is a flowchart of the method of the present invention. Detailed Embodiment

[0021] The present invention will be further described in detail below with reference to the drawings. The examples given are only used to explain the present invention and are not intended to limit the scope of the present invention.

[0022] Before the transformation, it is necessary to obtain the validation document of XML first. To improve the transformation efficiency from XML to OWL, it is necessary to clarify the specific definitions and restrictions of each element in the XML document first. A legal XML document must correspond to a validation document as the definition of the content elements in the XML document. The validation document here mainly refers to the XSD document, and other types of validation documents can also be used in the present invention. The XSD validation document defines the specific elements and the hierarchical nesting relationships between the elements in the corresponding XML document. These two parts of content will be used as the basis for the specific transformation next.

[0023] The specific process of the method of the present invention is as follows:

[0024] Step1: Determination of the nodes of the directed graph. To determine a directed graph, it is first necessary to determine the nodes therein. The complex elements defined in the XSD document usually can contain other elements or have more attributes, and are mostly used to describe a class of entities in real life. Therefore, the complex elements in the XSD document are all transformed into nodes. Secondly, the XSD document will also define some custom element types used in the corresponding XML document. These element types are usually defined to describe specific real entities, contain the knowledge content required to solve the corresponding problems, and are used to describe the attributes of the entities. Therefore, they are also transformed into nodes. However, for the sake of distinction, the nodes transformed from complex elements are used as class nodes, corresponding to the classes (owl:Class) in OWL, and the nodes transformed from the elements used to describe entity attributes are used as data property nodes, corresponding to the data properties (owl:DataProperty) in OWL.

[0025] Step2: Determination of the edges of the directed graph. After obtaining the nodes, it is necessary to determine the link relationship between the nodes, that is, the edges. In the XSD document, there is a hierarchical nesting relationship between the elements. Therefore, it needs to be transformed into a directed graph, and the original hierarchical information is retained in the form of directed edges. The generation of the directed graph here is different from the traversal of the XSD structure tree, but focuses on the nesting relationship itself. The original tree structure is re-integrated according to the nesting relationship, which is a further processing of the XSD structure tree. There will be loops in the finally obtained directed graph. In the specific transformation, for the nesting relationship N(a,b), where a is the parent element and b is the child element, it is transformed into a directed edge pointing from a to b, named "has" + the name of b, corresponding to the object property (owl:ObjectProperty) in OWL. Similar to the nodes, in order to distinguish the knowledge content and knowledge functions of different XML elements, the above-mentioned edges are also distinguished. The edge between a and b that satisfies the following conditions in the nesting relationship is called a class edge. The specific conditions that b needs to meet are: (1) It is a node under the class node; (2) There are specific constraint conditions in the XSD for restriction, that is, there exists <restriction>Restrictions are imposed on the elements, which indicates that these elements have certain practical meanings, and these constraints are the formal expressions of these practical meanings; (3) is the named node. In an XML document, node elements may not have names. These nodes may only serve to provide a nested hierarchical structure, but have no practical meaning and contain little knowledge content. The naming of named child nodes usually reflects the actual business object corresponding to this element and needs to be distinguished from other nodes.

[0026] Step 3: Merge similar nodes. First, semantic embedding needs to be performed on the nodes. Algorithms such as GraphSAGE can be used for learning to obtain the vector representation of the nodes. After that, nodes with high semantic similarity can be merged. Here, the cosine similarity calculation method is mainly used to calculate the similarity of the semantic embedding vector results of each node, and nodes with a similarity greater than the threshold are merged into the same class of nodes. The threshold here is adjusted manually according to the specific business needs. The specific method of merging into the same class of nodes is as follows: (1) Obtain the node clusters with a similarity greater than the threshold. That is, if the similarities between a and b, b and c, and d and e are all greater than the threshold, then a, b, and c form a node cluster, and d and e form a node cluster; perform clustering on each node cluster, create a cluster node for each clustering cluster. If there are multiple cluster nodes, select the node whose corresponding element is closest to the root element in the original XML structure tree as the cluster node, and the node is named manually according to the actual meaning of the clustering result; (2) Establish a directed edge from the cluster node to the group of nodes with a similarity greater than the threshold in (1), named "hasMember" + the name of the node in the cluster.

[0027] Step 4: Generation of the OWL ontology. After completing the above steps, the node cluster structure is added to the originally directly transformed directed graph, certain semantic information is obtained, and it is expanded. Then, the transformation from the directed graph to OWL can be directly performed to obtain the ontology of the original resource knowledge content described in the OWL language. The transformation rules are shown in Table 1.

[0028] Table 1 shows the transformation rules from the directed graph to OWL

[0029]

[0030] For class nodes or cluster nodes, they are directly transformed into classes (owl:Class) in OWL language. For class edges, that is, the edges between class nodes, they are transformed into object properties (owl:ObjectProperty) in OWL language. Since class edges are all directed edges, the source node of the directed edge is transformed into the domain of the object property, the target node is transformed into the range, and the name of the directed edge is transformed into the naming of the object property. For non-class nodes, the name of the node is transformed into the naming of a data property (owl:DataProperty) in OWL language. The class nodes connected by edges to the non-class nodes serve as the domain of the data property, and the data type of the source XSD element corresponding to the non-class node is transformed into the range of the data property. Thus, the knowledge content represented by a directed graph can be transformed into a description using OWL language, obtaining the ontology expression of the knowledge content of the original knowledge resource.

[0031] Although specific embodiments of the present invention are disclosed for illustrative purposes, which are intended to help understand the content of the present invention and implement it accordingly, those skilled in the art can understand that: within the spirit and scope of the present invention and the appended claims, various substitutions, changes, and modifications are possible. Therefore, the present invention should not be limited to the content disclosed in the best embodiments, and the scope of protection claimed by the present invention shall be defined by the scope defined in the claims.< / restriction>

Claims

1. A method for intelligent generation of an ontology, the steps of which include: 1) Convert the elements used to describe entities in the XSD document to be processed into class nodes; Convert the elements describing entity attributes in the XSD document to be processed into data attribute nodes; 2) Determine the edges between the nodes corresponding to each element according to the nested hierarchical relationship between the elements in the XSD document to be processed, and generate a directed graph corresponding to the XSD document to be processed; 3) Generate semantic embedding vectors for each node in the directed graph, and calculate the semantic similarity between nodes according to the semantic embedding vectors of the nodes; Merge the nodes with semantic similarity greater than the set threshold into cluster nodes; 4) Obtain the ontology of resource knowledge content described in OWL language according to the directed graph processed in step 3).

2. The method according to claim 1, characterized in that The method for generating the directed graph is as follows: For the nested relationship N(a, b) in the XSD document to be processed, a is the parent element and b is the child element; Generate a directed edge from the node corresponding to element a to the node corresponding to element b according to N(a, b), and name the directed edge "has”+b; If element b satisfies any one of conditions (1) to (3), the edge between the node corresponding to element a and the node corresponding to element b is called a class edge; Where conditions (1) to (3) are: (1) The node corresponding to element b is a node under the class node; (2) Element b has specific constraint conditions in the XSD document to be processed for restriction; (3) Element b is a named node in the XSD document to be processed, that is, element b is an actual business object.

3. The method according to claim 1, characterized in that The method for merging nodes with semantic similarity greater than the set threshold into cluster nodes is as follows: 1) Generate the XML structure tree of the processed XSD document; Put the nodes with semantic similarity greater than the set threshold into the same node clique, perform clustering on each node clique, and select a node from each clustering cluster I as the cluster node, where the node corresponding to the element closest to the root node of the XML structure tree among the elements corresponding to each node in the clustering cluster I is selected as the cluster node of the clustering cluster I; 2) Establish a directed edge from the cluster node in the clustering cluster I to other nodes in the clustering cluster I, and name it "hasMember”+node name.

4. The method according to claim 1 or 2 or 3, characterized in that The method for obtaining the ontology of resource knowledge content described in OWL language according to the directed graph processed in step 3) is as follows: Convert the class nodes or cluster nodes in the directed graph into classes in OWL language; Convert the directed edges between class nodes into object properties in OWL language, convert the source node of the directed edge into the domain of the object property, the target node into the range, and convert the name of the directed edge into the name of the object property; Convert the names of non-class nodes in the directed graph into the names of data properties in OWL language, use the class nodes connected by the non-class nodes as the domain of the data property, and convert the data type of the element corresponding to the non-class node into the range of the data property.

5. The method according to claim 1 or 2 or 3, characterized in that Use the GraphSAGE algorithm to generate semantic embedding vectors for each node in the directed graph.

6. A server, characterized in that It includes a memory and a processor, the memory stores a computer program, the computer program is configured to be executed by the processor, and the computer program includes instructions for executing each step in any one of claims 1 to 5.

7. A computer-readable storage medium, on which a computer program is stored, characterized in that When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Heterogeneous knowledge resource intelligent fusion method

    CN115391550A