A Fine-Grained Knowledge Extraction Method for Open Domains
The method leverages multiple knowledge graphs and deep learning to iteratively refine semantic tags, addressing the reliance on experts and semantic ambiguity in ontology knowledge extraction, achieving precise and automated fine-grained knowledge extraction.
Patent Information
- Application Number
- CN202210313413.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-28
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-03-28
AI Technical Summary
The existing ontology knowledge extraction technology relies on expert extraction rules, and there are problems of knowledge bias and semantic ambiguity, making it difficult to achieve automated and efficient fine-grained knowledge extraction.
By obtaining the ontology of primary and secondary domain libraries, establishing unlabeled vocabulary and labeled corpus, using the LSTM+CRF deep learning model for training, automatically labeling words and establishing fine-grained knowledge elements, ensuring semantic label consistency through similarity comparison, reaching 95% or more, forming a three-level labeled corpus.
It improves the accuracy of knowledge extraction, reduces semantic ambiguity, realizes automated fine-grained knowledge extraction, and reduces dependence on expert knowledge.
Smart Images

Figure CN114676265B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and particularly to a fine-grained knowledge extraction method for open domains. Background Art
[0002] Ontology knowledge systems, as the most important industrialized and commercialized products of the artificial intelligence discipline, assist the field of computer science to develop towards a more intelligent direction. To construct ontology knowledge, many methods have been explored to help extract knowledge from unstructured text data. Since the data and knowledge contained in Internet pages are rich, they provide valuable resources for ontology knowledge construction. The tabular data in Internet pages, due to its structured organization form, is conducive to realizing the mapping between knowledge and data. By extracting web tabular data for ontology knowledge construction, it will effectively help complete the ontology knowledge construction process.
[0003] Existing ontology knowledge extraction technologies mainly focus on the overall implementation of the ontology knowledge construction process, paying more attention to the system or device itself, only providing a human-computer interaction interface to assist in completing each process of ontology knowledge construction, and less involving the innovation of knowledge automatic extraction technology. Knowledge extraction mostly relies on experts to sort out extraction rules or training data. The existing technology is essentially a semi-automatic extraction system that assists in manual sorting work, not a truly automatic extraction, and there is a risk of subsequent errors due to knowledge deviations between experts and data, and semantic ambiguity is likely to occur. Summary of the Invention
[0004] Aiming at the deficiencies of the existing technology, the present invention provides a fine-grained knowledge extraction method for open domains, which solves the risk of subsequent errors caused by knowledge deviations between experts and data and the problem of easy semantic ambiguity.
[0005] Technical Solution
[0006] To achieve the above objectives, the present invention is implemented through the following technical solutions: A fine-grained knowledge extraction method for the open domain, including obtaining the ontology of the primary domain library, and determining each primary domain type in the ontology of the primary domain library according to the existing open knowledge graph; through the ontology of the primary domain library, establishing an unlabeled word library, including the association relationship between the primary domain type and the unlabeled word library; calculating the primary labeled word library through the unlabeled word library and the primary domain type; obtaining the words in the primary labeled word library through the primary labeled word library; based on the ontology of the primary domain library, using a tool to query the words in the primary labeled word library in the ontology of the primary domain library, establishing the association relationship between the words and the primary domain type, and obtaining the primary labeled corpus; based on the primary labeled corpus, establishing a primary training model to train the ontology of the primary domain library; extracting the words with the highest assigned probability as the semantic labels of the primary domain type to obtain the ontology of the secondary domain library; based on the ontology of the secondary domain library, determining the secondary domain types in the ontology of the secondary domain library according to the existing open knowledge graph; calculating the secondary labeled word library through the primary labeled word library and the secondary domain type; obtaining the words in the secondary labeled word library through the secondary labeled word library; based on the ontology of the secondary domain library, using a tool to query the words in the secondary labeled word library in the ontology of the secondary domain library, establishing the association relationship between the words and the secondary domain type, and obtaining the secondary labeled corpus; based on the secondary labeled corpus, establishing a secondary training model to train the ontology of the secondary domain library; extracting the words with the highest assigned probability as the semantic labels of the secondary domain type; comparing the similarity between the semantic labels of the primary domain type and the semantic labels of the secondary domain type through a tool, and when the similarity reaches 95% or more, obtaining the ontology of the tertiary domain library and the corresponding tertiary domain types, tertiary labeled word library, and tertiary labeled corpus; based on the knowledge elements composed of the words with semantic labels in the tertiary labeled word library, extracting fine-grained knowledge through a tool.
[0007] Preferably, the determination of the domain type refers to obtaining a fine-grained knowledge element type table according to the domain requirements and obtaining an unlabeled domain word library.
[0008] Preferably, the open knowledge graph can be at least three of DBpedia, Yago, Wikidata, BabelNet, ConceptNet, Microsoft Concept Graph, and OpenKG used in common.
[0009] Preferably, the labeled domain word library is specifically obtained by calculating the semantic similarity between the information in the domain library ontology and the unlabeled word library through the domain type, and automatically labeling the unlabeled domain word library based on the semantic similarity.
[0010] Preferably, both the primary training model and the secondary training model adopt the combination of LSTM+CRF deep learning and machine learning. The primary training model and the secondary training model respectively divide the first-level annotated corpus and the second-level annotated corpus into training sets, development sets, and test sets in terms of words.
[0011] Preferably, when comparing the semantic tags of the primary domain type with those of the secondary domain type, if the similarity fails to reach 95% or above, it will re-enter the action of establishing the ontology of the secondary domain library and loop, and continuously compare with the semantic tags of the domain type established last time until the similarity reaches 95% or above.
[0012] Preferably, the process of training the primary domain library refers to converting the ontology of the primary domain library into the form of an annotated corpus and inputting it into the primary training model to predict the probability that each word is assigned to each primary domain type.
[0013] Preferably, the process of training the secondary domain library refers to converting the ontology of the secondary domain library into the form of an annotated corpus and inputting it into the secondary training model to predict the probability that each word is assigned to each secondary domain type. Beneficial effects
[0014] The present invention provides a fine-grained knowledge extraction method for the open domain, which has the following beneficial effects:
[0015] Using at least three of DBpedia, Yago, Wikidata, BabelNet, ConceptNet, Microsoft ConceptGraph, and OpenKG as the open knowledge graph can greatly alleviate the problem of strong dependence on domain expert knowledge in the prior art when extracting knowledge elements.
[0016] By repeatedly comparing the semantic tags of the current domain type with those of the previous round of domain types until the corresponding conditions are met, the established tertiary domain type, tertiary annotation word library, and tertiary annotated corpus are made more accurate, avoiding the problem of semantic ambiguity and improving the accuracy of knowledge extraction. Brief description of the drawings
[0017] Figure 1 It is a schematic diagram of the working process of the present invention. Detailed implementation manners
[0018] Next, in combination with the accompanying drawings in the embodiments of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention. Embodiment
[0019] As Figure 1 shown, the embodiment of the present invention provides a fine-grained knowledge extraction method for the open domain, including the following steps:
[0020] S1. Obtain the primary domain library ontology, determine each primary domain type in the primary domain library ontology according to the existing open knowledge graph, establish an unlabeled word library through the primary domain library ontology, including the association relationship between the primary domain type and the unlabeled word library, calculate the primary labeled word library through the unlabeled word library and the primary domain type, and obtain the words in the primary labeled word library through the primary labeled word library;
[0021] S2. Based on the primary domain library ontology, use a tool to query the words in the primary labeled word library in the primary domain library ontology, establish the association relationship between the words and the primary domain type, and obtain the primary labeled corpus;
[0022] S3. Based on the primary labeled corpus, establish a primary training model, train the primary domain library ontology, extract the words with the highest distribution probability as the semantic labels of the primary domain type, and obtain the secondary domain library ontology;
[0023] S4. Based on the secondary domain library ontology, determine the secondary domain types in the secondary domain library ontology according to the existing open knowledge graph; calculate the secondary labeled word library through the primary labeled word library and the secondary domain type; obtain the words in the secondary labeled word library through the secondary labeled word library;
[0024] S5. Based on the secondary domain library ontology, use a tool to query the words in the secondary labeled word library in the secondary domain library ontology, establish the association relationship between the words and the secondary domain type, and obtain the secondary labeled corpus;
[0025] S6. Based on the secondary annotation corpus, establish a secondary training model to train the ontology of the secondary domain library; extract the words with the highest distribution probability as the semantic labels of the secondary domain types; use tools to compare the similarity between the semantic labels of the primary domain types and the semantic labels of the secondary domain types. When the similarity reaches 95% or more, obtain the ontology of the tertiary domain library and the corresponding tertiary domain types, tertiary annotation word library, and tertiary annotation corpus. When the similarity between the semantic labels of the primary domain types and the semantic labels of the secondary domain types fails to reach 95% or more, re-enter the action of establishing the ontology of the secondary domain library and loop, and continuously compare with the semantic labels of the domain types established last time until the similarity reaches 95% or more.
[0026] S7. Based on the knowledge elements composed of the words with semantic labels in the tertiary annotation word library, extract fine-grained knowledge through tools.
[0027] Although the embodiments of the present invention have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A fine-grained knowledge extraction method for the open domain, characterized in that: It includes obtaining the primary domain library ontology and determining each primary domain type in the primary domain library ontology according to the existing open knowledge graph; Through the primary domain library ontology, an unlabeled word library is established, including the association relationship between the primary domain type and the unlabeled word library; The primary labeled word library is calculated through the unlabeled word library and the primary domain type; Through the primary labeled word library, the words in the primary labeled word library are obtained; Based on the primary domain library ontology, a tool is used to query the words in the primary labeled word library in the primary domain library ontology, establish the association relationship between the words and the primary domain type, and obtain the primary labeled corpus; Based on the primary labeled corpus, a primary training model is established to train the primary domain library ontology; Extract the word with the highest assignment probability as the semantic label of the primary domain type to obtain the secondary domain library ontology; Based on the secondary domain library ontology, determine the secondary domain types in the secondary domain library ontology according to the existing open knowledge graph; The secondary labeled word library is calculated through the primary labeled word library and the secondary domain type; Through the secondary labeled word library, the words in the secondary labeled word library are obtained; Based on the secondary domain library ontology, a tool is used to query the words in the secondary labeled word library in the secondary domain library ontology, establish the association relationship between the words and the secondary domain type, and obtain the secondary labeled corpus; Based on the secondary labeled corpus, a secondary training model is established to train the secondary domain library ontology; Extract the word with the highest assignment probability as the semantic label of the secondary domain type; Through a tool, compare the similarity between the semantic label of the primary domain type and the semantic label of the secondary domain type. When the similarity reaches 95% or more, obtain the tertiary domain library ontology and the corresponding tertiary domain types, tertiary labeled word library, and tertiary labeled corpus; Based on the knowledge elements composed of the words with semantic labels in the tertiary labeled word library, fine-grained knowledge is extracted through a tool.
2. The fine-grained knowledge extraction method for the open domain according to claim 1, wherein: The domain type refers to obtaining a fine-grained knowledge element type table according to domain requirements and obtaining an unlabeled domain word library.
3. The fine-grained knowledge extraction method for the open domain according to claim 2, wherein: The open knowledge graph is at least three of DBpedia, Yago, Wikidata, BabelNet, ConceptNet, Microsoft ConceptGraph, and OpenKG used in common.
4. A fine-grained knowledge extraction method for the open domain according to claim 3, characterized in that: The labeled domain word library specifically calculates the semantic similarity between the information in the domain library ontology and the unlabeled word library through the domain type, and automatically labels the unlabeled domain word library based on the semantic similarity.
5. A fine-grained knowledge extraction method for the open domain according to claim 4, characterized in that: Both the primary training model and the secondary training model combine LSTM+CRF deep learning and machine learning. The primary training model and the secondary training model respectively divide the primary labeled corpus and the secondary labeled corpus into training sets, development sets, and test sets in terms of words.
6. A fine-grained knowledge extraction method for the open domain according to claim 5, characterized in that: When comparing the semantic label of the primary domain type with the semantic label of the secondary domain type, when the similarity is less than 95% or more, re-enter the action of establishing the secondary domain library ontology and loop, and continuously compare with the semantic label of the domain type established last time until the similarity reaches 95% or more.
7. A fine-grained knowledge extraction method for the open domain according to claim 6, characterized in that: The primary domain library training process refers to converting the primary domain library ontology into the form of labeled corpus, inputting it into the primary training model, and predicting the probability that each word is assigned to each primary domain type.
8. A fine-grained knowledge extraction method for the open domain according to claim 7, characterized in that: The secondary domain library training process refers to converting the secondary domain library ontology into the form of labeled corpus, inputting it into the secondary training model, and predicting the probability that each word is assigned to each secondary domain type.
Citation Information
Patent Citations
Medical record labeling method and device and storage medium
CN112860842A
Legal knowledge graph construction method and related equipment
CN113590846A