An ontology constraint-based knowledge extraction method and device and storage medium

By parsing the government affairs knowledge graph ontology and generating a node tree from HTML files, an extraction mapping model and rules are generated, solving the problem of redundant data in the updating of government affairs knowledge graphs and achieving efficient updating of government affairs information.

CN115757818BActive Publication Date: 2025-11-28CETC BIGDATA RES INST CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211384144.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-07
Publication Date
2025-11-28
Estimated Expiration
2042-11-07

AI Technical Summary

Technical Problem

The existing methods of crawling government websites result in the processing of redundant data information when updating government knowledge graphs, which reduces update efficiency.

Method used

By parsing the government affairs ontology, generating triplet files, obtaining the node tree of the HTML file, and generating extraction mapping models and rules based on the description information, the target knowledge is accurately extracted and mapped to the government affairs ontology, reducing the extraction of redundant data.

Benefits of technology

It improves the efficiency of updating government knowledge graphs, reduces the cost of processing redundant data, and enables real-time updates of government information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115757818B_ABST
    Figure CN115757818B_ABST
Patent Text Reader

Abstract

The application discloses a knowledge extraction method and device based on ontology constraint and a storage medium, and is used for improving the efficiency of updating government affair knowledge graph. The application comprises the following steps: analyzing a government affair graph ontology, generating a triple file after the analysis, and determining first description information of each concept, attribute and relationship according to the triple file; obtaining an HTML file of a government affair webpage, generating a node tree according to the HTML file, and obtaining second description information of leaf nodes in the node tree; generating an extraction mapping model according to the first description information and the second description information; setting an extraction rule according to the extraction mapping model; extracting target knowledge in the node tree according to the extraction rule, and mapping the target knowledge to the government affair graph ontology.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of knowledge graph, and particularly relates to a knowledge extraction method and device based on ontology constraint and a storage medium. BACKGROUND

[0002] With the continuous development of e-government construction, a large amount of government data related to public life exists in the web pages of government websites at all levels. In order to obtain perfect government information required by the government graph from the web pages of government websites at all levels, the knowledge extraction technology is applied to the government web pages. As a key step of constructing a knowledge graph, knowledge extraction extracts the knowledge required by the graph from various structured and unstructured text data in the government website, thereby constructing a knowledge graph related to the government at all levels.

[0003] In the process of using the knowledge extraction in the existing government web pages, the knowledge extraction task is usually divided into two parts, and a method of combining web crawlers and text data knowledge extraction is adopted, that is, first, the government information data is completely crawled from the government website through the crawler, and then the triple knowledge is extracted from the government information data through the knowledge extraction, and then the knowledge graph of the government web page is constructed through the triple knowledge.

[0004] However, the existing web crawler method completely crawls the government information data, which causes a lot of redundant data information processing when updating the government knowledge graph and processing data, thereby reducing the efficiency of updating the government knowledge graph. SUMMARY

[0005] In order to solve the above technical problems, the present application provides a knowledge extraction method and device based on ontology constraint and a storage medium, which improves the efficiency of updating the government knowledge graph.

[0006] The first aspect of the present application provides a knowledge extraction method based on ontology constraint, comprising:

[0007] analyzing the government graph ontology, generating a triple file after analysis, and determining the first description information of each concept, attribute and relationship according to the triple file;

[0008] obtaining the HTML file of the government web page, generating a node tree according to the HTML file, and obtaining the second description information of the leaf node in the node tree;

[0009] setting an extraction rule according to the first description information and the second description information;

[0010] extracting target knowledge in the node tree according to the extraction rule, and mapping the target knowledge to the government graph ontology.

[0011] Optionally, the setting the extraction rule according to the extraction mapping model comprises:

[0012] The extraction mapping model sets the extraction rule by similarity calculation or rule mapping.

[0013] Optionally, the extracting the target knowledge in the node tree according to the extraction rule comprises:

[0014] locating the target knowledge in the node tree according to the extraction rule.

[0015] Optionally, the obtaining the HTML file of the government affair webpage comprises:

[0016] respectively obtaining the HTML files corresponding to government affair service items, integrated items of one thing, government-citizen interaction and item list.

[0017] Optionally, the generating the node tree according to the HTML file comprises:

[0018] generating the node tree according to the format tags of HTML in the HTML file.

[0019] The second aspect of the application provides a knowledge extraction device based on ontology constraint, comprising:

[0020] a parsing unit configured to parse a government affair graph ontology, generate a triple file after parsing, and determine first description information of each concept, attribute and relationship according to the triple file;

[0021] an obtaining unit configured to obtain an HTML file of a government affair webpage, generate a node tree according to the HTML file, and obtain second description information of leaf nodes in the node tree;

[0022] a generating unit configured to generate an extraction mapping model according to the first description information and the second description information;

[0023] a setting unit configured to set an extraction rule according to the extraction mapping model;

[0024] a processing unit configured to extract target knowledge in the node tree according to the extraction rule, and map the target knowledge to the government affair graph ontology.

[0025] Optionally, the setting unit comprises:

[0026] a setting module configured to set the extraction rule by the extraction mapping model through similarity calculation or rule mapping.

[0027] Optionally, the processing unit comprises:

[0028] locating the target knowledge in the node tree according to the extraction rule.

[0029] Optionally, the obtaining unit comprises:

[0030] The obtaining module is configured to obtain HTML files corresponding to government affair service items, integrated items of one thing, government-citizen interaction, and item lists, respectively.

[0031] Optionally, the obtaining unit comprises:

[0032] The generating module is configured to generate the node tree according to the format tags of HTML in the HTML files.

[0033] The third aspect of the present application provides a knowledge extraction device based on ontology constraint, comprising:

[0034] The central processor, the memory, the input and output interface, the wired or wireless network interface, and the power supply;

[0035] The memory is a transitory storage memory or a persistent storage memory;

[0036] The central processor is configured to communicate with the memory and execute the instruction operation in the memory to perform the mode of any one of the first aspect and the optional modes of the first aspect.

[0037] The fourth aspect of the present application provides a computer readable storage medium comprising instructions, when the instructions are run on a computer, the computer executes the mode of any one of the first aspect and the optional modes of the first aspect.

[0038] From the above technical solutions, the present application has the following effects:

[0039] The present application first analyzes the government affair spectrum ontology, determines the first description information of each concept, attribute, and relationship in the government affair spectrum ontology, then obtains the HTML file of the government affair web page and the second description information of the leaf node in the node tree corresponding to the HTML, then generates an extraction mapping model according to the first description information and the second description information, sets an extraction rule according to the extraction mapping model, finally extracts the target knowledge in the node tree according to the extraction rule, and maps the target knowledge to the government affair spectrum ontology, thereby realizing the construction of the government affair spectrum, realizing the updating of the real-time government affair information on the government affair web page to the government affair spectrum ontology, so that the target knowledge is extracted from the government affair web page based on the government affair spectrum ontology, the target knowledge is the knowledge required in the government affair spectrum ontology, the extraction is performed for the knowledge required in the government affair spectrum ontology in the knowledge extraction process, the extraction of unnecessary redundant data is reduced, the cost of processing redundant data is reduced, and the efficiency of updating the government affair knowledge spectrum is improved. BRIEF DESCRIPTION OF DRAWINGS

[0040] In order to more clearly illustrate the technical solutions in the present application, the drawings needed to be used in the embodiments will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative effort on the basis of these drawings.

[0041] Figure 1 An example diagram of a knowledge extraction method based on ontology constraints according to the present application;

[0042] Figure 2 An example diagram of a knowledge extraction device based on ontology constraints according to the present application;

[0043] Figure 3 Another example diagram of a knowledge extraction device based on ontology constraints according to the present application;

[0044] Figure 4 Another example diagram of a knowledge extraction device based on ontology constraints according to the present application;

[0045] Figure 5 An example diagram of a triple in a knowledge extraction method based on ontology constraints according to the present application;

[0046] Figure 6 An example diagram of a node tree in a knowledge extraction method based on ontology constraints according to the present application. DETAILED DESCRIPTION

[0047] The present application provides a knowledge extraction method, device and storage medium based on ontology constraints, which is used to improve the efficiency of updating a government knowledge graph.

[0048] It should be noted that the knowledge extraction method based on ontology constraints provided by the present application can be applied to a terminal, a system or a server. For example, the terminal can be a smart phone or a computer, a tablet computer, a portable computer terminal, or a desktop computer. For the convenience of description, the terminal is taken as an example in the present application. In addition, the present application can be applied not only to the construction of a government knowledge graph, but also to the construction of other knowledge graphs, such as a company knowledge graph.

[0049] Please refer to Figure 1 , Figure 1 An example diagram of a knowledge extraction method based on ontology constraints according to the present application, which comprises:

[0050] 101. The terminal parses the government affairs spectrum ontology, generates a triple file after parsing, and determines the first descriptive information of each concept, attribute, and relationship based on the triple file.

[0051] The government affairs ontology is pre-constructed by humans. Staff members set up the government affairs ontology according to their needs, constructing the ontology's concepts, attributes, and relationships between concepts. A concept can also be called an ontology. Parsing the government affairs ontology mainly involves analyzing the concepts, concept attributes, relationships between concepts, constraints between relationships, and the hierarchical structure between concepts contained in the ontology file.

[0052] A triplet file consists of several triples, represented as "ontology, relation, instance". For example, in a government affairs graph, the triples are "basic information, processing method, online processing". In this embodiment, the terminal identifies several triples existing in the government affairs graph ontology to form a triplet file. Then, based on this triplet file, it determines the first descriptive information for each concept, attribute, and relation. The format of the triplet file can be found in [reference needed]. Figure 5 As shown, the descriptive information originates from the description in the government affairs graph ontology. The first descriptive information refers to the sum of several pieces of information, attributes, and relationships possessed by the ontology.

[0053] In this embodiment, the terminal determines the first descriptive information through a manually constructed government affairs map ontology. The first descriptive information is used to guide the knowledge crawling process of government affairs web pages, that is, to guide the knowledge to be extracted from government affairs web pages through the first descriptive information.

[0054] 102. The terminal obtains the HTML file of the government affairs webpage, generates a node tree based on the HTML file, and obtains the second description information of the leaf nodes in the node tree.

[0055] Each provincial, municipal, and district government has its own official website. The government information on each website is different, and the webpages display a wealth of government information. These documents can be used to construct a government information map. The content displayed on the webpages is based on HTML files.

[0056] In this embodiment, the terminal obtains the HTML file of a government affairs webpage and generates a node tree, also known as a DOM tree, based on the HTML file. The leaf nodes in the node tree represent the government affairs knowledge to be extracted, and the intermediate nodes are descriptions of that knowledge. The terminal obtains the second description information of the leaf nodes in the node tree, and through this second description information, it can determine which leaf nodes correspond to which knowledge needs to be extracted. This application combines the node tree generated from the government affairs webpage with... Figure 6 As shown.

[0057] Optionally, in some implementable manners, the terminal respectively acquires HTML files corresponding to the government affair service item, the one-thing integrated item, the government-citizen interaction, and the item list on the government affair webpage.

[0058] Optionally, when generating the node tree, the terminal generates the node tree according to the format tags of HTML in the HTML file. The format tags are, for example, "div", "body", and the like. The node tree is generated according to the corresponding content under the format tags. A plurality of tags correspond to a plurality of node trees.

[0059] 103. The terminal generates an extraction mapping model according to the first description information and the second description information.

[0060] The knowledge corresponding to the leaf node is the knowledge to be extracted. In order to determine which knowledge on the leaf node is required in the government affair graph ontology and accurately extract the target knowledge to be extracted of the leaf node into the government affair graph ontology, the terminal generates an extraction mapping model according to the first description information and the second description information in this embodiment. The extraction mapping model establishes a one-to-one mapping relationship between the leaf node and the knowledge of the government affair graph ontology, thereby facilitating accurate extraction of the knowledge of the leaf node into the government affair graph ontology.

[0061] 104. The terminal sets an extraction rule according to the extraction mapping model.

[0062] The result mapped by the extraction mapping model is the extraction rule. The extraction rule is used in the process of actually extracting the knowledge corresponding to the leaf node into the government affair graph ontology. Actually, the process of mapping the result by the extraction mapping model is as follows: optionally, the extraction mapping model sets the extraction rule by similarity calculation or rule mapping. The similarity calculation is used to calculate the similarity of the first description information and the second description information. When the similarity is greater than a preset threshold, the terminal determines that the second description information is similar to the first description information. At this time, the terminal determines that the knowledge of the leaf node corresponding to the second description information is the target knowledge to be extracted. For example, when the preset threshold is 80%, the first description information is "handling form", and the second description information is "handling method", the terminal calculates the similarity of "handling form" and "handling method". The calculation shows that the similarity of "handling form" and "handling method" is 80%. Then, the terminal determines that the first description information is similar to the second description information, and determines that the knowledge corresponding to the second description information is the knowledge required in the government affair graph ontology. In addition, for the attribute information of a plurality of ontologies in the government affair graph ontology and the relationship information between the ontologies, the first description information after parsing of the government affair graph ontology is calculated for similarity with the description information of the leaf node in the HTML file. If the similarity is greater than 80%, a mapping relationship between the ontology attribute or relationship and the current node in the HTML file is established. The knowledge is extracted through the mapping relationship.

[0063] It can be understood that the preset threshold is set artificially, for example, the similarity is greater than 80% or the similarity is greater than 90%, and the specific threshold is not limited in the application.

[0064] 105、The terminal extracts target knowledge in the node tree according to the extraction rule, and maps the target knowledge to the government spectrum graph ontology.

[0065] After the terminal sets the extraction rule according to the extraction mapping model, the terminal extracts target knowledge in the node tree according to the extraction rule, the target knowledge is the knowledge that needs to be extracted into the government spectrum graph ontology, and the terminal maps the extracted target knowledge to the government spectrum graph ontology. For example, the "handling mode" in the node tree is similar to the "handling form" in the government spectrum graph ontology, so the terminal extracts the "online handling" corresponding to the "handling mode" into the knowledge corresponding to the "handling form" in the government spectrum graph ontology.

[0066] Optionally, in the process of extracting target knowledge in the node tree according to the extraction rule, the terminal also extracts the position of the target knowledge in the node tree according to the extraction rule, and after determining the accurate position of the target knowledge in the node tree, the terminal extracts the target knowledge for the accurate position, which can improve the accuracy of extraction.

[0067] In this embodiment, first, the government spectrum graph ontology is parsed, the first description information is obtained through the government spectrum graph ontology, then the HTML file on the government web page is obtained in real time, and the node tree is generated according to the HTML file, then the second description information is obtained according to the node tree, then the extraction mapping model is generated according to the first description information and the second description information, and the extraction rule is set according to the extraction mapping model, finally the knowledge corresponding to the leaf node in the node tree is extracted into the government spectrum graph ontology according to the extraction rule, and the construction of the government spectrum graph is realized. In this way, the application extracts the real-time government web page based on the government spectrum graph ontology, which can accurately extract the knowledge required in the government spectrum graph ontology, avoid the extraction of redundant data, reduce the cost of processing redundant data, and improve the efficiency of constructing and updating the government spectrum graph.

[0068] Please refer to Figure 2 , Figure 2 is a schematic diagram of a knowledge extraction device based on ontology constraint provided by the application, which comprises:

[0069] The parsing unit 201 is used for parsing the government spectrum graph ontology, generating a triple file after parsing, and determining the first description information of each concept, attribute and relationship according to the triple file;

[0070] The acquisition unit 202 is configured to acquire an HTML file of a government affair webpage, generate a node tree according to the HTML file, and acquire second description information of leaf nodes in the node tree;

[0071] The generation unit 203 is configured to generate an extraction mapping model according to the first description information and the second description information;

[0072] The setting unit 204 is configured to set an extraction rule according to the extraction mapping model;

[0073] The processing unit 205 is configured to extract target knowledge in the node tree according to the extraction rule, and map the target knowledge to the government affair spectrum graph ontology.

[0074] In this embodiment, first, the parsing unit 201 parses the government affair spectrum graph ontology, generates a triple file after parsing, and determines first description information of each concept, attribute and relationship according to the triple file. Then, the acquisition unit 202 acquires an HTML file of a government affair webpage, generates a node tree according to the HTML file, and acquires second description information of leaf nodes in the node tree. Then, the generation unit 203 generates an extraction mapping model according to the first description information and the second description information, and the setting unit 204 sets an extraction rule from the extraction mapping model. Finally, the processing unit 205 extracts target knowledge in the node tree according to the extraction rule, and maps the target knowledge to the government affair spectrum graph ontology. In this way, knowledge extraction is directly based on the government affair spectrum graph ontology, the extracted knowledge is the knowledge required by the government affair spectrum graph ontology, the extraction of unnecessary redundant data is reduced, the cost of processing redundant data is reduced, and the efficiency of updating the government affair knowledge spectrum graph is improved.

[0075] Please refer to Figure 3 , Figure 3 Another schematic view of an ontology-constrained knowledge extraction device provided in the present application is provided, and the knowledge extraction device comprises:

[0076] The parsing unit 301 is configured to parse a government affair spectrum graph ontology, generate a triple file after parsing, and determine first description information of each concept, attribute and relationship according to the triple file;

[0077] The acquisition unit 302 is configured to acquire an HTML file of a government affair webpage, generate a node tree according to the HTML file, and acquire second description information of leaf nodes in the node tree;

[0078] The acquisition unit 302 comprises:

[0079] The acquisition module 3021 is configured to acquire HTML files corresponding to government affair service items, integrated items of a thing, government-citizen interaction and item lists, respectively;

[0080] The generating module 3022 is configured to generate a node tree according to the format tag of HTML in the HTML file

[0081] The generating unit 303 is configured to generate an extraction mapping model according to the first description information and the second description information.

[0082] The setting unit 304 is configured to set an extraction rule according to the first description information and the second description information.

[0083] The setting unit 304 includes:

[0084] The setting module 3041 is configured to set the extraction rule by the extraction mapping model in a manner of similarity calculation or rule mapping.

[0085] The processing unit 304 is configured to extract target knowledge in the node tree according to the extraction rule, and map the target knowledge to a government affair spectrum graph ontology.

[0086] The processing unit 305 includes:

[0087] The positioning module 3051 is configured to position the target knowledge in the node tree according to the extraction rule.

[0088] Please refer to Figure 4 , Figure 4 Another schematic diagram of the knowledge extraction device based on ontology constraint provided in the present application is shown in the figure, which includes:

[0089] The central processing unit 402, the memory 401, the input and output interface 403, the wired or wireless network interface 404 and the power supply 405.

[0090] The memory 401 is a transitory storage memory or a persistent storage memory.

[0091] The central processing unit 402 is configured to communicate with the memory 401, and execute the instruction operation in the memory 401 to perform the steps in the foregoing Figure 1 The embodiment shown in the figure.

[0092] The present application provides a computer readable storage medium, including instructions, when the instructions run on the computer, make the computer execute the steps in the foregoing Figure 1 The embodiment shown in the figure.

[0093] Those skilled in the art can clearly understand that, for the convenience and brevity of description, the specific working process of the system, device and unit described above can refer to the corresponding process in the foregoing method embodiment, which will not be repeated here.

[0094] In several embodiments provided in the present application, it should be understood that the disclosed system, device and method can be implemented in other manners. For example, the described device embodiments are merely schematic. For example, the division of the units is only a logical function division. There can be another division manner for the actual implementation, for example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual couplings or direct couplings or communication connections between different units, can be indirect couplings or communication connections through some interfaces, devices or units, and can be in electrical, mechanical or other forms.

[0095] The units described as separated components can or can not be physically separated, and the components displayed as units can or can not be physical units, i.e., can be located in one place, or can be distributed on a plurality of network units. Some or all of the units can be selected according to actual needs to achieve the purposes of the embodiments of the present application.

[0096] In addition, each functional unit in the embodiments of the present application can be integrated in one processing unit, or each unit can exist physically, or two or more units can be integrated in one unit. The integrated unit can be implemented in the form of hardware, or in the form of a software functional unit.

[0097] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer readable storage medium. Based on such an understanding, the technical solutions of the present application essentially, or the part that makes a contribution to the prior art, or all or a part of the technical solutions can be embodied in the form of a software product. The computer software product is stored in a storage medium, and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in the embodiments of the present application. The foregoing storage medium includes: a U disk, a mobile hard disk, a read-only memory (ROM, read-only memory), a random access memory (RAM, random access memory), a magnetic disk or an optical disk, and various other media that can store program codes.

Claims

1. An ontology constraint based knowledge extraction method, characterized in that, The method comprises the following steps: analyzing a government affair spectrum ontology, generating a triple file after the analysis, and determining first description information of each concept, attribute, and relationship according to the triple file; obtaining an HTML file of a government affair webpage, generating a node tree according to the HTML file, and obtaining second description information of leaf nodes in the node tree; generating an extraction mapping model according to the first description information and the second description information; setting an extraction rule according to the extraction mapping model; extracting target knowledge in the node tree according to the extraction rule, and mapping the target knowledge into the government affair spectrum ontology; the setting of the extraction rule according to the extraction mapping model comprises: the extraction mapping model sets the extraction rule through a similarity calculation mode.

2. The knowledge extraction method of claim 1, wherein, the extraction of the target knowledge in the node tree according to the extraction rule comprises: locating the target knowledge in the node tree according to the extraction rule.

3. The knowledge extraction method of claim 1, wherein, the obtaining of the HTML file of the government affair webpage comprises: respectively obtaining HTML files corresponding to government affair service items, integrated items of a thing, interaction between the government and the public, and a list of items.

4. The knowledge extraction method of claim 1, wherein, the generation of the node tree according to the HTML file comprises: generating the node tree according to the format tag of HTML in the HTML file.

5. An ontology constraint based knowledge extraction apparatus, characterized by, The method comprises the following steps: an analyzing unit is configured to analyze a government affair spectrum ontology, generate a triple file after the analysis, and determine first description information of each concept, attribute, and relationship according to the triple file; an obtaining unit is configured to obtain an HTML file of a government affair webpage, generate a node tree according to the HTML file, and obtain second description information of leaf nodes in the node tree; a generating unit is configured to generate an extraction mapping model according to the first description information and the second description information; a setting unit is configured to set an extraction rule according to the extraction mapping model; the setting unit comprises: a setting module is configured to set the extraction rule through a similarity calculation mode by the extraction mapping model; a processing unit is configured to extract target knowledge in the node tree according to the extraction rule, and map the target knowledge into the government affair spectrum ontology.

6. The knowledge extraction apparatus according to claim 5, characterized by the processing unit comprises: a locating module is configured to locate the target knowledge in the node tree according to the extraction rule.

7. An ontology constraint based knowledge extraction apparatus, characterized by, The device comprises: a central processing unit, a memory, an input and output interface, a wired or wireless network interface, and a power supply; the memory is a transitory storage memory or a persistent storage memory; the central processing unit is configured to communicate with the memory, and execute instruction operations in the memory to perform the method in any one of claims 1 to 4.

8. A computer readable storage medium comprising instructions which, when executed on a computer, cause the computer to perform the method of any one of claims 1 to 4.

Citation Information

Patent Citations

  • Chinese knowledge graph construction method and system

    CN108376160A

  • Semantic-based electric power measurement data processing method and device and computer equipment

    CN113641884A