Domain ontology construction method

By querying prompt words in multiple public big models, building and integrating domain candidate ontology diagrams, the problem of inaccurate top-level concept selection when building domain ontology from top to bottom is solved, and a more accurate and consistent domain ontology construction is achieved.

CN120045710AActive Publication Date: 2025-05-27LESHAN NORMAL UNIV
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510166337.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-27
Estimated Expiration
2045-02-14

AI Technical Summary

Technical Problem

When constructing a domain ontology from top to bottom, the choice of top-level concepts depends on the subjective judgment of domain experts, which may lead to inaccurate or comprehensive enough, affecting the accuracy and completeness of ontology construction.

Method used

By obtaining the prompt words for domain ontology construction, and querying them in the user interfaces of multiple public large models, obtaining feedback text sets, and text parsing to construct domain candidate ontology diagrams. The final domain ontology diagram is generated using directed graph merging algorithm and confidence evaluation.

Benefits of technology

Effectively utilize the knowledge capabilities of multiple large models, integrate the knowledge of multiple different large models, reduce dependence on domain expert knowledge, improve the accuracy and consistency of domain ontology, and avoid knowledge and axiom errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120045710A_ABST
    Figure CN120045710A_ABST
Patent Text Reader

Abstract

The invention discloses a domain ontology construction method, and relates to the technical field of natural language processing. Comprising the following steps: obtaining cue words, and querying in user interfaces provided by a plurality of public large models to obtain a domain ontology feedback text set; performing text analysis on each feedback text in the field ontology feedback text set to obtain field concepts, relationships and attribute vocabularies in the text; taking the domain concepts and the attribute vocabularies as nodes, and taking the relationship between the domain concepts and the relationship between the domain concepts and the attribute vocabularies as edges to construct domain candidate ontology graphs to obtain a domain candidate ontology graph set; combining the updated nodes and edges to the initialized empty graph with the weight to obtain a domain candidate ontology graph; and performing confidence evaluation on the domain candidate ontology graph, and selecting nodes and edges of which the confidence is greater than a set threshold in the graph to obtain a final domain ontology graph. According to the method, knowledge capabilities of a plurality of different public large models can be fully integrated, and domain expert knowledge does not need to be introduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a method for constructing a domain ontology. Background Art

[0002] Ontology is a standardized and formalized description of objective things. Constructing an ontology for a specific domain helps to acquire and describe knowledge in related fields, provide a common understanding of domain knowledge, and facilitate the sharing and reuse of domain knowledge. With the development and application of knowledge graphs, ontology, as part of the core architecture of knowledge graphs, has received widespread attention. The goal of constructing a domain ontology is to determine the entity types and attributes in the domain knowledge system, as well as the hierarchical structure and association relationships between them, and provide a standardized semantic framework for the organization of domain knowledge.

[0003] At present, the main approach to constructing domain ontology is from top to bottom. The specific implementation includes: first, led by a team of experts with deep domain knowledge, starting from the macro level, define the core concepts and top-level classifications of the domain to form the basic framework of the ontology; follow this top-level architecture and gradually refine the conceptual system, and through consulting a large amount of literature, data and expert interviews, deeply explore and define each sub-concept, attribute and the complex relationships between them; in order to ensure the accuracy and practicality of the ontology, internal reviews and external verifications are conducted regularly; after multiple rounds of iterations and improvements, a domain ontology with a clear structure, rich content and the ability to fully reflect the characteristics of the domain is constructed.

[0004] The starting point of building a domain ontology from top to bottom is to determine the top-level concepts, which often relies on the subjective judgment and experience of domain experts, resulting in inaccurate or incomplete selection of top-level concepts, affecting the accuracy and completeness of subsequent ontology construction. Summary of the invention

[0005] Based on this, it is necessary to provide a domain ontology construction method to address the above technical issues.

[0006] An embodiment of the present invention provides a method for constructing a domain ontology, including:

[0007] Get the prompt word Prompt for domain ontology construction;

[0008] The prompt word Prompt is used as the query text to query in the user interface provided by multiple public large models, and a domain ontology feedback text set Responses composed of the feedback text output by each large model is obtained; wherein the multiple public large models are represented as: LLMs = {m 1 ,m 2 ,...m i ,...,m K}, K is the total number of public large models, mi is the i-th publicly available large model; the domain ontology feedback text set is specifically Responses = {res 1 , res 2 ,... res i ,..., res K}}, where res i is the feedback text output by the i-th publicly available large model;

[0009] Perform text parsing on each feedback text in the domain ontology feedback text set Responses to obtain the domain concepts, relationships, and attribute vocabulary in the text; use the domain concepts and attribute vocabulary as nodes, and the relationships between domain concepts and the relationships between domain concepts and attribute vocabulary as edges to construct the domain candidate ontology graph canGraph i , and obtain the domain candidate ontology graph set CanGraphs;

[0010] Update the repetition degree and weight of each node and each edge in the domain candidate ontology graph set CanGraphs, and use the directed graph merging algorithm combined with similarity to merge the updated nodes and edges into the initialized weighted empty graph to obtain the domain candidate ontology graph MergedGraph;

[0011] Perform confidence evaluation on the domain candidate ontology graph MergedGraph, and select the nodes and edges with confidence greater than the set threshold in the graph to obtain the final domain ontology graph OntologyGraph.

[0012] Optionally, obtain the prompt Prompt for domain ontology construction, specifically including:

[0013] Decompose the domain ontology construction task into three information extraction subtasks including concept classification, relationship classification, and attribute classification;

[0014] Design output templates for the three information extraction subtasks respectively;

[0015] Select multiple publicly available large models, test the output of each publicly available large model and fine-tune the prompt, and select a set of K available large models with domain knowledge LLMs = {m 1 , m 2 ,... m i ,..., m K};

[0016] Determine the prompt Prompt according to the output templates of the information extraction subtasks and the large model set.

[0017] Optionally, construct the domain candidate ontology graph canGraph i , specifically including:

[0018] For each text res in Responses in sequence i Perform text parsing on it to obtain the set of concepts concepts i contained in the text res i , the set of attributes attributes i and the set of relations relations i ;

[0019] Based on the set of concepts concepts i , the set of attributes attributes i and the set of relations relations i , construct the domain candidate ontology graph canGraph i corresponding to the text res i ;

[0020] Add the domain candidate ontology graph to the set of domain candidate ontology graphs CanGraphs.

[0021] Optionally, for each text res in Responses in sequence i Perform text parsing, specifically including:

[0022] Initialize the set of concepts concepts i , the set of attributes attributes i and the set of relations relations i as empty sets;

[0023] Scan each line resSentence in the text res i in sequence, and judge the type of resSentence i,j according to the output template, where the type includes concept classification, relation classification or attribute classification; ij If resSentence i,j

[0024] belongs to the concept classification output template, initialize the set of concept words tmpConcepts as an empty set, extract the concept words in the text of resSentence i,j according to the output template and add them to the set tmpConcepts; scan each concept word tmpConcepts in the set of concept words tmpConcepts t in sequence, look up each word tmpConcepts i in the set of concepts concepts t , if it does not exist, add the word tmpConcepts t to the set of concepts concepts​i ; Taking the first concept term tmpConcepts in the set of concept terms tmpConcepts 1 as the subject, and all the remaining concept terms tmpConcepts t as the object, with the predicate or relationship type being include, respectively form relationship triples and add them to the relationship set relations i ;

[0025] If resSentence i,j belongs to the relationship classification output template, initialize the relationship set tmpRelations as an empty set, extract the relationship name relName and several relationship pairs from resSentence according to the output template and add them to the set tmpRelations; sequentially scan the relationship set tmpRelations, taking each relationship pair tmpRelations ij in the set as the subject and object with the two concept terms tmpConcept t and tmpConcept 1 respectively, and the predicate or relationship type being relName, form relationship triples and add them to the relationship set relations 2 ; i ;

[0026] If resSentence i,j belongs to the attribute classification output template, initialize the attribute set tmpAttributes as an empty set, extract the concept term conName and all attribute terms from resSentence according to the output template and add them to the set tmpAttributes; sequentially scan each attribute term tmpAttributes in the attribute term set tmpAttributes ij , add the term tmpAttributes t to the attribute node set attributes t ; i ;

[0027] Sequentially scan each attribute term tmpAttributes in the attribute term set tmpAttributes t , taking the concept term conName as the subject and the attribute term tmpAttributes t as the object, with the predicate or relationship type being hasAttribute, respectively form relationship triples and add them to the relationship set relations i ;

[0028] Repeat the iteration until the text resi until all text lines are processed.

[0029] Optionally, construct the text res i The corresponding domain candidate ontology graph canGraph i , specifically including:

[0030] Initialize the domain candidate ontology graph canGraph i as an empty graph, with the node set nodes i and the edge set edges i both being empty sets;

[0031] Scan each concept concepts i in the concept set concepts i,j , and add a node to the node set nodes i of the domain candidate ontology graph canGraph i , with the node label being concepts i,j and the type being conType;

[0032] Scan each attribute attributes i in the attribute set attributes i,j , and add a node to the node set nodes i of the domain candidate ontology graph canGraph i , with the node label being attributes i,j and the type being attType;

[0033] Scan each relation relations i in the relation set relations i,j , and assume its corresponding relation triple is <head, relType, tail>. Find the corresponding concept node or attribute node headNode and tailNode of head and tail in the graph canGraph i , and add a directed edge from headNode to tailNode to the edge set edges i of the domain candidate ontology graph canGraph i , with the edge type being relType.

[0034] Optionally, obtain the domain candidate ontology graph MergedGraph, specifically including:

[0035] Initialize the weighted domain candidate ontology graph MergedGraph as an empty graph, with the node set Nodes and the edge set Edges both being empty sets;

[0036] Copy the first directed graph canGraph in the set CanGraphs 1 to MergedGraph. The node set Nodes of MergedGraph = nodes 1 , and the edge set Edges = edges 1 . The multiplicity and weight of each node and each edge in the set are initialized to 1;

[0037] Scan each remaining directed graph CanGraph in the set CanGraphs in turn i , and use the node merging algorithm that combines similarity to merge the nodes in CanGraph i into MergedGraph, and update the multiplicity and weight of the nodes;

[0038] Scan each remaining directed graph CanGraph in the set CanGraphs in turn i , and use the edge merging algorithm that combines similarity to merge the edges in CanGraph i into MergedGraph, and update the multiplicity and weight of the edges.

[0039] Optionally, update the multiplicity and weight of the nodes, specifically including:

[0040] Scan each node tmpNode in the domain candidate ontology graph CanGraph i in turn, and obtain the node label tmpNodeTag and the node type tmpNodeType;

[0041] Search for the node sameNode with the same node label as tmpNode or the most similar node similarNode in the node set Nodes of MergedGraph. Among them, the most similar node similarNode is obtained by using the similarity calculation method of vocabulary. Taking the node label as the vocabulary, the similarity value similarity of the node similarNode > θ and the value is the largest, where θ is the preset similarity threshold;

[0042] If the same node sameNode exists, judge whether the node types of tmpNode and sameNode are also the same; if they are the same, update the multiplicity and weight of sameNode in the MergedGraph, and add 1 to both the multiplicity and weight; if they are different, add the node tmpNode to the node set Nodes of MergedGraph, and set both the multiplicity and weight of the node to 1;

[0043] If the same node sameNode does not exist but the most similar node similarNode exists, determine whether the types of tmpNode and similarNode are also the same; if they are the same, update the repetition degree and weight of similarNode in the MergedGraph. The repetition degree is incremented by 1, and the weight is increased by the similarity value similarity; update the node label of tmpNode in CanGraph to be the node label of similarNode; if they are not the same, add the node tmpNode to the node set Nodes of MergedGraph, and set both the repetition degree and weight of the node to 1. i If both sameNode and similarNode do not exist, add the node tmpNode to the node set Nodes of MergedGraph, and set both the repetition degree and weight of the node to 1.

[0044] Repeat the iteration until all nodes in the domain candidate ontology graph CanGraph

[0045] i are processed.

[0046] Optionally, update the repetition degree and weight of the edges, specifically including:

[0047] Scan the edges tmpEdge in the domain candidate ontology graph CanGraph in sequence, obtain the subject and object node labels tmpHeadTag and tmpTailTag attached to the edge, and the edge type tmpEdgeType; i

[0048] Find the edges in the edge set Edges of MergedGraph that have the same subject and object node labels as tmpEdge, and form a set sameHeadTails;

[0049] If no edge exists in the sameHeadTails set, add the edge tmpEdge to the edge set Edges of MergedGraph, and set both the repetition degree and weight of the edge to 1;

[0050] If an edge exists in the sameHeadTails set, find the edge sameEdge or the most similar edge similarEdge with the same edge type as tmpEdge in this set. The most similar edge similarEdge is obtained by using the similarity calculation method of vocabulary, with the edge type as the vocabulary; the similarity value similarity of the edge similarEdge > λ and is the largest, where λ is a preset similarity threshold;

[0051] ​​If the same edge sameEdge exists, update the repeatability and weight of sameEdge in the MergedGraph, both the repeatability and weight are incremented by 1;

[0052] If the same edge sameEdge does not exist but the most similar edge similarEdge exists, update the repeatability and weight of similarEdge in the MergedGraph, the repeatability is incremented by 1, and the weight is incremented by the similarity value similarity;

[0053] If neither sameNode nor similarNode exists, add the edge tmpEdge to the edge set Edges of MergedGraph, and both the repeatability and weight of the edge are set to 1;

[0054] Repeat the iteration until all edges in the domain candidate ontology graph CanGraph i are processed.

[0055] Optionally, obtain the final domain ontology graph OntologyGraph, specifically including:

[0056] Initialize the domain ontology graph OntologyGraph as an empty graph;

[0057] For each node Nodes in the node set Nodes of the merged domain candidate ontology graph MergedGraph i Calculate its confidence nodeConfidence i =(repeatability i *weight i ) / K 2 , where repeatability i is the repeatability of the node Nodes i , weight i is the weight value of the node Nodes i , and K is the number of large models;

[0058] For each edge Edges in the edge set Edges of the graph MergedGraph i Calculate its confidence edgeConfidence j =nodeConfidence head *nodeConfidence tail *(repeatability j *weight j ) / K 2 , where repeatability jFor the edges j The multiplicity, weight j For the edges j The weight value, K is the number of large models, nodeConfidence head For the edges j The confidence of the corresponding main body node, nodeConfidence tail For the edges j The confidence of the corresponding object node;

[0059] Sort all the nodes in the node set Nodes in descending order of confidence, and scan each node in the node set Nodes in descending order of confidence; if the confidence requirement is met, add the node, select the edges and adjacent nodes that meet the confidence requirement attached to the node, and add the nodes and edges that meet the confidence requirement to the final domain ontology graph OntologyGraph.

[0060] Optionally, add the nodes and edges that meet the confidence requirement to the final domain ontology graph OntologyGraph, specifically including:

[0061] Scan each node tmpNode in the node set Nodes in descending order of confidence, and determine whether its confidence is greater than the set threshold α;

[0062] If not, the construction of the domain ontology graph OntologyGraph ends;

[0063] If it holds, first add the node tmpNode to the domain ontology graph OntologyGraph; then scan each edge tmpEdge attached to the node tmpNode in turn, and determine whether its confidence is greater than the set threshold β; if it holds, add the edge tmpEdge and the other node attached to the edge to the domain ontology graph OntologyGraph;

[0064] Repeat the iteration until all the edges attached to the node tmpNode are scanned;

[0065] Continue to scan the next node in the node set Nodes until all the nodes that meet the confidence requirement are scanned, and the final domain ontology graph OntologyGraph is composed of all the added nodes and edges.

[0066] The above domain ontology construction method provided by the embodiments of the present invention has the following beneficial effects compared with the prior art:

[0067] The present invention constructs a domain ontology based on the knowledge of general large models, forms a think tank by introducing multiple large models, and uses a directed graph merging algorithm to integrate the knowledge of multiple experts to achieve the goal of complementing each other's strengths in the think tank. It can effectively utilize the knowledge breadth and natural language understanding ability of publicly available large models, fully integrate the knowledge capabilities of multiple different publicly available large models, and does not require the introduction of domain expert knowledge.

[0068] In addition, a confidence evaluation is performed on the domain candidate ontology graph MergedGraph, and the nodes and edges in the graph with a confidence greater than the set threshold are selected to analyze the consistency and accuracy of the domain ontology knowledge, thereby avoiding knowledge-based and axiomatic errors in the domain ontology. Brief Description of the Drawings

[0069] Figure 1 It is a flowchart of a domain ontology construction method provided in an embodiment. Detailed Embodiment

[0070] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0071] First, the domain knowledge comes from the results generated by general large models, and the accuracy and authority of the data can be guaranteed to a certain extent. Secondly, for the concepts and relationships involved in the domain ontology, the degree of lexical abstraction is high, and the knowledge breadth of the large model can better cover various ontologies such as the top layer, common sense, and tasks involved in the construction of the domain ontology. Finally, the reasoning and thinking abilities possessed by the large model can effectively decompose and execute the complex task of domain ontology construction.

[0072] In one embodiment, a domain ontology construction method is provided, and the method includes:

[0073] 1. Design a domain ontology construction prompt Prompt.

[0074] 2. Use the prompt as the query text and execute the query in the user interfaces provided by several publicly available large models to obtain the corresponding feedback text of the large models; the large model set is represented as: LLMs = {m 1 , m 2 ,... m i ,..., m K}, m i is the i-th (1 ≤ i ≤ K) large model, K is the number of selected large models, and the domain ontology feedback text set is represented as: Responses = {res 1 , res 2 ,... res i,...,res K}, res i is the text output for model m i .

[0075] 3. For each text res in the large model feedback text set Responses i perform text parsing to obtain domain concepts, relationships, and attribute vocabulary in the text. Using domain concepts and attribute vocabulary as nodes, and using the relationships between domain concepts and the relationships between domain concepts and attribute vocabulary as edges, construct a domain candidate ontology graph canGraph i , thus forming a set of domain candidate ontology graphs CanGraphs.

[0076] 4. Update the repetition degree and weight of each node and each edge in the set of domain candidate ontology graphs CanGraphs, and use a directed graph merging algorithm that combines similarity to merge the updated nodes and edges into an initialized weighted empty graph to obtain a domain candidate ontology graph MergedGraph.

[0077] 5. Perform confidence evaluation on the merged domain candidate ontology graph MergedGraph, select the nodes and edges in the graph with confidence greater than the set threshold, and construct the final domain ontology graph OntologyGraph.

[0078] Example 1

[0079] The present invention provides a method for constructing a domain ontology based on a large model. By introducing multiple large models as "domain experts", the natural language understanding and knowledge capabilities of the large models are used to assist in obtaining domain ontology knowledge. The example takes scenic area tourism as the target domain, and uses domestic mainstream large models such as Zhipu Qingyan, Wenxin Yiyan, ByteDance Doubao, and iFlytek Spark as "domain experts", and designs prompts with a chain of thought to obtain large model text feedback. Based on these text feedbacks, several candidate domain ontology graphs are constructed, and algorithms such as semantic similarity calculation, graph merging, and confidence evaluation are used to integrate and evaluate the large model domain ontology knowledge to generate the final domain ontology, as Figure 1 shown, including the following steps:

[0080] S1. Design a domain ontology construction prompt Prompt, and the specific steps are as follows:

[0081] S11. Decompose the domain ontology construction task into three information extraction subtasks: concept (entity type) classification, relationship (entity relationship) classification, and attribute (entity attribute relationship) classification;

[0082] S12. Design output templates for the three information extraction subtasks, and their corresponding indicator words are as follows:

[0083] Concept classification subtask output template:

[0084] List the scenic spot tourism ontology vocabulary by category in the following format:

[0085] <Concept category 1> includes <Concept category 2>, <Concept category 3>, etc.

[0086] <Concept category 2> includes <Concept vocabulary 1>, <Concept vocabulary 2>, etc.

[0087] Relationship classification subtask output template:

[0088] List the scenic spot tourism ontology relationships by category in the following format:

[0089] <Relationship category 1>: (<Concept category 1>, <Concept category 2>), etc.

[0090] <Relationship category 2>: (<Concept vocabulary 1>, <Concept vocabulary 2>), etc.

[0091] Attribute classification subtask output template:

[0092] Please list the attributes of each concept in scenic spot tourism in the following format:

[0093] The attributes of <Concept category 1> include <Concept attribute 1>, <Concept attribute 2>, etc.

[0094] The attributes of <Concept vocabulary 1> include <Concept attribute 3>, <Concept attribute 4>, etc.

[0095] S13. Select several publicly available general large models, test the outputs of the large models and fine-tune the prompting words, and select a set of K available large models with domain knowledge LLMs = {m 1 , m 2 ,... m i ,..., m K}, and combine with the subtask output template to determine the prompting word Prompt for domain ontology construction, such as:

[0096] Please give the scenic spot tourism ontology structure according to the following requirements, and it is required to add general ontologies such as time, location, and person without examples.

[0097] First step, list the scenic spot tourism ontology vocabulary by category in the following format:

[0098] <Concept category 1> includes <Concept category 2>, <Concept category 3>, etc.

[0099] <Concept category 2> includes <Concept vocabulary 1>, <Concept vocabulary 2>, etc.

[0100] Second step, list the scenic spot tourism ontology relationships by category in the following format:

[0101] <Relationship Category 1>: (<Concept Category 1>, <Concept Category 2>) etc.

[0102] <Relationship Category 2>: (<Concept Vocabulary 1>, <Concept Vocabulary 2>) etc.

[0103] <Relationship Category 3>: (<Concept Category 3>, <Concept Vocabulary 3>) etc.

[0104] <Relationship Category 3>: (<Concept Vocabulary 4>, <Concept Category 4>) etc.

[0105] Step 3: Please list the attributes of each concept category in the following format:

[0106] The attributes of <Concept Category 1> include <Concept Attribute 1> <Concept Attribute 2> etc.

[0107] The attributes of <Concept Vocabulary 1> include <Concept Attribute 3> <Concept Attribute 4> etc.

[0108] S2. Use the prompt as the query text and execute the query in the user interfaces provided by several public large models to obtain the corresponding feedback text from the large models; the set of large models is represented as: LLMs = {m 1 , m 2 ,... m i ,..., m K}, where m i is the i-th (1 ≤ i ≤ K) large model, and K is the number of selected large models. The set of domain ontology feedback texts is represented as: Responses = {res 1 , res 2 ,... res i ,..., res K}, where res i is the text output of model m i ; for example, after querying ByteDance Doubao, some of the obtained feedback texts are:

[0109] <Scenic Area> includes <Natural Landscape Scenic Area>, <Cultural Landscape Scenic Area>, <Comprehensive Scenic Area>, etc.

[0110] <Scenic Spot> includes <Natural Scenic Spot>, <Cultural Scenic Spot>, <Entertainment Scenic Spot>, etc.

[0111] S3. Parse each text res i in the large model feedback text set Responses to construct the corresponding domain candidate ontology graph canGraph i (where the nodes and edges in the graph correspond to concepts and relationships respectively), thus forming the domain candidate ontology graph set CanGraphs. The specific steps include:

[0112] S31. For each text res in Responses in sequence i perform text parsing to obtain the concept set concepts i contained in the text res i , the attribute set attributes i and the relationship set relations i , and the specific steps include:

[0113] S311. Initialize the concept set concepts i , the attribute set attributes i and the relationship set relations i as empty sets;

[0114] S312. Scan each line resSentence in the text res in sequence i , and judge the type of resSentence according to the output template, and the types include concept classification, relationship classification or attribute classification; i,j ij i,j

[0115] S313. If resSentence belongs to the concept classification output template, initialize the concept vocabulary set tmpConcepts as an empty set, and extract the concept vocabulary in the text of resSentence according to the output template and add it to the set tmpConcepts; i,j t i t t

[0116] S314. Scan each concept vocabulary tmpConcepts in the concept vocabulary set tmpConcepts in sequence t , search for each vocabulary tmpConcepts in the concept set concepts i , if it does not exist, add the vocabulary tmpConcepts t t to the concept set concepts i 1 ;

[0117] S315. Take the first concept vocabulary tmpConcepts in the concept vocabulary set tmpConcepts as the subject, and the remaining all concept vocabularies tmpConcepts 1 t as the object, and the predicate or relationship type is include, and respectively form a relationship triple (that is, <tmpConcepts 1 , include, tmpConcepts i >) and add it to the relationship set relations​i ; For example, in the text "<scenic area> includes <natural landscape scenic area>, <human landscape scenic area>, <comprehensive scenic area>, etc.", "scenic area", "natural landscape scenic area", and "human landscape scenic area" are all conceptual words. Among them, the concept "scenic area" is the subject, and the "natural landscape scenic area" is the object, forming a relationship triple "scenic area, include, natural landscape scenic area".

[0118] S316. If resSentence i,j belongs to the relationship classification output template, initialize the relationship set tmpRelations as an empty set, and extract the relationship name relName and several relationship pairs from resSentence according to the output template ij and add them to the set tmpRelations;

[0119] S317. Scan the relationship set tmpRelations in turn, and use each relationship pair tmpRelations t in the set as the subject and object of two conceptual words tmpConcept 1 and tmpConcept 2 respectively, and the predicate or relationship type is relName, forming a relationship triple (i.e., <tmpConcept 1 , relName, tmpConcept 2 >) and add it to the relationship set relations i , for example: in the text "<subordination relationship>: (<tourism organization management category>, <travel agency>), (<travel agency>, <tour guide>), (<scenic area management department>, <operation management section>)", there is a "subordination" relationship between "travel agency" and "tour guide", so a triple <travel agency, subordination, tour guide> can be formed;

[0120] S318. If resSentence i,j belongs to the attribute classification output template, initialize the attribute set tmpAttributes as an empty set, and extract the conceptual word conName and all attribute words from resSentence according to the output template ij and add them to the set tmpAttributes;

[0121] S319. Scan each attribute word tmpAttributes in the attribute word set tmpAttributes in turn t , and add the word tmpAttributes t to the attribute node set attributes i ;

[0122] S3110. Sequentially scan each attribute term tmpAttributes in the attribute term set tmpAttributes t , taking the concept term conName as the subject, the attribute term tmpAttributes t as the object, and the predicate or relationship type as hasAttribute, respectively form relationship triples (i.e., <conName, hasAttribute, tmpNode i >) and add them to the relationship set relations i . For example: "The attributes of <scenic area> include <scenic area area>, <number of scenic spots>, <landscape features>, <ticket price>, <opening policy>, etc.", "scenic area area" and "number of scenic spots", etc. are all attributes of the concept "scenic area", so triples such as "scenic area, hasAttribute, scenic area area" can be formed.

[0123] S3111. Repeat the above steps until all text lines in the text res i are processed.

[0124] S32. Construct the domain candidate ontology graph canGraph i corresponding to the text res i from the concept set concepts i , the attribute set attributes i and the relationship set relations i . The specific steps include:

[0125] S321. Initialize the domain candidate ontology graph canGraph i as an empty graph, that is, its node set nodes i and edge set edges i are both empty sets;

[0126] S322. Scan each concept concepts i in the concept set concepts i,j , add a node to the node set nodes i of the domain candidate ontology graph canGraph i , and the label of this node is concepts i,j , and the type is conType;

[0127] S323. Scan each attribute attributes i in the attribute set attributes i,j , in the domain candidate ontology graph canGraph iThe node set nodes i Add a node to it, and the label of this node is attributes i,j , and the type is attType;

[0128] S324. Scan the relationship set relations i For each relationship relations in it i,j , assume its corresponding relationship triple is <head, relType, tail>. Respectively find the concept node or attribute node headNode and tailNode corresponding to head and tail in the graph canGraph i , and in the domain candidate ontology graph canGraph i Add a directed edge from headNode to tailNode to the edge set edges i , and the type of the edge is relType.

[0129] S33. Add the domain candidate ontology graph to the domain candidate ontology graph set CanGraphs.

[0130] S4. Use the directed graph merging algorithm combined with similarity to merge the domain candidate ontology graph set CanGraphs to construct the merged domain candidate ontology graph MergedGraph. The specific steps are as follows:

[0131] S41. Initialize the weighted domain candidate ontology graph MergedGraph as an empty graph, that is, both its node set Nodes and edge set Edges are empty sets;

[0132] S42. Copy the first directed graph canGraph in the set CanGraphs 1 to MergedGraph, that is, the node set Nodes of MergedGraph = nodes 1 , and the edge set Edges = edges 1 , and the repeatability and weight of each node and each edge in the set are both initialized to 1;

[0133] S43. Sequentially scan each remaining directed graph CanGraph in the set CanGraphs i , and use the node merging algorithm combined with similarity to merge the nodes in CanGraph i into MergedGraph, and update the repeatability and weight of the nodes. The specific steps are as follows:

[0134] S431. Sequentially scan the domain candidate ontology graph CanGraph iFor each node tmpNode in it, obtain the node label tmpNodeTag and the node type tmpNodeType;

[0135] S432. Search for the node sameNode with the same node label as tmpNode or the most similar node similarNode in the node set Nodes of MergedGraph, where the most similar node similarNode is obtained by using the lexical similarity calculation method (taking the node label as the vocabulary), and the similarity value similarity of the node similarNode is > θ and the value is the largest (θ is the preset similarity threshold);

[0136] S433. If the same node sameNode exists, determine whether the node types of tmpNode and sameNode are also the same. If they are the same, update the repetition degree and weight of sameNode in the MergedGraph (that is, both the repetition degree and the weight are incremented by 1); if they are different, add the node tmpNode to the node set Nodes of MergedGraph, and both the repetition degree and the weight of the node are set to 1;

[0137] S434. If the same node sameNode does not exist but the most similar node similarNode exists, determine whether the types of tmpNode and similarNode are also the same. If they are the same, update the repetition degree (that is, the repetition degree is incremented by 1) and weight (that is, the weight is incremented by the similarity value similarity) of similarNode in the MergedGraph, and update the node label of tmpNode in CanGraph i to be the node label of similarNode; if they are different, add the node tmpNode to the node set Nodes of MergedGraph, and both the repetition degree and the weight of the node are set to 1;

[0138] S435. If neither sameNode nor similarNode exists, add the node tmpNode to the node set Nodes of MergedGraph, and both the repetition degree and the weight of the node are set to 1;

[0139] S436. Repeat the above process until all nodes in the domain candidate ontology graph CanGraph i have been processed.

[0140] S44. Sequentially scan each remaining directed graph CanGraph in the set CanGraphs i , and use the edge merging algorithm combined with similarity to merge CanGraph iMerge the edges in [Graph Name] into MergedGraph, and update the multiplicity and weight of the edges. The specific steps are as follows:

[0141] S441. Sequentially scan the edges tmpEdge in the domain candidate ontology graph CanGraph i to obtain the subject and object node labels tmpHeadTag and tmpTailTag attached to the edge, as well as the edge type tmpEdgeType;

[0142] S442. In the edge set Edges of MergedGraph, find the edges with the same subject and object node labels as tmpEdge, and form a set sameHeadTails (multiple relationships between the subject and object are allowed);

[0143] S443. If there is no edge in the sameHeadTails set, add the edge tmpEdge to the edge set Edges of MergedGraph, and set the multiplicity and weight of the edge to 1;

[0144] S444. If there is an edge in the sameHeadTails set, find the edge sameEdge or the most similar edge similarEdge with the same edge type as tmpEdge in this set. The most similar edge similarEdge is obtained by using the lexical similarity calculation method (with the edge type as the vocabulary), and the similarity value similarity of the edge similarEdge > λ and is the largest (λ is a preset similarity threshold);

[0145] S445. If the same edge sameEdge exists, update the multiplicity and weight of sameEdge in the MergedGraph (that is, both the multiplicity and weight are incremented by 1);

[0146] S446. If the same edge sameEdge does not exist but the most similar edge similarEdge exists, update the multiplicity (that is, the multiplicity is incremented by 1) and weight (that is, the weight is incremented by the similarity value similarity) of similarEdge in the MergedGraph;

[0147] S447. If neither sameNode nor similarNode exists, add the edge tmpEdge to the edge set Edges of MergedGraph, and set the multiplicity and weight of the edge to 1;

[0148] S448. Repeat the above process until all the edges in the domain candidate ontology graph CanGraph i are processed.

[0149] S5. Evaluate the confidence of the merged domain candidate ontology graph MergedGraph, and select the nodes and edges in the graph with confidence greater than the set threshold to construct the final domain ontology graph OntologyGraph. The specific steps are as follows:

[0150] S51. Initialize the domain ontology graph OntologyGraph as an empty graph;

[0151] S52. For each node Nodes in the node set Nodes of the merged domain candidate ontology graph MergedGraph i Calculate its confidence nodeConfidence i =(repeatability i *weight i ) / K 2 , where repeatability i is the repeatability of the node Nodes i , weight i is the weight value of the node Nodes i , and K is the number of large models;

[0152] S53. For each edge Edges in the edge set Edges of the graph MergedGraph i Calculate its confidence edgeConfidence j =nodeConfidence head *nodeConfidence tail *(repeatability j *weight j ) / K 2 , where repeatability j is the repeatability of the edge Edges j , weight j is the weight value of the edge Edges j , K is the number of large models, nodeConfidence head is the confidence of the subject node corresponding to the edge Edges j , and nodeConfidence tail is the confidence of the object node corresponding to the edge Edges j ;

[0153] S54. Sort all the nodes in the node set Nodes in descending order of confidence;

[0154] S55. Scan each node in the node set Nodes in descending order of confidence. If the confidence requirement is met, add the node, select the edges and adjacent nodes that meet the confidence requirement attached to this node, and add these nodes and edges to the final domain ontology graph OntologyGraph. The specific steps are as follows:

[0155] S551. Scan each node tmpNode in the node set Nodes in descending order of confidence, and judge whether its confidence is greater than the set threshold α;

[0156] S552. If not, the construction of the domain ontology graph OntologyGraph ends;

[0157] S553. If it holds, first add the node tmpNode to the domain ontology graph OntologyGraph;

[0158] S554. Then scan each edge tmpEdge attached to the node tmpNode in turn, and judge whether its confidence is greater than the set threshold β;

[0159] S555. If it holds, add the edge tmpEdge and the other node attached to this edge to the domain ontology graph OntologyGraph;

[0160] S556. Repeat the above two steps until all the edges attached to the node tmpNode are scanned;

[0161] S557. Continue to scan the next node in the node set Nodes until all the nodes that meet the confidence requirement are scanned;

[0162] S558. All the nodes and edges added in the above process together form the final domain ontology graph OntologyGraph.

[0163] The present invention provides a method for constructing a domain ontology based on a large model. First, construct the domain ontology based on the knowledge of the general large model. The knowledge breadth and migration ability of the large model provide a strong guarantee for the accuracy and recall rate of the domain ontology knowledge. Second, design the model prompt words in the way of the thought chain, and effectively decompose and execute the task of constructing the domain ontology by means of the natural language understanding and reasoning ability of the large model. Third, form a "think tank" by introducing multiple large model "domain experts", and adopt a directed graph merging algorithm to realize the integration of multi-expert knowledge, so as to achieve the goal of "learning from each other's strengths" of the think tank. Finally, analyze the consistency and accuracy of the domain ontology knowledge by means of semantic similarity calculation and confidence evaluation, etc., so as to avoid possible knowledge-based and axiomatic errors in the domain ontology.

[0164] The above-described embodiments merely represent several implementation manners of the present invention. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all fall within the protection scope of the present invention.

Claims

1. A method for constructing a domain ontology, characterized in that: include: Get hint words for domain ontology construction ; The prompt word As query text, query in the user interface provided by multiple public big models to obtain the domain ontology feedback text set composed of the feedback text output by each big model ; Wherein, the multiple public large models are represented as: , is the total number of public models, For the The domain ontology feedback text collection is specifically: , For the A feedback text of the public large model output; Feedback text collection for domain ontology Each feedback text in the text is parsed to obtain the domain concepts, relations and attribute words in the text; the domain candidate ontology graph is constructed with domain concepts and attribute words as nodes and the relations between domain concepts and the relations between domain concepts and attribute words as edges. , get the domain candidate ontology graph set ; Domain candidate ontology graph set The duplication and weight of each node and edge in the domain are updated, and the updated nodes and edges are merged into the initialized weighted empty graph using the directed graph merging algorithm combined with similarity to obtain the domain candidate ontology graph. ; Domain Candidate Ontology Graph Perform confidence assessment, select nodes and edges in the graph whose confidence is greater than the set threshold, and obtain the final domain ontology graph .

2. A method for constructing a domain ontology according to claim 1, characterized in that: The prompt words for obtaining domain ontology construction , specifically including: The domain ontology construction task is decomposed into three information extraction subtasks including concept classification, relationship classification and attribute classification; Design output templates for three information extraction subtasks respectively; Select multiple public models, test each public model output and fine-tune the prompt words, select A large collection of models available with domain knowledge ; Determine the prompt words based on the output template of the information extraction subtask and the large model set .

3. A method for constructing a domain ontology according to claim 1, characterized in that: The construction of a candidate domain ontology graph , specifically including: Sequentially Each text in Perform text parsing to obtain text The concept set contained in , attribute collection and relationship set ; According to the concept collection , attribute collection and relationship set , construct the text Corresponding domain candidate ontology graph ; Add the domain candidate ontology graph to the domain candidate ontology graph collection .

4. A method for constructing a domain ontology as claimed in claim 3, characterized in that: The following Each text in Perform text parsing, including: Initialize concept set , attribute collection and relationship set is an empty set; Scan the text sequentially Each row , according to the output template The type includes a concept classification, a relationship classification or an attribute classification; like Belongs to the concept classification output template, then initialize the concept vocabulary set Empty set, extracted according to output template Concept words in the text and add to the collection ; Scan the concept vocabulary set in sequence Each concept word in , in the concept set Look up each word , if it does not exist, then the word Add to concept collection ; Concept vocabulary collection The first concept word As the main body, all other conceptual words is the object, and the predicate or relation type is , respectively forming relation triples and adding them to the relation set ; like Belongs to the relation classification output template, then initialize the relation set Empty set, extracted according to output template Relation names in text and several relationship pairs and add them to the collection ; Scan the relation set in sequence , for each relation pair in the set Two Concept Vocabularies and are the subject and object respectively, and the predicate or relation type is , forming a relation triple and adding it to the relation set ; like Belongs to the attribute classification output template, then initialize the attribute set Empty set, extracted according to output template Conceptual vocabulary in the text and all attribute words and add to the collection ; Scan the attribute vocabulary set in sequence Each attribute word in , the vocabulary Add to attribute node collection ; Scan the attribute vocabulary set sequentially Each attribute word in , with conceptual vocabulary as the subject, attribute vocabulary is the object, and the predicate or relation type is , respectively forming relation triples and adding them to the relation set ; Repeat until the text Until all text lines in are processed.

5. A method for constructing a domain ontology as claimed in claim 3, characterized in that: The constructed text Corresponding domain candidate ontology graph , specifically including: Initialize the domain candidate ontology graph Empty graph, node set and edge set are all empty sets; Scanning concept set Each concept , in the domain candidate ontology graph The node set Add a node with the label , type ; Scan attribute collection Each attribute , in the domain candidate ontology graph The node set Add a node with the label , type ; Scanning a Relationship Collection Each relationship in , let its corresponding relation triple be , respectively in Figure Found and Corresponding concept node or attribute node and , in the domain candidate ontology graph The edge set Add a line from arrive The directed edge of .

6. A method for constructing a domain ontology according to claim 1, characterized in that: The candidate ontology graph of the domain is obtained , specifically including: Initialize the weighted domain candidate ontology graph Empty graph, node set and edge set are all empty sets; Will gather The first directed graph in Copy to , The node set , edge set , the repetition and weight of each node and each edge in the set are initialized to 1; Scan the collection sequentially Each of the remaining directed graphs in , using the node merging algorithm combined with similarity The nodes in are merged into , and update the duplication and weight of the nodes; Scan the collection sequentially Each of the remaining directed graphs in , using the edge merging algorithm combined with similarity The edges in are merged into , and update the edge duplication and weight.

7. A method for constructing a domain ontology as claimed in claim 6, characterized in that: The repetition and weight of the updating node specifically include: Scan the domain candidate ontology graph sequentially Each node in , get the node label and node type ; exist The node set Find and Nodes with the same node label or most similar node , where the most similar node The similarity calculation method of vocabulary is used to obtain the node label. The similarity value of And the value is the largest, is a preset similarity threshold; If the same node Existence, judgment and Are the node types of the same? If they are the same, update In the picture The repetition and weight of the repetition and weight are increased by 1; if they are different, The node set Add a node , the duplication and weight of the nodes are set to 1; If the same node Does not exist but the most similar node exists ,judge and Are the types of the same? If they are the same, update In the picture The repetition and weight of the repetition is 1, and the weight is added with the similarity value. ;renew middle The node label is Node labels; if different, then The node set Add a node , the duplication and weight of the nodes are set to 1; like and If neither exists, then The node set Add a node , the duplication and weight of the nodes are set to 1; Repeat the iteration until the domain candidate ontology graph All nodes in the process are processed.

8. A method for constructing a domain ontology according to claim 6, characterized in that: The update edge repetition and weight specifically include: Scan the domain candidate ontology graph sequentially The edge , get the subject and object node labels to which the edge is attached and , and the type of edge ; exist The edge set Find and The subject and object node labels of the same edge form a set ; like There is no edge in the set, then The edge set Add Edges , the edge duplication and weight are set to 1; like If there is an edge in the set, then find the edge in the set that matches Edges of the same type or most similar edge , where the most similar edge The similarity calculation method of words is used to obtain the edge type, which is the word; the edge The similarity value of And the value is the largest, is a preset similarity threshold; If the same side If exists, update In the picture The repetition and weight of the , repetition and weight are both increased by 1; If the same side Does not exist but the most similar edge exists , then update In the picture The repetition and weight of the repetition is 1, and the weight is added with the similarity value. ; like and If neither exists, then The edge set Add Edges , the edge duplication and weight are set to 1; Repeat the iteration until the domain candidate ontology graph All edges in are processed.

9. A method for constructing a domain ontology according to claim 1, characterized in that: The final domain ontology graph is obtained , specifically including: Initialize the domain ontology graph is an empty graph; The merged domain candidate ontology graph The node set Each node in Calculate its confidence ,in For Node The repetition rate, For Node The weight value of is the number of large models; Pair The edge set Each edge in Calculate its confidence ,in For edge The repetition rate, For edge The weight value of is the number of large models, For edge The confidence of the corresponding subject node, For edge The confidence of the corresponding object node; Node Set All nodes in the are sorted from high to low by confidence, and the node set is scanned in order from high to low by confidence. Each node in the domain ontology; if it meets the confidence requirement, add the node, select the edges and adjacent points attached to the node that meet the confidence requirement, and add the nodes and edges that meet the confidence requirement to the final domain ontology graph .

10. A method for constructing a domain ontology according to claim 9, characterized in that: The nodes and edges that meet the confidence requirements are added to the final domain ontology graph , specifically including: Scan the node set from high to low confidence Each node in , determine whether its confidence is greater than the set threshold ; If not, then the domain ontology diagram The construction is finished; If established, first in the domain ontology diagram Add a node ; Then scan the nodes attached to it in turn Each edge of , determine whether its confidence is greater than the set threshold ; If true, then in the domain ontology graph Add Edges and the other node to which the edge is attached; Repeat the iteration until it is attached to the node All edge scans are completed; Continue scanning the node set The next node is added until all nodes that meet the confidence requirements are scanned. All the added nodes and edges together form the final domain ontology graph. .

Citation Information

Patent Citations

  • Concept map construction method and device, computer equipment and storage medium

    CN112395391A

  • Knowledge graph automatic expansion method for constructing audit domain ontology framework

    CN115203429A

  • Auxiliary diagnosis method, device and equipment and storage medium

    CN116705299A

  • News data knowledge graph construction method based on artificial intelligence

    CN118396092A

  • Method and system for multi-level artificial intelligence supercomputer design

    US12001462B1