A method, device and equipment for constructing a domain concept map based on scientific and technological literature

By using scientific and technological literature to construct a field concept map, the problem of difficulty in constructing an accurate field concept map in the existing technology is solved, the depth and accuracy of the field knowledge are achieved, and the professionalism and accuracy of the knowledge are ensured.

CN119807445BActive Publication Date: 2025-05-23LESHAN NORMAL UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510293427.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-13
Publication Date
2025-05-23
Estimated Expiration
2045-03-13

AI Technical Summary

Technical Problem

It is difficult for the prior art to accurately construct concept maps for specific fields, and general concept maps cannot meet the needs of specific fields in terms of the depth and accuracy of domain knowledge.

Method used

By obtaining scientific and technological literature in different fields, we construct a collection of field seed concepts, a collection of field literature vocabulary collections and a collection of field concepts, and use methods such as cosine similarity and confidence clustering to determine the synonyms, similarities and superior-level relationships to construct a field concept map.

Benefits of technology

The concept map construction for specific fields is realized, the accuracy and depth of field concept knowledge is improved, and the professionalism and accuracy of knowledge is ensured.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119807445B_ABST
    Figure CN119807445B_ABST
Patent Text Reader

Abstract

The present invention discloses a method, device and equipment for constructing a domain concept map based on scientific and technological literature, and relates to the technical field of natural language processing. The present invention firstly takes scientific and technological literature as the basis, and constructs a domain seed concept set and a domain concept set with keywords and vocabulary in the scientific and technological literature. This process enables the domain concept knowledge to all come from the scientific and technological literature, and does not require the introduction of expert knowledge, thereby ensuring the accuracy of domain concept extraction; using existing keywords in the domain literature as seed concepts, using similarity calculation methods of different granularities to respectively measure the synonymy, similarity and hyponymy relationships between concepts, so as to construct a domain concept synonymy relationship set, a domain concept similarity relationship set and a domain concept hyponymy relationship set. This process can clearly determine the knowledge boundaries between entities and concepts in the domain concept set, and construct synonymy relationships, similarity relationships and hyponymy relationships, thereby constructing a domain concept map, and greatly improving the accuracy of the constructed domain concept map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of natural language processing technology, and in particular to a method, device, equipment and medium for constructing a domain concept map based on scientific and technological literature. Background Art

[0002] Concept graph is a special form of knowledge graph, which mainly includes entities, concepts and "isA" relationships. Different from traditional knowledge graphs, each entity in the concept graph has a corresponding concept, and concepts are generally connected through "isA" relationships to form a complete concept network. Therefore, the core technologies for constructing concept graphs mainly include concept extraction and synonymy, similarity and hierarchical relationships of concepts. For example, if four concepts such as "Persian cat", "Garfield cat", "cat" and "animal" are extracted, it is necessary to identify a similarity relationship ("Persian cat", Similar, "Garfield cat") and three hierarchical relationships ("Garfield cat", isA, "cat"), ("Persian cat", isA, "cat") and ("cat", isA, "animal").

[0003] Concept graphs have a wide range of application scenarios. The concepts, entities and concept relationships in the graphs can provide references for the construction of large-scale knowledge graphs, and can also provide accurate and reliable knowledge guarantees for knowledge reasoning and induction. In particular, for the "illusion" problem exposed by the current large model technology in specific fields, an accurate domain knowledge system is particularly important for solving this problem. Currently, most common general concept graphs are based on open corpora and extract a large number of concepts and entities, such as Big Word Forest, WordNet and OpenConcepts. However, for specific fields, these concept graphs pay more attention to the breadth of knowledge, and cannot take into account the depth and accuracy of domain knowledge, and cannot accurately and comprehensively cover the knowledge system of specific fields. In addition, the knowledge in the open corpus comes from an open network platform, and the accuracy of its knowledge is limited by the reliability of the data source, such as user comments. Therefore, it is very necessary to build a comprehensive and accurate concept graph for special fields.

[0004] From the perspective of the core technologies of concept graphs, the main methods for concept extraction and relationship identification include rule-based, machine learning-based, or deep learning-based methods. Rule-based methods rely on experts to manually define rule templates, while machine learning-based or deep learning-based methods mostly use supervised methods to learn concept features from a large amount of annotated corpus. However, such methods often rely on expert knowledge, and expert knowledge is usually vague in determining the knowledge boundaries between entities and concepts in a domain. Therefore, current methods find it difficult to accurately construct domain concept graphs. Summary of the invention

[0005] The embodiments of the present invention provide a method, device and equipment for constructing a domain concept map based on scientific and technological literature, which can solve the problem in the prior art that the current methods are difficult to accurately construct a domain concept map.

[0006] The embodiment of the present invention provides a method for constructing a domain concept map based on scientific and technological literature, comprising the following steps:

[0007] Acquire scientific and technological literature in different fields and form a scientific and technological literature collection; extract keywords from scientific and technological literature and construct a domain seed concept collection based on the keywords; extract words contained in sentences in scientific and technological literature, construct a domain literature vocabulary collection, and construct a domain concept collection based on the domain literature vocabulary collection;

[0008] Add all concepts in the domain seed concept set to the domain concept set, use cosine similarity to determine the concept synonymy relationship in the domain concept set, and construct a domain concept synonymy relationship set;

[0009] The concepts in the domain seed concept set are set as initial seeds, and the concept similarity relationships in the domain concept set are clustered using confidence; based on the concept similarity relationships in the domain concept set, a domain concept similarity relationship set is constructed, and the clustered similar concept clusters are formed into a concept cluster set;

[0010] The semantic generalization evaluation and the evaluation method of the hierarchical relationship value of the concept cluster are used to obtain the hierarchical relationship between different clusters in the concept cluster set; according to the hierarchical relationship between different clusters, the domain concept hierarchical relationship set is constructed;

[0011] The concepts in the domain concept set are taken as nodes, and the relationships in the domain concept synonymy set, domain concept similarity set, and domain concept hyponymy set are taken as edges to construct a domain concept graph.

[0012] Preferably, the construction of a domain seed concept set includes:

[0013] Download scientific and technological literature data from scientific and technological literature websites, and divide the scientific and technological literature data into a scientific and technological literature collection SP in a specific field and a scientific and technological literature collection OP in other fields;

[0014] For each document in the scientific literature collection SP of a specific field, all the keywords and text marked in the document are extracted respectively, and the extracted keywords are used to construct the domain seed concept set IC, and the extracted document text is used to construct the domain document set SD.

[0015] Preferably, the constructing of a domain concept set includes:

[0016] Extract the text of each document in the scientific literature collection OP of other fields to construct the non-domain document collection OD;

[0017] For each document in the domain document set SD and the non-domain document set OD, perform sentence splitting, word segmentation and part-of-speech tagging on the sentences, retain the words with noun or verb part-of-speech, and construct the domain document vocabulary set SW and other domain document vocabulary sets OW respectively;

[0018] The word distribution statistical analysis method is used to filter the common words in scientific and technological literature in the domain literature word set SW, and the remaining words in the domain literature word set are phrase expanded to form a domain concept set SC.

[0019] Preferably, constructing a domain concept synonym relationship set includes:

[0020] Add the domain document set SD to the background corpus, and use the neural network model Word2Vec to train and learn the distributed representation of vocabulary;

[0021] Scan the domain concept set SC, given any pair of concepts and , construct candidate relations candiER, and find concepts in distributed representation and The distributed representation vectors are and , calculate the cosine similarity of two vectors ;

[0022] If the cosine similarity , is a pre-designed threshold, then the confidence of the relationship is calculated ,in Indicates the number of co-occurrences of two concepts in sentences. It indicates the number of times two concepts appear at the same time and conform to the concept synonymy relationship template ERPattern;

[0023] If the confidence , is a pre-designed threshold, candiER is added to the concept synonymy relationship set ER to complete the construction of the domain concept synonymy relationship set ER.

[0024] Preferably, the constructing of a set of domain concept homogeneous relationships includes:

[0025] Construct a concept graph G=[Vertex, Edge], where Vertex is the set of vertices in the graph, each concept corresponds to a node, and Edge represents the set of edges in the graph. Initially, there are no edges.

[0026] Take the concepts in the domain seed concept set as the initial seeds, scan each seed concept node in turn, and calculate the similarity between the current seed concept node and the remaining seed concept nodes. The similarity calculation adopts the vector cosine similarity calculation method, and the similarity threshold is ;

[0027] Sort the similarities from high to low and select the unconnected concept nodes with the highest similarity. and ,like and If there is no similar relationship, then construct a candidate relationship candiSR and calculate its relationship confidence ,in Indicates the number of co-occurrences of sentences of two concepts, It indicates the number of times two concepts appear at the same time and conform to the concept similarity relation template SRPattern;

[0028] If the confidence , is a pre-designed threshold, then in the concept and Add an edge between them, that is, the two concepts are in the same cluster; otherwise, no operation is performed; and the confidence is calculated continuously until no similarity with the seed concept is found that is greater than the threshold and separate the isolated nodes in the concept graph into a cluster;

[0029] The concept similarity relationships within each similar concept cluster after clustering are constructed as a domain concept similarity relationship set SR.

[0030] Preferably, the step of constructing a set of domain concept hyponyms and hyponyms includes:

[0031] After clustering, the similar concept clusters together form the concept cluster set CC;

[0032] Scan any pair of concept clusters in the concept cluster set CC and ,The semantic generalization-based method is used to determine the semantic generalization relationship between clusters;

[0033] like and If there is a semantic generalization relationship, scan and Each concept pair and , and judge the concept pair and Is there a hierarchical relationship between them?

[0034] After the semantic generalization relationship between all concept clusters is determined, the hierarchical relationship between different clusters in the concept cluster set CC is obtained; according to the hierarchical relationship between different clusters, the domain concept hierarchical relationship set CR is constructed.

[0035] Preferably, any pair of concept clusters in the scanning concept cluster set CC and , a semantic generalization-based method is used to determine the semantic generalization relationship between clusters, including:

[0036] Scan the concept cluster set CC and set any pair of concept clusters and , respectively calculate the semantic generalization index of the concept cluster and ;

[0037] Concept Cluster Each concept ,statistics The context words and the frequency of each word that appear in the specified window in the original text, the vocabulary set is , computing concept clusters Semantic generalization index ;

[0038] Concept Cluster Each concept ,statistics The context words and the frequency of each word that appear in the specified window in the original text, the vocabulary set is , computing concept clusters Semantic generalization index ;

[0039] Computational Concept Clusters and Jaccard distance ;like , θ 3 is the pre-designed Jaccard distance threshold, then the concept clusters are calculated and The hierarchical relationship value of and ,judge or Is it established? is a pre-set threshold value, whose value is close to 0; if If established, it means and There is a hierarchical relationship, in which For the lower position, For the upper position, if If established, it means and There is a hierarchical relationship, in which For the upper position, is the lower position, otherwise, and There is no superior-subordinate relationship.

[0040] Preferably, the judgment concept is and Whether there is a hierarchical relationship between them, including:

[0041] Scanning concept clusters and ,exist and Take one concept from each cluster and , construct the candidate relation candiCR;

[0042] Calculate the relationship confidence of the candidate relationship candiCR ,in Indicates the number of co-occurrences of sentences of two concepts, Indicates the number of times two concepts appear at the same time and conform to the concept-hypernym relationship template CRPattern;

[0043] If the confidence , is a pre-designed threshold, then the candidate relation candiCR is a concept and The superior-subordinate relationship.

[0044] The embodiment of the present invention also provides a device for constructing a domain concept map based on scientific and technological literature, including:

[0045] The concept matching module is used to obtain scientific and technological documents in different fields and form a collection of scientific and technological documents; extract keywords from scientific and technological documents and construct a collection of domain seed concepts based on the keywords; extract the words contained in the sentences in the scientific and technological documents, construct a collection of domain document words, and construct a collection of domain concepts based on the collection of domain document words;

[0046] The concept relationship module is used to add all concepts in the domain seed concept set to the domain concept set, determine the concept synonymy relationship in the domain concept set using cosine similarity, and construct a domain concept synonymy relationship set;

[0047] The concepts in the domain seed concept set are set as initial seeds, and the concept similarity relationships in the domain concept set are clustered using confidence; based on the concept similarity relationships in the domain concept set, a domain concept similarity relationship set is constructed, and the clustered similar concept clusters are formed into a concept cluster set;

[0048] The semantic generalization evaluation and the evaluation method of the hierarchical relationship value of the concept cluster are used to obtain the hierarchical relationship between different clusters in the concept cluster set; according to the hierarchical relationship between different clusters, the domain concept hierarchical relationship set is constructed;

[0049] The graph construction module is used to construct a domain concept graph by taking each concept in the domain concept set as a node, and each relationship in the domain concept synonym relationship set, the domain concept similar relationship set, and the domain concept hyponym relationship set as an edge.

[0050] An embodiment of the present invention further provides an electronic device, including a memory and a processor;

[0051] The memory is used to store computer programs;

[0052] The processor is used to implement the steps of the method for constructing a domain concept map based on scientific and technological literature when executing the computer program stored in the memory.

[0053] An embodiment of the present invention also provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the steps of a method for constructing a domain concept map based on scientific and technological literature as described above.

[0054] The embodiment of the present invention provides a method, device and apparatus for constructing a domain concept map based on scientific and technological literature. Compared with the prior art, the beneficial effects thereof are as follows:

[0055] The present invention firstly builds a domain seed concept set based on scientific and technological literature and keywords in the scientific and technological literature, and at the same time extracts the words contained in the sentences in the scientific and technological literature to build a domain document vocabulary set, and builds a domain concept set according to the domain document vocabulary set, sets the concepts in the domain seed concept set as the initial concept seeds, and respectively builds a domain concept synonymy relationship set, a domain concept similar relationship set and a domain concept hyponymy relationship set, and builds a domain concept map; this process makes the domain concept knowledge all come from scientific and technological literature, fully utilizes the advantages of strong professional domain and high knowledge accuracy of scientific and technological literature, and does not require the introduction of additional expert knowledge, thereby ensuring the accuracy of domain concept extraction; at the same time, using the keywords unique to scientific and technological literature as concept seeds, using similarity calculation methods of different granularities to respectively measure the synonymy, similarity and hyponymy relationships between concepts, so as to build a domain concept synonymy relationship set, a domain concept similar relationship set and a domain concept hyponymy relationship set, which can clearly determine the knowledge boundaries between entities and concepts in the domain concept set, and greatly improve the accuracy of the constructed domain concept map. BRIEF DESCRIPTION OF THE DRAWINGS

[0056] Figure 1A schematic diagram of the overall process of a method for constructing a domain concept map based on scientific and technological literature provided in an embodiment of the present invention. DETAILED DESCRIPTION

[0057] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the specific embodiments of the present invention are described in detail below in conjunction with the accompanying drawings. In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without violating the connotation of the present invention, so the present invention is not limited by the specific embodiments disclosed below.

[0058] See also Figure 1 , the embodiment of the present invention provides a method for constructing a domain concept map based on scientific and technological literature. The present invention downloaded 3,600 scientific and technological documents from websites such as HowNet and Wanfang (including 1,200 in the fields of autism, rice pests and diseases, and cultural tourism), and used them as training and test corpora to verify the effectiveness of the method provided by the present invention; the embodiment on this data set shows that the domain concepts extracted by this method have a high recall rate, the domain concept relationship is accurately portrayed, and the constructed concept map has a wide knowledge coverage and a deep knowledge hierarchy relationship. The process specifically includes the following steps:

[0059] Step S1: Download scientific and technological literature data from scientific and technological literature websites such as HowNet and Wanfang, and divide them into a scientific and technological literature collection SP (Special Papers) in a specific field and a scientific and technological literature collection OP (Other Papers) in several other fields. Perform preprocessing operations on the literature collection SP and OP respectively to construct a domain seed concept collection, a domain literature vocabulary collection, and a collection of other domain literature vocabulary. Specifically, it includes:

[0060] Step S11: For each document in the scientific and technological literature collection SP (Special Papers) in a specific field, extract all the keywords and texts given by the author in the document, use the extracted keywords to construct the specific field seed concept set IC (Initial Concepts), and use the extracted document text to construct the field document set SD (Special Docs).

[0061] Step S12: For each document in the scientific and technological document set OP (Other Papers) in other fields, extract the main text in the document and construct a non-domain document set OD (Other Docs).

[0062] Step S13: For each document in the SD and OD sets, firstly split the sentences, then perform word segmentation and part-of-speech tagging on the sentences, retain the words with noun or verb as the part of speech, and construct the domain literature vocabulary set SW (Special Words) and other domain literature vocabulary set OW (Other Words).

[0063] Step S2: Use the vocabulary distribution learning method to filter the common scientific and technological literature vocabulary in the field literature vocabulary set SW (vocabulary with a high probability of appearing in SW and OW, such as "paper", "research", "propose", "clarify", etc.). Specifically include:

[0064] Step S21: Scan each word in the domain literature vocabulary set SW , calculate the inverse document frequency of the word in the domain document set SD ,in, Include vocabulary in SD The number of documents, is the number of documents in the field.

[0065] Step S22: Scan each word in the domain document vocabulary set SW , calculate the document frequency of the word in the non-domain document collection ,in, Include vocabulary for non-domain document collections The number of documents, The number of all documents.

[0066] Step S23: Scan each word in the domain document vocabulary set SW , calculate the category distinction of the vocabulary ,like (α is a pre-designed threshold), then the retained vocabulary , otherwise delete the word from the set SW.

[0067] At the same time, it should be noted that for domain vocabulary , its inverse document frequency in the domain document collection The larger the value is, the greater the domain category distinction of the word is; its document frequency in non-domain document collections is The larger the value is, the smaller the domain category distinction of the word is; therefore, The smaller it is, the greater the distinction between vocabulary categories.

[0068] In particular, the method of vocabulary distribution learning is not limited to the above method, and the topic model LDA can also be used to learn the topic (domain) distribution of vocabulary using LDA, and finally retain the vocabulary with higher topic probability in the domain literature vocabulary set SW.

[0069] Step S3: Perform phrase expansion operations on the words in the domain literature vocabulary set SW to construct a domain concept set SC (Special Concepts). Specifically including:

[0070] Step S31: Scan each remaining word in the domain document vocabulary set SW , locate each sentence in the domain document SD , using a syntactic analyzer to analyze the sentence After syntactic analysis, each sentence is obtained The syntax tree of .

[0071] Step S32: Find the word , starting from this vocabulary node in the syntax tree Backtrack to find the nearest ancestor NP or VP node to the word .

[0072] Step S33: If found, extract the node All leaf nodes under the Combined into phrases corresponding to the vocabulary , and then the phrase Add domain concept set SC.

[0073] Step S34: If not found, Directly join the domain concept set SC.

[0074] For example, the two adjacent words “narrative” and “ability” in a sentence fragment were incorrectly split during preprocessing and expanded into a more accurate domain concept “narrative ability” through phrase expansion.

[0075] Step S4: Add the domain document set SD to the background corpus, and learn the distributed representation of vocabulary after training using a neural network model (such as Word2Vec).

[0076] Step S5: Add all concepts in the domain seed concept set IC to the domain concept set SC, and use a method combining word similarity calculation and rules to determine the concept synonym relationship set ER (EqualRelations) in the domain concept set SC. Specifically, it includes:

[0077] Step S51: Design a concept synonym relationship template ERPattern, for example:

[0078] sepcialConcept1[,]{that is|also called|also known as|also known as}sepcialConcept2.

[0079] Step S52: Scan the domain concept set SC, given any pair of concepts and , construct candidate relations candiER, and search for concepts in the learned distributed representation and The distributed representation vectors are and , calculate the cosine similarity of two vectors .

[0080] Step S53: If the cosine similarity ( is a pre-designed threshold), then the confidence of the relationship is calculated ,in Indicates the number of co-occurrences of two concepts in sentences. Indicates the number of times two concepts appear at the same time and conform to the concept synonym relationship template ERPattern.

[0081] Step S54: If the confidence ( is a pre-designed threshold), candiER is added to the concept synonym relationship set ER.

[0082] Step S6: Using the concepts in the domain seed concept set IC (InitialConcepts) as the initial seeds, a method combining parameterless clustering and rules is used to determine the concept similarity relationship set SR (SimilarRelations) in the domain concept set SC. The clustered similar concept clusters together form the concept cluster set CC (ConceptClusters). Specifically include:

[0083] Step S61: Design a Chinese semantic homogeneous relationship template SRPattern, for example:

[0084] {Includes|contains} sepcialConcept1{,|and|and} sepcialConcept2.

[0085] Step S62: Using the concepts in the domain seed concept set IC as initial seeds, cluster the domain concept set SC using a method combining parameterless clustering (such as the Whisper algorithm) and rules to obtain a domain concept cluster set CC. Each concept cluster Contains concepts with high semantic similarity. Specifically including:

[0086] Step S621: construct a concept graph G=(Vertex, Edge), where Vertex is a set of vertices in the graph, each concept corresponds to a node, and Edge represents an edge set of the graph, and initially no edge exists.

[0087] Step S622: Take the concepts in the domain seed concept set IC as the initial seeds, scan each seed concept node in turn, and calculate the similarity between the seed node and the remaining nodes (the node vector uses the learned distributed representation, the similarity calculation uses the vector cosine similarity calculation method, and the similarity threshold is ).

[0088] Step S623: After sorting by similarity from high to low, select the unconnected concept node with the highest similarity and ,like and If there is no similar relationship, then construct a candidate relationship candiSR and calculate its relationship confidence ,in Indicates the number of co-occurrences of sentences of two concepts, Indicates the number of times two concepts appear at the same time and conform to the concept similarity relation template SRPattern.

[0089] Step S624: If the confidence ( is the pre-designed threshold), then in the concept and An edge is added between them, that is, the two concepts are in the same cluster; otherwise, no operation is performed.

[0090] Step S625: Repeat the above 3 steps until no similarity with the seed concept is found that is greater than the threshold. to the conceptual node.

[0091] Step S626: At this point, if there are still isolated nodes in the graph (i.e., the concept node is a cluster of its own), repeat the above process for the remaining isolated nodes (each node can be used as a seed node) until no nodes with a similarity greater than the threshold are found. The concept is correct so far.

[0092] Step S627: All concept nodes on each connected component in the graph are regarded as the same concept cluster, and all connected components constitute a concept cluster set CC.

[0093] Step S63: Scan each concept cluster , add similar relationships to each pair of concepts in the cluster and add them to the concept similar relationship set SR.

[0094] Step S7: Use a method based on semantic generalization and rules to measure the hierarchical relationship between different clusters in the concept cluster set CC (ConceptClusters), and thereby determine the hierarchical relationship between the superordinate cluster concept and the subordinate cluster concept, thereby obtaining the concept hierarchical relationship set CR (ConceptRelations). Specifically including:

[0095] Step S71: Design a Chinese semantic hyponym template CRPattern, for example:

[0096] sepcialConcept1{includes|contains}{ sepcialConcept2,}*[etc.|etc.].

[0097] sepcialConcept1{such as|for example|like|like}{ sepcialConcept2,}*.

[0098] sepcialConcept2 is {a|a kind of|a type of} sepcialConcept1.

[0099] Step S72: Scan any pair of concept clusters in the concept cluster set CC and use a semantic generalization-based method to determine the semantic generalization relationship between clusters. Specifically, it includes:

[0100] Step S721: Scan the concept cluster set CC, given any pair of concept clusters and , respectively calculate the semantic generalization index of the concept cluster and .

[0101] Step S722: Concept clusters Each concept ,statistics The context words and the frequency of each word that appear in the specified window in the original text (the vocabulary set is ), computing concept clusters Semantic generalization index .

[0102] Step S723: Concept clusters Each concept ,statistics The context words and the frequency of each word that appear in the specified window in the original text (the vocabulary set is ), computing concept clusters Semantic generalization index .

[0103] Step S724: Calculate the Jaccard distance of the concept clusters. .

[0104] Step S725: If , θ 3 is the pre-designed Jaccard distance threshold, then the concept clusters are calculated and The hierarchical relationship value of and ,judge or ( is a pre-set threshold value, whose value is close to 0) is established, if If established, it means and There is a hierarchical relationship ( For the lower position, is the upper position), if If established, it means and There is a hierarchical relationship ( For the upper position, is the lower position), otherwise, and There is no superior-subordinate relationship.

[0105] Step S73: If and If there is a semantic generalization relationship, scan and Each concept pair in the , and determine whether there is a hierarchical relationship between the concept pairs. Specifically include:

[0106] Step S731: Scan concept clusters and ,exist and Take one concept from each cluster and , construct the candidate relation candiCR.

[0107] Step S732: Calculate the relationship confidence of candiCR ,in Indicates the number of co-occurrences of sentences of two concepts, Indicates the number of times two concepts appear at the same time and conform to the concept-hypernym relationship template CRPattern.

[0108] Step S733: If the confidence ( is a pre-designed threshold), then candiCR is added to the concept hyponym set CR.

[0109] Step S734: Repeat the above process until the concept cluster and Until all concept pairs are scanned.

[0110] Step S74: Repeat the scanning of inter-cluster relations until all inter-cluster semantic generalization relations are judged.

[0111] Step S8: Using each concept in the domain concept set SC as a node, and each relationship in the concept synonymy relationship set ER, the concept homonymy relationship set SR, and the concept hyponymy relationship set CR as an edge, a domain concept graph is constructed.

[0112] The domain concept knowledge in the present invention comes from scientific and technological literature, and the accuracy and authority of the data are more guaranteed than that of open network platforms; at the same time, the semi-structured information in scientific and technological literature (such as titles and keywords, etc.) can be regarded as natural crowdsourcing data, providing an effective data basis for the extraction of domain concepts and the identification of conceptual relationships; on the one hand, the present invention aims to build a comprehensive concept knowledge network for a specific field, effectively improving the accuracy and depth of domain knowledge; on the other hand, it effectively utilizes the data characteristics of scientific and technological literature, pays attention to the automation and unsupervised nature of the method, and tries to reduce the introduction of expert knowledge or manual annotation work.

[0113] The present invention constructs a domain concept map based on scientific and technological literature. The authority of scientific and technological literature provides strong data guarantee for the accuracy of the concept map knowledge system. The present invention introduces scientific and technological literature data from other fields as counterexamples to effectively filter common terms in scientific and technological literature and ensure the professionalism of domain concepts. The present invention uses keywords unique to scientific and technological literature as concept seeds, without the need to introduce additional expert knowledge, which can not only effectively guarantee the accuracy of domain concept extraction, but also help to define the knowledge boundaries of domain concepts. The present invention uses different semantic calculation methods such as distributed semantics, Jaccard distance and semantic generalization to measure synonymous, similar and hyponymous relationships between concepts, and combines template matching to evaluate the confidence of concept semantic relationships, so as to accurately characterize and measure different concept semantic relationships, thereby ensuring the accuracy of the concept map structure.

[0114] The method of the present invention, firstly, makes full use of the advantages of strong professional domain and high knowledge accuracy of scientific and technological literature, and provides strong data guarantee for the accuracy of the concept map knowledge system; secondly, by introducing scientific and technological literature data in other fields as counterexamples and utilizing this comparative learning mechanism, it effectively overcomes the interference of general terms and concepts, and has a clearer definition of the knowledge boundaries of concepts in specific fields; thirdly, by utilizing the keywords unique to scientific and technological literature as concept seeds, it effectively improves the accuracy and recall rate of domain concept extraction, and ensures the accuracy and comprehensiveness of the domain concept map; finally, semantic computing methods of different granularities are used to measure the synonymous, similar and hyponymous relationships in the concept map, and template matching is used to evaluate the confidence of different semantic relationships, effectively ensuring the structural accuracy of the concept map.

[0115] The embodiment of the present invention also provides a device for constructing a domain concept map based on scientific and technological literature, including:

[0116] The concept matching module is used to obtain scientific and technological literature in different fields and form a scientific and technological literature collection; extract keywords from scientific and technological literature and build a domain seed concept collection based on the keywords; extract the words contained in the sentences in the scientific and technological literature, build a domain literature vocabulary collection, and build a domain concept collection based on the domain literature vocabulary collection.

[0117] The concept relationship module is used to add all concepts in the domain seed concept set to the domain concept set, and use cosine similarity to determine the concept synonymy relationship in the domain concept set, and to construct a domain concept synonymy relationship set.

[0118] The concepts in the domain seed concept set are set as initial seeds, and the concept similarity relationships in the domain concept set are clustered using confidence; based on the concept similarity relationships in the domain concept set, a domain concept similarity relationship set is constructed, and the clustered similar concept clusters are combined into a concept cluster set.

[0119] The semantic generalization evaluation and the evaluation method of the hierarchical relationship value of concept clusters are used to obtain the hierarchical relationship between different clusters in the concept cluster set; according to the hierarchical relationship between different clusters, a domain concept hierarchical relationship set is constructed.

[0120] The graph construction module is used to construct a domain concept graph by taking each concept in the domain concept set as a node, and each relationship in the domain concept synonym relationship set, the domain concept similar relationship set, and the domain concept hyponym relationship set as an edge.

[0121] An embodiment of the present invention further provides an electronic device, including a memory and a processor.

[0122] The memory is used to store computer programs.

[0123] When the processor is used to execute the computer program stored in the memory, it implements the steps of the above method for constructing a domain concept map based on scientific and technological literature.

[0124] An embodiment of the present invention also provides a computer-readable storage medium for storing a computer program, which, when executed by a processor, implements the steps of the above method for constructing a domain concept map based on scientific and technological literature.

[0125] The above-mentioned embodiments only express several implementation methods of the present invention, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the invention patent. It should be pointed out that, for ordinary technicians in this field, several variations and improvements can be made without departing from the concept of the present invention, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the attached claims.

Claims

1. A method for constructing a domain concept map based on scientific and technological literature, characterized in that: The following steps are involved: Acquire scientific and technological literature in different fields and form a scientific and technological literature collection; extract keywords from scientific and technological literature and construct a domain seed concept collection based on the keywords; extract words contained in sentences in scientific and technological literature, construct a domain literature vocabulary collection, and construct a domain concept collection based on the domain literature vocabulary collection; Add all concepts in the domain seed concept set to the domain concept set, use cosine similarity to determine the concept synonymy relationship in the domain concept set, and construct a domain concept synonymy relationship set; The concepts in the domain seed concept set are set as initial seeds, and the concept similarity relationships in the domain concept set are clustered using confidence; based on the concept similarity relationships in the domain concept set, a domain concept similarity relationship set is constructed, and the clustered similar concept clusters are formed into a concept cluster set; The semantic generalization evaluation and the evaluation method of the hierarchical relationship value of the concept cluster are used to obtain the hierarchical relationship between different clusters in the concept cluster set; according to the hierarchical relationship between different clusters, the domain concept hierarchical relationship set is constructed; Take each concept in the domain concept set as a node, and each relationship in the domain concept synonymy relationship set, domain concept homogeneity relationship set, and domain concept hyponymy relationship set as an edge to construct a domain concept graph; The constructing of a domain concept hyponym set includes: After clustering, the similar concept clusters together form the concept cluster set CC; Scan any pair of concept clusters in the concept cluster set CC and , a semantic generalization-based method is used to determine the semantic generalization relationship between clusters; specifically, it includes: Scan the concept cluster set CC and set any pair of concept clusters and , respectively calculate the semantic generalization index of the concept cluster and ; Concept Cluster Each concept ,statistics The context words and the frequency of each word that appear in the specified window in the original text, the vocabulary set is , computing concept clusters Semantic generalization index ; Concept Cluster Each concept ,statistics The context words and the frequency of each word that appear in the specified window in the original text, the vocabulary set is , computing concept clusters Semantic generalization index ; Computational Concept Clusters and Jaccard distance ;like , θ 3 is the pre-designed Jaccard distance threshold, and the concept clusters are calculated and The hierarchical relationship value of and ,judge or Is it established? is a pre-set threshold value, whose value is close to 0; if If established, it means and There is a hierarchical relationship, in which For the lower position, For the upper position, if If established, it means and There is a hierarchical relationship, in which For the upper position, is the lower position, otherwise, and There is no superior-subordinate relationship; like and If there is a semantic generalization relationship, scan and Each concept pair and , and judge the concept pair and Whether there is a hierarchical relationship between them; specifically including: Scanning concept clusters and ,exist and Take one concept from each cluster and , construct the candidate relation candiCR; Calculate the relationship confidence of the candidate relationship candiCR ,in Indicates the number of co-occurrences of two concepts in sentences. Indicates the number of times two concepts appear at the same time and conform to the concept-hypernym relationship template CRPattern; If the confidence , is a pre-designed threshold, then the candidate relation candiCR is a concept and The relationship between superior and subordinate; After the semantic generalization relationship between all concept clusters is determined, the hierarchical relationship between different clusters in the concept cluster set CC is obtained; according to the hierarchical relationship between different clusters, the domain concept hierarchical relationship set CR is constructed.

2. According to claim 1, a method for constructing a domain concept map based on scientific and technological literature is characterized in that: The construction domain seed concept set includes: Download scientific and technological literature data from scientific and technological literature websites, and divide the scientific and technological literature data into a scientific and technological literature collection SP in a specific field and a scientific and technological literature collection OP in other fields; For each document in the scientific literature collection SP of a specific field, all the keywords and text marked in the document are extracted respectively, and the extracted keywords are used to construct the domain seed concept set IC, and the extracted document text is used to construct the domain document set SD.

3. According to claim 2, a method for constructing a domain concept map based on scientific and technological literature is characterized in that: The construction domain concept set includes: Extract the text of each document in the scientific literature collection OP of other fields to construct the non-domain document collection OD; For each document in the domain document set SD and the non-domain document set OD, perform sentence splitting, word segmentation and part-of-speech tagging on the sentences, retain the words with noun or verb part-of-speech, and construct the domain document vocabulary set SW and other domain document vocabulary sets OW respectively; The word distribution statistical analysis method is used to filter the common words in scientific and technological literature in the domain literature word set SW, and the remaining words in the domain literature word set are phrase expanded to form a domain concept set SC.

4. According to claim 3, a method for constructing a domain concept map based on scientific and technological literature is characterized in that: The constructing of a domain concept synonym relationship set includes: Add the domain document set SD to the background corpus, and use the neural network model Word2Vec to train and learn the distributed representation of vocabulary; Scan the domain concept set SC, given any pair of concepts and , construct candidate relations candiER, and find concepts in distributed representation and The distributed representation vectors are and , calculate the cosine similarity of two vectors ; If the cosine similarity , is a pre-designed threshold, then the confidence of the relationship is calculated ,in Indicates the number of co-occurrences of two concepts in sentences. It indicates the number of times two concepts appear at the same time and conform to the concept synonymy relationship template ERPattern; If the confidence , is a pre-designed threshold, candiER is added to the concept synonymy relationship set ER to complete the construction of the domain concept synonymy relationship set ER.

5. According to claim 4, a method for constructing a domain concept map based on scientific and technological literature is characterized in that: The construction of the domain concept homogeneous relationship set includes: Construct a concept graph G=[Vertex, Edge], where Vertex is the set of vertices in the graph, each concept corresponds to a node, and Edge represents the set of edges in the graph. Initially, there are no edges. Take the concepts in the domain seed concept set as the initial seeds, scan each seed concept node in turn, and calculate the similarity between the current seed concept node and the remaining seed concept nodes. The similarity calculation adopts the vector cosine similarity calculation method, and the similarity threshold is ; Sort the similarities from high to low and select the unconnected concept nodes with the highest similarity. and ,like and If there is no similar relationship, then construct a candidate relationship candiSR and calculate its relationship confidence ,in Indicates the number of co-occurrences of two concepts in sentences. It indicates the number of times two concepts appear at the same time and conform to the concept similarity relation template SRPattern; If the confidence , is a pre-designed threshold, then in the concept and Add an edge between them, that is, the two concepts are in the same cluster; otherwise, no operation is performed; and the confidence is calculated continuously until no similarity with the seed concept is found that is greater than the threshold and separate the isolated nodes in the concept graph into a cluster; The concept similarity relationships within each similar concept cluster after clustering are constructed as a domain concept similarity relationship set SR.

6. A device for constructing a domain concept map based on scientific and technological literature, characterized in that: include: Concept matching module, used to obtain scientific and technological literature in different fields and form a collection of scientific and technological literature; Extract keywords from scientific and technological documents and construct a domain seed concept set based on the keywords; extract words contained in sentences in scientific and technological documents, construct a domain document vocabulary set, and construct a domain concept set based on the domain document vocabulary set; The concept relationship module is used to add all concepts in the domain seed concept set to the domain concept set, determine the concept synonymy relationship in the domain concept set using cosine similarity, and construct a domain concept synonymy relationship set; The concepts in the domain seed concept set are set as initial seeds, and the concept similarity relationships in the domain concept set are clustered using confidence; based on the concept similarity relationships in the domain concept set, a domain concept similarity relationship set is constructed, and the clustered similar concept clusters are formed into a concept cluster set; The semantic generalization evaluation and the evaluation method of the hierarchical relationship value of the concept cluster are used to obtain the hierarchical relationship between different clusters in the concept cluster set; according to the hierarchical relationship between different clusters, the domain concept hierarchical relationship set is constructed; A graph construction module is used to construct a domain concept graph by taking each concept in the domain concept set as a node, and each relationship in the domain concept synonymy relationship set, the domain concept homogeneity relationship set, and the domain concept hyponymy relationship set as an edge; The constructing of a domain concept hyponym set includes: After clustering, the similar concept clusters together form the concept cluster set CC; Scan any pair of concept clusters in the concept cluster set CC and , a semantic generalization-based method is used to determine the semantic generalization relationship between clusters; specifically, it includes: Scan the concept cluster set CC and set any pair of concept clusters and , respectively calculate the semantic generalization index of the concept cluster and ; Concept Cluster Each concept ,statistics The context words and the frequency of each word that appear in the specified window in the original text, the vocabulary set is , computing concept clusters Semantic generalization index ; Concept Cluster Each concept ,statistics The context words and the frequency of each word that appear in the specified window in the original text, the vocabulary set is , computing concept clusters Semantic generalization index ; Computational Concept Clusters and Jaccard distance ;like , θ 3 is the pre-designed Jaccard distance threshold, and the concept clusters are calculated and The hierarchical relationship value of and ,judge or Is it established? is a pre-set threshold value, whose value is close to 0; if If established, it means and There is a hierarchical relationship, in which For the lower position, For the upper position, if If established, it means and There is a hierarchical relationship, in which For the upper position, is the lower position, otherwise, and There is no superior-subordinate relationship; like and If there is a semantic generalization relationship, scan and Each concept pair and , and judge the concept pair and Whether there is a hierarchical relationship between them; specifically including: Scanning concept clusters and ,exist and Take one concept from each cluster and , construct the candidate relation candiCR; Calculate the relationship confidence of the candidate relationship candiCR ,in Indicates the number of co-occurrences of two concepts in sentences. Indicates the number of times two concepts appear at the same time and conform to the concept-hypernym relationship template CRPattern; If the confidence , is a pre-designed threshold, then the candidate relation candiCR is a concept and The relationship between superior and subordinate; After the semantic generalization relationship between all concept clusters is determined, the hierarchical relationship between different clusters in the concept cluster set CC is obtained; according to the hierarchical relationship between different clusters, the domain concept hierarchical relationship set CR is constructed.

7. An electronic device, characterized in that: include: Memory and processor; The memory is used to store computer programs; The processor, when used to execute the computer program stored in the memory, implements the steps of a method for constructing a domain concept map based on scientific and technological literature as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Ontology model construction method of domain knowledge graph

    CN111177322A

  • Concept knowledge graph construction method and device

    CN112131401A

  • Method and device for carrying out concept expansion on science and technology concept atlas

    CN115062622A