Biological Domain Knowledge Graph Construction Method and Device

Through the biological domain knowledge graph construction method, lexical analysis, deep learning and PCNN relationship extraction technologies are used to solve the problem of difficult integration and inference of biological relationships in the existing technology, and efficient infectious disease tracing and infection source identification are achieved.

CN114897167BActive Publication Date: 2025-06-13INNER MONGOLIA UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210522728.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-13
Publication Date
2025-06-13
Estimated Expiration
2042-05-13

AI Technical Summary

Technical Problem

The existing technology is difficult to effectively integrate and infer the relationship between organisms, and lacks concise and intuitive tools to study the relationship between organisms, affecting the tracing of infectious diseases and the identification of transmission pathways.

Method used

A biological domain knowledge graph construction method is adopted, and the initial entity set is obtained through lexical analysis tools, combined with the predetermined entity library for expansion and naming, deep learning and PCNN relationship extraction and acquisition of triples, and bidirectional supervision and iterative fusion representation learning method is used to obtain biological domain knowledge graph.

Benefits of technology

Improve the efficiency of infectious disease traceability and infectious source identification, providing a concise and intuitive tool to integrate and reason the relationship between organisms, supporting the summary of commonalities of endangered species and rapid response in the outbreak of infectious disease.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114897167B_ABST
    Figure CN114897167B_ABST
Patent Text Reader

Abstract

The present invention discloses a method and apparatus for constructing a knowledge graph in the biological field. The method includes: using a lexical analysis tool to perform word segmentation, part-of-speech tagging, and entity recognition on the text content in the biological field to obtain an initial entity set, obtaining an extended entity library through the entity set, and naming and classifying entities in the extended entity library; classifying entity relationships of the classified entities in the extended entity library according to predetermined rules and dictionaries and based on a deep learning method, and extracting relationships of the entities in the extended entity library through relationship extraction based on PCNN to obtain triples; fusing the entities in the triples through a representation learning method based on bidirectional supervision and iterative fusion and an embedding model based on the knowledge graph to obtain a knowledge graph in the biological field. The present invention obtains useful entities in the biological field text, performs entity recognition, relationship classification, and knowledge fusion, so as to obtain a knowledge graph in the biological field.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of atlas construction, and particularly to a method and device for constructing a knowledge atlas in the biological field. Background Art

[0002] In order to study the relationships between organisms, artificial intelligence has gradually entered the laboratory. With the progress of science and technology, reasoning ability has become an important feature of artificial intelligence, and knowledge graphs have also become an effective method for computers to understand semantic information. It is an important production material in the intelligent society. It can not only display complex knowledge fields and knowledge systems through data mining, information processing, knowledge measurement, and graphic drawing, but also map the human cognitive way of the world and help institutions achieve business intelligence. At present, the research in the biological field urgently needs a concise and intuitive tool to integrate the relationships between organisms, providing convenience for the reasoning and retrieval of relationships. Therefore, the knowledge graph based on the biological field has emerged as the times require. Since the organization form of the knowledge graph is extremely similar to the biological chain in nature, such a knowledge graph based on the biological field can reflect the position of any organism in the ecological chain, and can summarize the commonalities of endangered species by horizontally comparing endangered species on the earth. Moreover, when infectious diseases such as the COVID-19 pandemic and plague break out, if we use the biological knowledge graph to master the relationships between organisms in advance, then tracing the source of infection and mastering the transmission route will become simple and fast. Summary of the Invention

[0003] The purpose of the present invention is to provide a method and device for constructing a knowledge atlas in the biological field, aiming to solve the above problems in the prior art.

[0004] The present invention provides a method for constructing a knowledge atlas in the biological field, including:

[0005] Using a lexical analysis tool to perform word segmentation, part-of-speech tagging, and entity recognition on the text content in the biological field to obtain an initial entity set, comparing the initial entity set with a predetermined entity library to obtain an extended entity library, naming the extended entity library, and classifying the entities in the named extended entity library;

[0006] According to predetermined rules and dictionaries and based on a deep learning method, classifying the entities in the extended entity library, and performing relationship extraction on the entities in the extended entity library through relationship extraction based on PCNN to obtain triples;

[0007] Fusing the entities in the triples through a representation learning method based on bidirectional supervision and iterative fusion and an embedding model based on the knowledge graph to obtain a knowledge atlas in the biological field.

[0008] The present invention provides a device for constructing a knowledge atlas in the biological field, including:

[0009] An entity recognition and classification module uses a lexical analysis tool to perform word segmentation, part-of-speech tagging, and entity recognition on the text content in the biological field to obtain an initial entity set, obtains an extended entity library by comparing the initial entity set with a predetermined entity library, names the extended entity library, and classifies the entities in the named extended entity library;

[0010] A relationship extraction module classifies entity relationships of the classified entities in the extended entity library by methods based on rules and dictionaries and methods based on deep learning, and extracts relationships of the entities in the extended entity library by relationship extraction based on PCNN to obtain triples;

[0011] A knowledge graph acquisition module fuses the entities in the triples by a representation learning method based on bidirectional supervision and iterative fusion and an embedding model based on the knowledge graph to obtain a biological field knowledge graph.

[0012] An embodiment of the present invention further provides a biological field knowledge graph construction device, including: a memory, a processor, and a computer program stored on the memory and executable on the processor, and when the computer program is executed by the processor, the steps of the above biological field knowledge graph construction method are implemented.

[0013] An embodiment of the present invention further provides a computer-readable storage medium, on which an information transmission implementation program is stored, and when the program is executed by a processor, the steps of the above biological field knowledge graph construction method are implemented.

[0014] By adopting the embodiment of the present invention, an extended entity library is obtained by using a lexical analysis tool, the entities in the extended entity library are classified, relationships of the entities in the extended entity library are extracted by relationship extraction based on PCNN to obtain triples, and finally the entities in the triples are fused by a representation learning method based on bidirectional supervision and iterative fusion and an embedding model based on the knowledge graph to obtain a biological field knowledge graph, improving the tracing efficiency of infectious diseases and the efficiency of obtaining the sources of infectious diseases.

[0015] The above description is only an overview of the technical solution of the present invention. In order to be able to understand the technical means of the present invention more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of the present invention more obvious and understandable, the following specifically describes the specific embodiments of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0016] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the following will briefly introduce the drawings required for the description of the specific embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0017] Figure 1 It is a flowchart of the method for constructing a knowledge graph in the biological field according to an embodiment of the present invention;

[0018] Figure 2 It is a technical roadmap of entity recognition according to an embodiment of the present invention;

[0019] Figure 3 It is a schematic diagram of entity category extraction according to an embodiment of the present invention;

[0020] Figure 4 It is a schematic diagram of the multi-page similarity calculation process according to an embodiment of the present invention;

[0021] Figure 5 It is a schematic diagram of logical rule tree learning according to an embodiment of the present invention;

[0022] Figure 6 It is a schematic diagram of biological relationship classification based on the "pre-training + fine-tuning" technology according to an embodiment of the present invention;

[0023] Figure 7 It is a schematic diagram of the improvement of the MASK method according to an embodiment of the present invention;

[0024] Figure 8 It is a schematic diagram of model distillation according to an embodiment of the present invention;

[0025] Figure 9 It is a flowchart of relationship extraction according to an embodiment of the present invention;

[0026] Figure 10 It is a syntactic analysis tree according to an embodiment of the present invention;

[0027] Figure 11 It is a schematic diagram of the automatic learning text feature framework according to an embodiment of the present invention;

[0028] Figure 12 It is a schematic diagram of the PCNN model with segment pooling according to an embodiment of the present invention;

[0029] Figure 13 It is a schematic diagram of knowledge representation learning according to an embodiment of the present invention;

[0030] Figure 14 It is a schematic diagram of the technical roadmap of two-way supervised training and iterative fusion according to an embodiment of the present invention;

[0031] Figure 15 It is a schematic diagram of the device for constructing a knowledge graph in the biological field according to an embodiment of the present invention. Detailed implementation manners

[0032] Next, the technical solutions of the present invention will be described clearly and completely in conjunction with the embodiments. Obviously, the described embodiments are some but not all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0033] Method embodiments

[0034] According to an embodiment of the present invention, a method for constructing a knowledge graph in the biological field is provided. Figure 1 It is a flowchart of the method for constructing a knowledge graph in the biological field according to an embodiment of the present invention. As Figure 1 shown, the method for constructing a knowledge graph according to an embodiment of the present invention specifically includes:

[0035] Step S101: Use a lexical analysis tool to perform word segmentation, part-of-speech tagging, and entity recognition on the text content in the biological field to obtain an initial entity set. Compare the initial entity set with a predetermined entity library to obtain an extended entity library, name the extended entity library, and classify the entities in the named extended entity library.

[0036] The entity types involved in the biological field are diverse, numerous, new entities are constantly emerging, and the entity composition structure is relatively complex. The length of some types of entity words has no certain limit. In addition, in different fields and scenarios, the extension of entities is different, which may cause ambiguity. Therefore, the present invention adopts a hybrid method of a rule- and knowledge-based method and a statistic-based method for entity recognition. Figure 2 It is a technical roadmap for entity recognition according to an embodiment of the present invention. The specific steps of entity recognition are as follows:

[0037] Entity category extraction: Since the general entities in biology have exceeded the five major categories and seven sub-categories of named entities (this is a relatively mature category system in public research, but we need to determine entity categories according to actual needs), and the amount of data contained is huge, the existing entity categories cannot fully cover the entire biological field (relationships are established on the basis of entities, so if entity recognition is inaccurate, it will affect the extraction of relationships). Therefore, we need to obtain more entity categories to improve the coverage of data. As Figure 3 shown, it is a schematic diagram of entity category extraction according to an embodiment of the present invention.

[0038] Word Segmentation and Part-of-Speech Tagging: As previously mentioned, the thulac tool, which is based on public data and can only identify three categories (person names, place names, and organization names), is used for word segmentation, part-of-speech tagging, and entity recognition. To identify specific entities in the biological field, we need to: (1) perform word segmentation, part-of-speech tagging, and named entity recognition; (2) put the labeled words into the entity library, and if they are not in the entity library, mark them; (3) for some non-entity parts, adopt certain word combination and part-of-speech rules, scan all the segmented words in O(n) time, and filter out the parts that cannot be biological entities according to their parts of speech; (4) for the remaining words and word combinations, match the entities in the entity library that have been classified. If no entity is matched, or the matched entity belongs to category 0 (non-entity), then filter it out.

[0039] Entity Classification: We split the crawled data into six parts: title, openTypeList, detail, image, baseInfoKeyList, and baseInfoValueList. Using the KNN algorithm as the classifier, we standardize four small similarities to obtain two page similarities, and select a label by comparing the similarity between the unlabeled words and the words in the entity library (select the one with the most occurrences as the label among the top K processed words). It can also be improved based on the characteristics of biological data. An obvious drawback of traditional KNN is that when the sample distribution density is uneven, only considering the first K nearest neighbors in order without considering their distances will cause misjudgment and affect the classification performance. Through experiments on biological data, we found that since some categories of training samples are easier to obtain than others, it often leads to an imbalance in the number of training samples among different categories. That is, even if the number of training samples in each category is basically close, due to the different sizes of the regions they occupy, the distribution of training samples is still uneven. Currently, the improved methods include uniformizing the sample distribution density; improving the decision rule of KNN to solve the problem of the decline in the classification performance of the KNN classifier when the data distribution of each category is uneven. Or using a large number of nearest neighbor sets to replace the single set in KNN, and obtaining a relatively reliable support value by accumulating the support degrees of the nearest neighbor data sets for different categories, thereby improving the nearest neighbor decision rule.

[0040] Among them, Figure 4 is a schematic diagram of the multi-page similarity calculation process of the embodiment of the present invention, as Figure 4As shown, the cosine similarity of word vectors is used for the title (the word vectors calculated by fasttext can avoid out-of-vocabulary). The sum of the IDF values of the baseInfoKeys with the same average cosine similarity of the word vectors between the openTypes of the two groups (since the contribution of attributes such as "Chinese name" should be relatively small), the number of baseInfoValues that are the same under the same baseInfoKey. When predicting a page, since KNN needs to compare this page with all pages in the training set, the complexity of each prediction is O(n), where n is the size of the training set. In this process, we can count the IDF values, means, variances, and standard deviations of the similarities of each category, and then standardize the four similarities: (x - mean) / variance. The weighted sum of the similarities of the four parts is the final similarity between the two pages, and the weights are controlled by the vector weight, which is obtained through 10-fold cross-validation + grid search.

[0041] Step S102: Classify the entities in the classified entity library according to a predetermined rule and a dictionary and based on a deep learning method, and extract relationships of the entities in the extended entity library through relation extraction based on PCNN to obtain triples.

[0042] Among them, entity relationship classification is a multi-classification task. Currently, pre-trained language models, such as BERT, have refreshed the best level in various NLP tasks. However, the baseline BERT has problems such as being sensitive to complex relationships, unable to understand entity semantics, and consuming a lot of time and hardware when facing biological data. Therefore, the embodiments of the present invention adopt a rule-based method to process complex relationships, and use entity-level MASK technology and a lightweight improved BERT model to adapt to the biological relationship classification task.

[0043] Acquiring triples based on rules and dictionaries: Rule-based methods mostly use manual tagging, and the selected features include statistical information, punctuation, keywords, indicator words and direction words, position words (such as tail words), center words, etc., with pattern and string matching as the main means. Most of these systems rely on the establishment of knowledge bases and dictionaries. Among various mature word segmentation tools, LAC has high support for quantifiers, so we first use the LAC tool to segment and tag the original text. We directly treat quantifiers and adjectives in the part of speech as relations, attributes as the objects of the triples, and the subjects as the names of the animals and plants we choose. The key part here is to design and specify regular templates. For example, when the word "color" is matched, we suspect that it is a word indicating color. When the word contains words containing quantity such as meters and kilograms, we suspect that it is a word indicating height and length. We treat all these suspicious words as correct relations first, and then clean out the correct ones. The traditional regularization method ends here, but we then find that:

[0044] Deep rule tree learning: Using the advantages of neural networks to make up for the shortcomings of traditional rules, we first use mathematical logic as a basis, such as defining a logic rule tree using "and" and "or" as examples, that is, using a similar neural network structure to express logic rules. Each layer contains two sublayers, which are used to represent logic and and or respectively. Different from neural networks, each node is no longer an activation unit, and the corresponding value is not a real value, but a logic unit and a 0 / 1 binary value. We call the tree structure constructed in this way a rule tree.

[0045] After having a rule tree, the next question is how to learn this rule tree. Figure 5 This is a schematic diagram of the learning of the logic rule tree of the embodiment of the present invention. The key to learning the rule tree is to learn the law of association from input to output at each layer, that is, the weight of each edge. Here, a multi-layer perceptron is used for learning, such as Figure 5 As shown on the right, the multilayer perceptron uses a fully connected approach, such as Links and all a1-a4, Links and all Then use the existing gradient descent method to learn. After learning is completed, it is converted into the logic rule tree on the left. Here, a threshold is set, such as 0.5. If it passes Figure 5 If the weight of the edge learned by the multilayer perceptron on the right is greater than the threshold, we believe that there should be an edge link in the logic tree, otherwise the corresponding edge is deleted.

[0046] The embodiments of the present invention adopt the method of "pre-training + fine-tuning" for self-supervised training tasks. On the BERT model, domain data is added to continue training the well-trained representation model to achieve domain adaptation and match the domain business scenarios. Then, these learned representation models are used for other tasks. Figure 6 It is a schematic diagram of biological relationship classification based on the "pre-training + fine-tuning" technology in the embodiments of the present invention.

[0047] The entity-level MASK technology is adopted. MASK is similar to the traditional cloze task. At the input end, some words are randomly "masked". At the output end, the model is required to predict these "masked" words. Initially, the model does not know which words to predict. Therefore, the embedding representation of each word output by it covers the semantic information of the context to accurately predict the masked words. As Figure 7 shown on the left, a single word "masked" by the baseline BERT. To learn the overall semantic representation of the entity, when randomly "masking", directly select the word corresponding to the entity to be "masked". Figure 7 It is an improved schematic diagram of the MASK method in the embodiments of the present invention. As Figure 7 shown, the key of entity-level mask is to find the accurate entity to determine the semantic range of the mask. Usually, the entity linking method is used to link to the entities in the KG. However, in this project, many common senses are composed of non-normalized texts or free texts such as situational vocabulary and user opinions. To facilitate the linking of common sense entity mentions and the common sense KG, and to avoid the integration of expensive knowledge retrieval and fine-tuning and reasoning, it helps to generalize other downstream tasks. At the same time, it greatly reduces the dependence on the performance of the linking algorithm. We only expose the LMs to structured information before training. Specifically as follows: (1) Construct a list of MASK candidate objects. First, use the concepts / phrases in the common sense KG and the named entities in the domain KG to form an entity "vocabulary", and then detect all entity mentions e i appearing in the text segment U from the corpus to form a set E of linked MASK candidate objects. Each entity mention e i is appended with the corresponding index label M ei . For fast detection, an effective inverted index in lemma-based fuzzy matching can be used as our entity linking system. (2) Calculate the importance of MASK entities, mainly based on measuring the importance of each token for downstream tasks. First, calculate the importance S(w i ) of the word on a small-scale fine-tuning dataset, which can be obtained in an active learning manner. S(w i) is obtained on the fine-tuning dataset and cannot be directly used for the pre-training dataset. It is necessary to apply the above method to generate a small-scale dataset with important tokens annotated, and fine-tune the GenePT model to learn the implicit rules for selecting important tokens in the data.

[0048] For fast training, this project plans to adopt a 4-level lightweight scheme - KG-mask filtering corpus, low-precision quantization, model pruning, and model distillation. Trivial and irreducible entities may be difficult to contribute to model learning. This project plans to use the entity importance filtering corpus obtained by the KG-mask method, enabling the model to be trained very effectively. Task MASK uses a mixed floating-point precision (FP16 or even INT8, binary network) representation instead of the original precision (FP32) representation in model training and inference. The sentences involved in this project contain limited semantic information, and a lower-layer Transformer structure, such as within 4 layers, can be considered during fine-tuning to significantly reduce the number of model parameters and training and fine-tuning time. Model distillation is to compress these large pre-trained models (teacher models) into small student models using knowledge distillation to reduce their storage and computational costs and speed up the inference time. Figure 8 It is a schematic diagram of model distillation for an embodiment of the present invention, as Figure 8 shown. The source domain and the target domain sampled from P act together on the objective function in the generator to generate the data required for pre-training. The trained data is respectively put into the teacher model and the student model. The number of layers of the student is half of that of the teacher model each time, and the student model is used to learn the prediction results of the teacher model. The predicted results are compared with the actual labels, and after being selected by the reinforcement selector, they are used to supervise the student model to learn from the teacher model and correct the objective function.

[0049] Figure 9 It is a flowchart of relation extraction for an embodiment of the present invention. Statement parsing mainly obtains the syntactic and lexical information of the statement by generating the syntactic analysis tree of the statement. Figure 10 It is the syntactic analysis tree for an embodiment of the present invention, as Figure 10 shown. The rule symbols in the analysis tree are annotated according to the syntactic analysis tree annotation set. nn represents a common noun, vc represents "is", QP represents a quantifier phrase, PU represents a punctuation mark, and NP represents a noun phrase. The phrase structure tree we created is as follows: for the sentence "Bamboo is a tropical plant", first, we can obtain the phrases [Banana], [is a tropical fruit], then we can get the phrases [is], [a tropical fruit], and so on. Finally, we can get [Banana][is][a][kind][tropical][fruit].

[0050] Step S103: Fuse the entities in the triple through a representation learning method based on bidirectional supervision and iterative fusion and a knowledge graph-based embedding model to obtain a biological domain knowledge graph.

[0051] We map existing knowledge to rich unstructured text to generate a large amount of training data. First, a feature vector is created for each type of relationship. That is, using word2vec, the features of the sentence are directly extracted. If there are 10 sentences containing this relationship and each sentence can extract 3 features, then the feature vector of this relationship contains 30 features. Even if some of the extracted features cannot express this relationship, there are many features in the feature vector. Perhaps most of the features or a combination of some features can still effectively express this relationship, which is the guarantee of the effectiveness of distant supervision. Train a classification model.

[0052] Distant supervision learning is an important method for relation extraction, but it has a major drawback: the problem of mislabeling. This kind of error will spread to the relation extraction model, and the spread and accumulation of errors will reduce the performance of the model. To solve this problem, we use a CNN model. However, traditional lexical features generally include the part-of-speech of the entity itself, the word sequence before the entity, the hypernym of the word, etc., and are relatively dependent on manual feature engineering. Therefore, we improve from the features at the lexical level and the features at the syntactic level. Figure 11 Schematic diagram of the automatic learning text feature framework for the embodiments of the present invention, as Figure 11 shown. Automatically learn text features by converting words into word vectors and using word vectors to represent the features at the lexical level. Considering the context features, use a sliding window. Assume the window size is 3. Each time, look at 3 words, and then slide one grid backward. Finally, a set of sentence-level features is obtained.

[0053] Learning from the advantages of segmental pooling, the embodiments of the present invention further improve the above model by adopting the segmental pooling idea of PCNN. Figure 12 PCNN model with segmental pooling for the embodiments of the present invention, as Figure 12 shown. The sentence is segmented into 3 segments according to the positions of the target entity pairs, and the maximum pooling operation is performed separately on each segment. Each segment obtains a maximum value, and finally all the maximum values are concatenated to form a feature vector.

[0054] Traditional fusion algorithms are mainly based on linguistic similarity, highly dependent on text information, ignore the structural information between entities, have a high time complexity of the algorithm, and lack scalability. The embodiments of the present invention adopt a knowledge representation learning algorithm. Figure 13 Schematic diagram of the knowledge representation learning for the embodiments of the present invention. The main research ideas include:

[0055] 1) Automatically mine entity features: Knowledge representation learning uses machine learning to represent the semantic information of entities as dense low-dimensional real-valued vectors, mining the structural features between entities. The similarity between entities can be calculated without manually setting matching attributes and similarity metric functions.

[0056] 2) Improve the fusion efficiency and be applicable to large-scale datasets: When the dataset is large, a representation learning model with a lower time complexity can be used, such as TransE. The calculation of its loss function only involves the calculation of the energy functions of legal triples and wrong triples.

[0057] 3) To avoid KG fusion, use Seeds data to perform bidirectional supervision on the KG to be fused, and continuously expand the Seeds training data through multiple iterative fusion operations, and use the newly fused entity pairs to further strengthen the fusion effect.

[0058] Knowledge representation learning: First, use existing representation learning models to train RDF triples. For example, typical embedding representation models such as TranE and PtransE map entities and relationships to the same low-dimensional vector space. Representation learning methods can also be designed separately for biological data features. In the embodiments of the present invention, the SimpLE model with good current effects and easy to train is adopted. SimpLE is proposed based on the tensor factorization method CP. It uses the inverse of a relationship to handle the independence of two vectors in CP. That is to say, for an entity e, there are two vectors he and te, and for a relationship r, there are two vectors v and v-1. For SimpLE: (1) Each entity is represented as two vectors: a head entity vector and a tail entity vector (each vector is independent); (2) Each relationship is represented as two relationships: a forward relationship and an inverse relationship vector. In the learning process, negative samples must be set. After taking out a small batch of positive samples each time, we randomly change the head vector or the tail vector to generate the corresponding negative samples. The significance of generating such negative samples is also for the necessity of learning. Because a supervision flag is required in the learning process, we generate the corresponding negative samples so that the score of the positive samples is higher than that of the negative samples, and use this as a standard to generate the corresponding loss.

[0059] Bidirectional supervision training: Use pre-fused entity pairs (Seeds) to perform bidirectional supervision on two knowledge graphs to be fused, and generate constraints on the training of their respective vectors. Before training, it is necessary to initialize the vectors of the pre-fused entity pairs, that is, first train the entities separately in their respective knowledge graphs and then take the average value. Figure 14 For the bidirectional supervision training and iterative fusion technical route schematic diagram of the embodiments of the present invention, the basic idea of the bidirectional supervision training algorithm is as Figure 14As shown on the left, the vector values of the pre-fused entities in the Seeds set remain unchanged during the training process, without gradient descent optimization. Only the vector values of other unfused entities are updated.

[0060] Iterative fusion: After each round of fusion operation, the Seeds set is updated. The fused entity pairs with higher confidence are inserted into the Seeds and used as the training data for the next round, as Figure 14 shown on the right. Through knowledge representation learning and bidirectional supervised training, the entities and relationships in the knowledge graph to be fused can be uniformly mapped to a vector space, and then the next fusion operation is performed according to the distance between the entity vectors. The distance calculation can use L1, L2 norms, etc. For each unfused entity e1 in KG1, the entity e2 with the closest distance to it is found in KG2. At the same time, a threshold is defined. If the distance is less than the threshold, it is reasonable to believe that e1 and e2 are two close entities. Conversely, the similarity between the two is low. Among them, (e1, e2) is called a new fused entity pair. Intuitively, the new fused entity pair can further improve the bidirectional supervised training and find more fused entities. Therefore, the embodiment of the present invention adopts an iterative fusion method to perform the fusion operation multiple times. The iterative stop condition is generally to reach the maximum number of iterations or no new fused entity pairs are found.

[0061] This algorithm needs to read in the pre-fused data (Seeds) and the knowledge base to be fused at the same time. Before the algorithm executes, the data needs to be preprocessed, that is, each relationship and entity is replaced with a unique integer value. If PtransE is used, the data preprocessing also includes path calculation and path confidence calculation. After the preprocessing is completed, the pre-fused data is used to perform bidirectional supervised training on the knowledge graphs KG1 and KG2 to be fused. The continuously expanding Seeds set is the final fusion result. The three main steps of the fusion algorithm are not sequential processes, but a combination relationship.

[0062] Device Embodiment 1

[0063] According to an embodiment of the present invention, a device for constructing a knowledge graph in the biological field is provided. Figure 15 It is a schematic diagram of the device for constructing a knowledge graph in the biological field according to an embodiment of the present invention, as Figure 15 shown. The device for constructing a knowledge graph in the biological field according to an embodiment of the present invention specifically includes:

[0064] An entity recognition and classification module that uses a lexical analysis tool to perform word segmentation, part-of-speech tagging, and entity recognition on the biological field text content to obtain an initial entity set, obtains an extended entity set by comparing the initial entity set with a predetermined entity library, names the extended entity set, and classifies the entities in the named extended entity set;

[0065] The relationship extraction module classifies entity relationships for the classified entities in the extended entity library through rule- and dictionary-based methods and deep learning-based methods, and performs relationship extraction on the entities in the extended entity library through PCNN-based relationship extraction to obtain triples;

[0066] The knowledge graph acquisition module fuses the entities in the triples through a representation learning method based on bidirectional supervision and iterative fusion and a knowledge graph embedding model to obtain a biological domain knowledge graph.

[0067] Device Embodiment II

[0068] An embodiment of the present invention provides a device for constructing a biological domain knowledge graph, including: a memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor 52, the steps described in the method embodiment are implemented.

[0069] Device Embodiment III

[0070] An embodiment of the present invention provides a computer-readable storage medium, on which an implementation program for information transmission is stored. When the program is executed by a processor, the steps described in the method embodiment are implemented.

[0071] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.

[0072] The above are only the embodiments of this document and are not used to limit this document. For those skilled in the art, various changes and modifications can be made to this document. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this document shall be included within the scope of the claims of this document.

Claims

1. A method for constructing a biological domain knowledge graph, characterized in that, it includes: Using a lexical analysis tool to perform word segmentation, part-of-speech tagging, and entity recognition on the biological domain text content to obtain an initial entity set, comparing the initial entity set with a predetermined entity library to obtain an extended entity library, naming the extended entity library, and classifying the entities in the named extended entity library; According to predetermined rules and dictionaries and based on a deep learning method, classify the entities in the extended entity library after classification, and perform relationship extraction on the entities in the extended entity library through relationship extraction based on PCNN to obtain triples; Fuse the entities in the triples through a representation learning method based on bidirectional supervision and iterative fusion and an embedding model based on the knowledge graph to obtain a biological domain knowledge graph; The specific process of classifying the entities in the extended entity library through a method based on rules and dictionaries and a deep learning method specifically includes: By setting a regular template, using the LAC tool, perform word segmentation and part-of-speech tagging on the original text, use the obtained numeral classifier and adjective as the relationship of the triple, and use the obtained attribute as the object of the triple; Construct a rule tree, and learn the rule tree through the gradient descent method to obtain the relationship between two objects; Use the BERT model for entity relationship classification; The specific process of performing relationship extraction on the entities in the extended entity library through relationship extraction based on PCNN to obtain triples specifically includes: Obtain training samples through a remote supervision method that integrates syntactic analysis, and improve the relationship model based on remote supervision learning by using the segment pooling method of PCNN to obtain triples.

2. The method according to claim 1, characterized in that, The specific process of obtaining the extended entity library by comparing the initial entity set with a predetermined entity library specifically includes: Put the words that have been labeled in the initial entity set into the entity library. If the words obtained through entity recognition are not in the predetermined entity library, label them to form unlabeled words; For the non-entity part in the initial entity set, use specific word combinations and part-of-speech rules to scan all word segments within a specific time complexity, and filter out the parts that cannot be biological entities; Filter out the remaining words and word combinations in the initial entity set.

3. The method according to claim 2, characterized in that, The specific process of classifying the entities in the named extended entity library specifically includes: By comparing the similarity between the unlabeled words and the entities in the extended entity library, label the unlabeled words.

4. The method according to claim 3, characterized in that, The further process of using the BERT model for entity relationship classification includes: By using the distillation method, compress the key pre-trained model of the BERT model into a small student model to reduce the number of parameters for training the BERT model.

5. The method according to claim 1, characterized in that, The specific process of fusing the entities in the triple through representation learning based on bidirectional supervision and iterative fusion and the embedding model based on the knowledge graph includes: Use the SimplE learning model to train the triple, fuse the entities in the triple, and obtain a specific number of knowledge graphs to be fused; Select two entities in two knowledge graphs to be fused as the pre-fusion entity pair, perform bidirectional supervision on the two knowledge graphs to be fused through the pre-fusion entity pair, continuously expand the pre-fusion entity pair through multiple iterative fusion operations, and use the expanded pre-new fusion entity pair to further enhance the fusion effect.

6. A device for constructing a biological domain knowledge graph, characterized in that, it includes: An entity recognition and classification module that uses a lexical analysis tool to segment, perform part-of-speech tagging, and entity recognition on the biological domain text content to obtain an initial entity set, obtains an extended entity library by comparing the initial entity set with a predetermined entity library, names the extended entity library, and classifies the entities in the named extended entity library; A relationship extraction module that classifies entity relationships of the classified entities in the extended entity library through methods based on rules and dictionaries and methods based on deep learning, and extracts relationships of the entities in the extended entity library through relationship extraction based on PCNN to obtain triples; A knowledge graph acquisition module that fuses the entities in the triple through a representation learning method based on bidirectional supervision and iterative fusion and an embedding model based on the knowledge graph to obtain a biological domain knowledge graph; The specific process of classifying entity relationships of the entities in the extended entity library through methods based on rules and dictionaries and methods based on deep learning includes: By setting a specified regular template and using the LAC tool, segment and perform part-of-speech tagging on the original text, use the obtained numeral classifier and adjective as the relationship of the triple, and use the obtained attribute as the object of the triple; Construct a rule tree and learn the rule tree through the gradient descent method to obtain the relationship between two objects; Use the BERT model for entity relationship classification; The specific process of extracting relationships of the entities in the extended entity library through relationship extraction based on PCNN to obtain triples includes: Obtain training samples through a remote supervision method that integrates syntactic analysis, and improve the relationship model based on remote supervision learning through the segment pooling method of PCNN to obtain triples.

7. A device for constructing a biological domain knowledge graph, characterized in that, it includes: A memory, a processor, and a computer program stored on the memory and executable on the processor. When the computer program is executed by the processor, it implements the steps of the biological domain knowledge graph construction method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, An information transfer implementation program is stored on the computer-readable storage medium. When the program is executed by the processor, it implements the steps of the biological domain knowledge graph construction method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method and system for automatically constructing knowledge maps for mass unstructured texts

    CN108875051A

  • Protein knowledge graph vectorization method

    CN113963748A