A knowledge graph system for intestinal microorganisms
By constructing a multi-theme, multi-source, and multi-modal intestinal microbial knowledge graph system, integrating multiple public databases and using knowledge graph embedding models and multi-modal annotations for uncertain reasoning, the existing system's problems of narrow knowledge coverage, low integrity and simple reasoning are solved, and the effects of wide knowledge coverage, high integrity and strong reasoning ability are achieved.
Patent Information
- Application Number
- CN202410601024.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-05-15
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2044-05-15
AI Technical Summary
The existing intestinal microbial knowledge graph system has problems such as narrow knowledge coverage, low knowledge integrity, fixed knowledge structure and simple reasoning, and cannot effectively perform general analysis of multiple diseases, insufficient data volume, inability to conduct multimodal reasoning, lack of confidence score assignment, and limited prediction range of traditional inference models.
A multi-themed, multi-source, multi-modal intestinal microbial knowledge graph system was constructed. By integrating multiple public databases, the ‘intestinal microbial knowledge base’, ‘intestinal microbial small molecule drug treatment association knowledge base’ and ‘clinical medical database’ was constructed, and the knowledge graph embedding model and multi-modal annotation were used for uncertainty reasoning.
A knowledge graph system with wide knowledge coverage, high integrity, rich modality and confidence score assignment is realized, which solves the problems of cold start and long-chain reasoning, improves the prediction range and interpretability, and overcomes various shortcomings of traditional systems.
Smart Images

Figure CN118506887B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of medicine, treatment, etc., and specifically, to a knowledge graph system and reasoning system in the field of intestinal microorganisms. Background Art
[0002] The intestinal flora is the largest microbial community in the human body. An adult carries about 200 grams of intestinal microbes. The genes they contain are about 30 times that of the human genome, and they can even be regarded as an important organ alone. Intestinal microbes act directly or indirectly on other important organs of the human body through pathways such as the gut-brain axis, gut-liver axis, and gut-lung axis. Therefore, the stability and imbalance of intestinal microbes are closely related to human health and disease. They can help human hosts improve their immunity, promote beneficial metabolism, and control weight gain, etc., and may also lead to the occurrence and development of diseases such as inflammatory bowel disease and diabetes. In recent years, with the emergence and popularization of high-throughput sequencing technology, the threshold for research on the human microbiome has been significantly lowered. Multi-omics and clinical research on intestinal microbes have become a very popular field at present. Many studies have conducted in-depth experimental verification and analytical discussions on various associations. However, due to the high variability of intestinal microbes, the current findings may only be the tip of the iceberg. More specific functional characteristics and mechanisms of intestinal microbes on human hosts still need further research and discovery.
[0003] In order to fully understand the important role of intestinal microorganisms in human health and disease, in addition to reliable but time-consuming and laborious experimental methods and efficient but highly limited microbiome methods, researchers should also pay more attention to biomedical big data. With the high frequency of publication of cutting-edge literature, the flourishing of professional databases and the massive accumulation of clinical data, biomedicine has gradually shifted from "evidence-based medicine" driven by subjective experience to "precision medicine" driven by objective data. At present, a huge biomedical knowledge body has been formed. How to effectively store them and conduct deep mining and fusion reasoning of knowledge will be a huge challenge. In recent years, the knowledge graph, as a heterogeneous network, has gradually replaced the non-relational network that cannot distinguish complex biological effects, and has been more used in the simulation of knowledge systems in the biomedical field. As a way of storing knowledge, the knowledge graph connects massive knowledge scattered in various carriers in the form of triples (head entity, relationship, tail entity), which has attracted widespread attention from academia and industry.
[0004] However, the real world is changing with each passing day, and the existing knowledge graphs are almost incomplete. Using existing knowledge to infer new reliable knowledge is very important for improving the integrity and extensibility of knowledge. Among them, inductive reasoning based on knowledge graph embedding has become the mainstream in fields such as biomedicine. It maps entities and relationships to a low-dimensional space, significantly improves computational efficiency, effectively alleviates data sparsity, and can also achieve the fusion of heterogeneous information. Knowledge graph embedding can be subdivided into translation models represented by RotatE, semantic models represented by HolE, and neural network models represented by ConvKB. Multimodal knowledge graph embedding is another major development direction, such as RSN combined with relationship paths, KEPLER combined with entity descriptions, TKRL combined with entity categories, and IKRL combined with entity images. In addition, there are UKGE and UKGSE that combine confidence scores for uncertainty knowledge graph embedding. They play an important role in downstream tasks such as drug reuse and disease gene association prediction.
[0005] In recent years, many teams have tried to build some Gut Microbe Knowledge Graphs (GMKGs), such as Food4healthKG, KG4NH, MiKG4MD and MGMLink, but there are no mature projects for implementation. Looking at these works, although they have inspired the development of this field in different ways, they have more or less the following defects:
[0006] (1) Most of them only analyze one or two diseases, especially mental health diseases, which account for the majority of the total number of diseases. The scope of application is very limited. The knowledge coverage should be expanded, no longer just for a single disease, but to be universal for multiple diseases as much as possible, because the knowledge systems of different diseases are sometimes interoperable;
[0007] (2) Most of them only extract structured knowledge from professional databases, or manually extract knowledge from a small amount of relevant literature, resulting in a small amount of data, insufficient knowledge integrity, and inability to obtain convincing conclusions;
[0008] (3) After the construction work is completed, most of them only visualize and query the knowledge graph based on graph databases such as Neo4j, but cannot obtain reliable new knowledge based on existing knowledge reasoning. Therefore, it is even more impossible to perform more complete and diverse multimodal reasoning based on annotation information other than triples;
[0009] (4) All that is constructed are deterministic knowledge graphs. A piece of knowledge has only two situations: existence or non-existence. There is no confidence score assigned to it based on the actual situation, or the setting of the confidence score is unreasonable. Therefore, it is even more impossible to perform more reasonable uncertainty reasoning based on the confidence score;
[0010] (5) Traditional knowledge graph reasoning models can only predict existing relationships, which means that the knowledge graph is incomplete and its prediction model is very fixed;
[0011] (6) Traditional knowledge graph reasoning models can only predict existing entities, that is, hot start reasoning, and their prediction range is relatively limited. Summary of the invention
[0012] The first aspect of the present invention relates to a method for constructing a knowledge graph
[0013] 1. Obtaining the original data of the “Gut Microbiome Knowledge Base” (Bacteria-Bacteria Knowledge Base)
[0014] Integrate the intestinal microbial knowledge base, obtain information from public databases, perform data cleaning and de-redundancy processing, and retain intestinal microbial related knowledge; extract entity tables and relationship tables, and further include extracting attribute tables; use entity tables to uniformly match entities in relationship tables and attribute tables to build a complete "intestinal microbial knowledge base".
[0015] Furthermore, the entity table and the relationship table may be extracted from the entire database or a part of the database may be selected to extract the entity table and the relationship table, and the entire database or a part of the database may be merged and de-redundant to extract the attribute table.
[0016] The entity table means professional terms in macromolecular fields such as intestinal microorganisms, genetic engineering, protein engineering, enzyme engineering, biochemistry, etc., including but not limited to genes, proteins, RNA, small molecules, pathways, reactions, nucleotides, mutations, etc., and can be expressed in various languages, such as Chinese, English, German, French, Japanese, etc. Specifically, each entity corresponds to at least one proprietary term in the biomedical field, such as a special noun; but it is not necessarily a noun such as "gene" or "protein", it can also be a verb, verb-object phrase, etc., such as "protein expression" or "gene modification", or a short sentence containing a few prepositions such as "purify the protein", etc., but they can all express a relatively complete meaning, but are not sufficient to express the meaning of the whole sentence; preferably, nouns are the preferred entities; preferably, Chinese and / or English are the preferred languages.
[0017] The relationship table refers to the horizontal associations between the entity contents in the entity table, including but not limited to the associations of therapeutic pathways, the associations of gene expression patterns, the associations of chemical modifications, the activation pathways of proteins, the inhibition pathways of proteins, the auxiliary associations of proteins, etc.;
[0018] The attribute table refers to the inherent attributes, superordinate attributes, subordinate attributes, etc. of each entity in the entity table, including but not limited to biological classification, text description and definition, genomic sequence, proteomic sequence, microbiological classification, 16S rRNA sequence, 168 rRNA sequence, etc.
[0019] The entity refers to the actual content in the entity table, relationship table, and attribute table, rather than the content that completely matches the text. Based on the general understanding of text, the entity includes various language translations of the actual content, synonyms, common aliases, common expressions, and order changes that do not affect the understanding of the meaning.
[0020] Specifically, the original database can be BioCyc, EcoCyc, MetaCyc, BacDive, BV-BRC, EMBL-EBI, NCBI-Taxonomy and MicrobeWiki, etc. Specifically, the retained intestinal microorganism-related knowledge is not limited to the number of intestinal microorganisms. If it needs to be limited, it can be 100-500, 200-400, 200-300, 200, 300 kinds of intestinal microorganism-related knowledge; just for example, the eight databases can be cleaned and de-redundanted first, and the relevant knowledge of 200 kinds of intestinal microorganisms can be retained. After the individual processing is completed, the five databases of BioCyc, EcoCyc, MetaCyc, BacDive and BV-BRC are merged and de-redundant, and the entity table and relationship table are extracted from them. At the same time, the three databases of EMBL-EBI, NCBI-Taxonomy and MicrobeWiki are merged and de-redundant, and the attribute table is extracted from them. Finally, the entity table is used to uniformly match the entities in the relationship table and the attribute table to form a complete "intestinal microbial knowledge base" (see Figure 2 , Figure 2 The "Zhuowei Knowledge Base" in the above is the specific name of the "Intestinal Microbial Knowledge Base" recorded above)
[0021] The "Gut Microbiome Knowledge Base" has been reviewed by experts in the field, so the confidence score is set to 1.0.
[0022] 2 Construction of the “Gut Microbiome Small Molecule Drug Therapy Association Knowledge Base” (Bacteria-Human Knowledge Base)
[0023] It includes searching relevant databases, indexing the literature, obtaining a small database after indexing, dividing it into a training set and a validation set, using a non-generative pre-trained language model for training and evaluation, obtaining a three-classification model and obtaining the probabilities of the three relationships through functions, and obtaining confidence scores through calculations.
[0024] Specifically, the database related to intestinal microorganisms and small molecule drugs for the treatment of human diseases is retrieved, related literature is downloaded, the literature is indexed, and the annotation tool platform, AI technology, and manual indexing are used for indexing and review (furthermore, it is preferred to use the annotation tool annotation platform for indexing, AI technology for automatic annotation, and manual review of the annotation results). After indexing, a small database is obtained, which is divided into a training set and a validation set. A non-generative pre-trained language model is used for training and evaluation to obtain a three-classification model and obtain the probabilities of the three relationships through functions. Confidence scores are obtained through calculation, and finally a "knowledge base on the association of intestinal microorganisms with small molecule drugs for treatment" is obtained.
[0025] The annotation refers to extracting the association between intestinal microbial entities and small molecule entities and disease entities based on the entities annotated by the annotation tool platform; specifically, it can be judged whether it is a potential entity for relationship extraction based on the position, interval, frequency of occurrence, etc. of the entity in the article. The position of the entity can be in the same paragraph (with a carriage return as a pause), a sentence (with a period as a pause), a paragraph type or a field (such as abstract, title, keyword, background technology, method, discussion, etc.); the interval can be a specific number of characters, such as 1-10, 2-6, 2-8 characters, or more characters, which can ensure that multiple entities are guaranteed in a specific order or an unlimited order; the frequency of occurrence refers to the frequency of the entity appearing in the article, such as 1 time, 2 times, 3 times... etc. For example, if it only appears once, it can be judged as having a low relevance. It can be in the full text, or in a specific paragraph type, or it can be combined, such as appearing once in the abstract and more than 5 times in the full text. You can choose one of the above multiple judgment criteria for annotation, or you can choose multiple judgment criteria for comprehensive judgment.
[0026] The comprehensive confidence score is ;
[0027] The last layer of the relation extraction model is a three-classification model. After a SoftMax function, the probability of the paragraph belonging to the three relations is obtained. This probability can be directly used as the confidence score of the non-Unrelated triple. Assume that the triple Total Source of the article, y i 、j i and i Respectively represent Articles The publication year of the source article, the journal impact factor, and the confidence score given by the relationship extraction model. The comprehensive confidence score of the triple is defined as , where k n , k yi and k ji Respectively based on ,y i and j i The linear segmentation weighting coefficients are monotonically increasing, they are all less than 1 and the turning points of the linear segmentation need to be manually defined empirically; finally, the Unrelated triples predicted to be unrelated are deleted, the non-Unrelated triples are triples predicted to be related, and the triples with confidence scores greater than or equal to a certain value (for example, greater than or equal to 0.5) are retained to obtain a four-tuple containing "head entity, relationship, tail entity, confidence score".
[0028] For example, the association between intestinal microorganisms and diseases can be defined as activate or inhibit. Non-Unrelated triples are examples such as (a certain bacteria, activate, a certain disease), while the quadruples obtained after further screening are examples such as (a certain bacteria, activate, a certain disease, confidence score 0.8).
[0029] The “database for searching the association between intestinal microorganisms and small molecule drugs for treating human diseases” may be a PubTator3 original xml file containing PubMed abstracts and PMC full texts;
[0030] The annotation tool platform may be PubTator3 officially developed by NCBI, which is used to annotate entities and some relationships in the article; the annotated entities include chemical small molecules, Disease diseases, Gene genes, Species species and Variant variants and other related entities and some relationships; but the annotated relationships do not include relevant knowledge of the "intestinal microbial knowledge base";
[0031] Specifically, the training set and the validation set are divided according to a specific ratio, for example, 6:2, 7:1, 8:1, 9:2, 9:1, 8:2;
[0032] Said Figure 3 The establishment of the “Gut Microbiome Small Molecule Drug Therapy Association Knowledge Base” can be found in Figure 3 .
[0033] 3. Building a “clinical medicine database” (human-human knowledge base)
[0034] The clinical medicine database uses molecular biology to provide data from authoritative clinical medicine databases, such as the PMapp database, to provide relevant knowledge such as disease diagnosis, medication guidelines, and disease prevention, providing knowledge support for the intestinal microbial knowledge map (see Figure 1 , Figure 1 The GMKG-200 shown in the figure is the "Gut Microbiome Knowledge Graph")
[0035] Furthermore, the database integrates 61 authoritative databases of molecular biology and clinical medicine. In view of the authority of the database and the fact that it has been reviewed by experts in the field, the confidence score is defined as 1.0.
[0036] 4. Construction of intestinal microbial knowledge graph
[0037] The three knowledge bases of "Gut Microbiome Knowledge Base", "Gut Microbiome Small Molecule Drug Therapy Association Knowledge Base" and "Clinical Medicine Database" are combined and connected to align small molecules, drugs and disease entities to obtain the final gut microbiome knowledge graph.
[0038] The second aspect of the present invention relates to a knowledge graph multimodal uncertainty reasoning system
[0039] The knowledge graph multimodal uncertain reasoning system of the "Gut Microbiome Knowledge Graph" is used to predict potential associated diseases, drugs, genes, etc. for intestinal microorganisms.
[0040] First, the head entity, tail entity, and relation each have a randomly initialized embedding vector, and each have an embedding vector based on multimodal annotations, and the vectors are normalized based on the expression matrix;
[0041] Secondly, a dynamic gating network of a hybrid expert system is used to fuse the two embedding vectors of each entity to obtain its final vector representation;
[0042] Finally, the final vector representation of the triplet is brought into the basic knowledge graph embedding model and the conversion function to generate the predicted confidence score, which is compared with the true confidence score to define the loss function.
[0043] Specifically, the initialization vector is , and ; The embedding vector is and ; The multimodal feature vector extraction methods for different entity types are different; the vector normalization processing can be vector normalization processing of intestinal microbial entities based on gene expression matrix, which means that entities perform multi-category vector average pooling processing based on ontology categories, or use encoding software for encoding processing.
[0044] For example, gene entities are processed by multi-category vector average pooling based on gene ontology categories, disease entities are encoded based on text descriptions using PubMedBERT or BioLinkBERT, protein entities are encoded based on amino acid sequences using ProteinBERT, and small molecule and drug entities are encoded based on SMILES using ChemBERTa.
[0045] Specifically, the head entity vector For example: ,in , is a relation-based weight vector, which is also randomly initialized. The calculation of is similar. For an entity that does not exist in the "Gut Microbiome Knowledge Graph" (GMKG-200), just add or Set to zero vector, or is equal to zero, or Will be completely based on or Therefore, as long as multimodal annotations are extracted for any entity, this model can achieve cold-start reasoning.
[0046] Specifically, the final vector of the triple can be represented as Substitute the basic knowledge graph embedding model and conversion function Generate confidence scores for predictions , and compare it with the true confidence score to define the loss function.
[0047] For example, TransE represents the distance after translation in plane space: ,RotatE represents the distance after rotation in complex space: ,in represents the product of vector elements, and concat represents the concatenation of the real and imaginary parts of a complex number.
[0048] In the "Gut Microbiome Knowledge Graph (GMKG-200)", gut microbes are directly connected to diseases and drugs, and predicting new potential associations is equivalent to performing a simple knowledge graph completion task. However, gut microbes are not directly related to human genes, and can only be indirectly associated through multiple steps of diseases, drugs, and small molecules, which cannot be achieved through traditional knowledge graph link prediction. By achieving homogeneity of existing and new inference relationships based on confidence scores, it is only necessary to comprehensively calculate the probability expectation to effectively achieve confidence score estimation and correlation ranking of long-chain relationships. (See Figure 5 )
[0049] The third aspect of the present invention relates to an intestinal microbial knowledge graph system
[0050] The intestinal microbial knowledge graph system includes an “intestinal microbial knowledge base”, a “intestinal microbial small molecule drug treatment related knowledge base”, and a “clinical medicine database”;
[0051] Furthermore, the "gut microbial knowledge base" contains knowledge association information between bacteria obtained after extraction and processing from the original database, and the "gut microbial small molecule drug treatment association knowledge base" contains information related to intestinal microorganisms and small molecule drugs for the treatment of human diseases obtained after extraction and processing from the original database.
[0052] Furthermore, the "gut microbial knowledge base" is constructed by the method for obtaining the original data of the "gut microbial knowledge base" recorded in the first aspect of the present invention, the "gut microbial small molecule drug therapy associated knowledge base" is constructed by the construction method of the "gut microbial small molecule drug therapy associated knowledge base" recorded in the first aspect of the present invention, and the "clinical medicine database" is constructed using the method of "constructing a clinical medicine database" recorded in the first aspect of the present invention.
[0053] Furthermore, the intestinal microorganism knowledge graph system also includes a knowledge graph multimodal uncertain reasoning system for predicting potential associated diseases, drugs, genes, etc. for intestinal microorganisms.
[0054] Furthermore, the intestinal microbial knowledge graph system also includes an image database that can be used for query, provide visual query results, and realize open source sharing of data.
[0055] The fourth aspect of the present invention relates to a knowledge graph system construction device
[0056] The knowledge graph system construction device comprises:
[0057] The data acquisition module is used to obtain the original data related to intestinal microorganisms, specifically, to obtain the original database information of the "intestinal microorganism knowledge base", "intestinal microorganism small molecule drug treatment related knowledge base" and "clinical medicine database";
[0058] Data processing module, including cleaning and removing redundancy from data, retaining target-related knowledge, extracting entity tables, relationship tables, and attribute tables; or using annotation tool platforms, AI technology, and manual indexing to index the acquired data;
[0059] The data training module is used to divide the labeled data into training sets and validation sets, and use a non-generative pre-trained language model for training and evaluation to obtain a three-classification model and obtain the probabilities of the three relationships through functions, and obtain confidence scores through calculations;
[0060] The data matching module is used to uniformly match the entity table with the entities in the relationship table and the attribute table; and to obtain matching data after data processing through confidence scores.
[0061] Furthermore, it also includes a data reasoning module for importing the acquired data into the knowledge graph multimodal uncertain reasoning system to predict potential associated diseases, drugs, genes, etc. for intestinal microorganisms.
[0062] Specifically, it includes a vector normalization processing module. The head entity, tail entity, and relationship each have a randomly initialized embedding vector, and each have an embedding vector based on multimodal annotations. This module performs vector normalization based on the expression matrix.
[0063] Specifically, it also includes a hybrid expert system module, which is used to fuse the two embedding vectors of each entity using a dynamic gating network of the hybrid expert system to obtain its final vector representation;
[0064] Specifically, it also includes a comparison module, which is used to bring the final vector representation of the triple into the basic knowledge graph embedding model and the conversion function to generate a predicted confidence score, and compare it with the actual confidence score to define a loss function.
[0065] Furthermore, the knowledge graph system construction device may also include a display module, which can be used for querying, providing visual query results, and realizing open source sharing of data.
[0066] A fifth aspect of the present invention relates to a computer program and a storage medium
[0067] A computer program for implementing the intestinal microbial knowledge graph construction program designed in the first to fourth aspects of the present invention;
[0068] A computer storage medium can store the intestinal microbial knowledge graph construction program designed to implement the first to fourth aspects of the present invention and the established intestinal microbial knowledge graph.
[0069] In summary, the present invention constructs a gut microbial knowledge graph (GMKG) with wide knowledge coverage (multiple topics), high knowledge integrity (multiple sources), rich knowledge modalities (multimodal) and confidence scores. And based on the knowledge graph embedding model, a knowledge graph multimodal uncertainty reasoning system is established, and multimodal annotations and confidence scores are used to achieve cold start and long-chain reasoning, thereby improving the prediction range and interpretability of traditional reasoning models. It can not only help intestinal microbial researchers summarize existing authoritative and cutting-edge knowledge (build a Neo4j graph database to facilitate query and search), but also serve as a large-scale screening tool to discover new research directions from a forward-looking perspective for subsequent in-depth experimental verification (publish confidence scores and ranking lists of new reasoning knowledge), and will also achieve more accurate and faster prevention and treatment of various diseases in clinical decision-making (publish the trained entity and relationship embedding vectors as pre-training vectors for downstream clinical tasks).
[0070] Beneficial Effects
[0071] The present invention solves two major difficulties in current biomedical reasoning: cold start and long-chain reasoning. The invention extracts multimodal annotations of biomedical entities and uses a hybrid expert system to supplement the knowledge gaps of triples to achieve cold start reasoning and enhance the prediction ability for small samples and zero samples. Long-chain reasoning is achieved by calculating the comprehensive confidence score of the relationship path through the confidence score. The obtained relationship path will enhance the interpretability of biomedical knowledge and play a similar effect to a pathway. It overcomes the four main defects of narrow knowledge coverage, low knowledge integrity, fixed knowledge structure and simple knowledge reasoning. BRIEF DESCRIPTION OF THE DRAWINGS
[0072] Figure 1 This is an overview of the intestinal microbiome knowledge graph system.
[0073] Figure 2 Processing data for the "Gut Microbiome Knowledge Base".
[0074] Figure 3 To establish a "knowledge base related to intestinal microbial small molecule drug therapy".
[0075] Figure 4 Unified knowledge base construction for small molecules, drugs, and disease entities.
[0076] Figure 5 It is a multimodal uncertainty reasoning system for knowledge graphs. DETAILED DESCRIPTION
[0077] For 200 common intestinal microorganisms in the human body and commonly used in detection, the intestinal microbial knowledge graph is constructed and the multimodal uncertainty reasoning system is constructed and operated.
[0078] (1) Building a knowledge base of intestinal microbes
[0079] Integrate 8 biological categories, molecular pathways and multi-omics public databases to build a large-scale comprehensive intestinal microbial knowledge base and name it "Zhuowei Knowledge Base". Zhuowei Knowledge Base will completely download and integrate the 8 original databases of BioCyc, EcoCyc, MetaCyc, BacDive, BV-BRC, EMBL-EBI, NCBI-Taxonomy and MicrobeWiki. First, the 8 databases are cleaned and de-redundanted separately, and only the relevant knowledge of 200 intestinal microorganisms is retained. After the individual processing is completed, the five databases of BioCyc, EcoCyc, MetaCyc, BacDive and BV-BRC are merged and de-redundant, and the entity table and relationship table are extracted from them. At the same time, the three databases of EMBL-EBI, NCBI-Taxonomy and MicrobeWiki are merged and de-redundant, and the attribute table is extracted from them. Finally, the entity table is used to uniformly match the entities in the relationship table and the attribute table to form a complete intestinal microbial knowledge base ( Figure 1 The Zhuowei Knowledge Base is shown in the figure as an example).
[0080] (2) Construction of the “Gut Microbiome Small Molecule Drug Therapy Association Knowledge Base”
[0081] There is little knowledge about gut microbial interactions and their associations with small molecules and diseases in existing databases. We downloaded the original xml file of PubTator3 containing PubMed abstracts and PMC full texts to supplement these relationships. PubTator3 is an annotation tool platform officially developed by NCBI. It annotates the Chemical, Disease, Gene, Species and Variant entities and some relationships in the article, but the annotated relationships do not include Species entities, which means there is no knowledge related to gut microbes.
[0082] First, according to the entities annotated by PubTator3, if an intestinal microorganism entity appears in the same paragraph as another intestinal microorganism entity, a small molecule entity, or a disease entity, it is considered as a set of potential entity pairs for relationship extraction. Then, 2,000 paragraphs of each of the three categories of eligible paragraphs are randomly selected for relationship annotation, and each is annotated with three relationships. In addition to Unrelated, the relationship between intestinal microorganisms and intestinal microorganisms is Symbiotic or Exclusion, between intestinal microorganisms and small molecules is Secrete or Consume, and between intestinal microorganisms and diseases is Activate or Inhibit.
[0083] GPT-4 was used for two rounds of automatic annotation, and the paragraphs with inconsistent results were manually reviewed to obtain the final annotation results. When using GPT-4 for automatic annotation, a prompt word needs to be added before the paragraph to assist in text generation, and entity annotation tokens are added before and after the entity context for marking (intestinal microbial entities are @gm-1$, @ / gm-1$, @gm-2$, and @ / gm-2$, small molecules are @sm-1$ and @ / sm-1$, and diseases are @di-1$ and @ / di-1$).
[0084] Then, the annotated small dataset is divided into a training set and a validation set in a ratio of 8:2, and a non-generative pre-trained language model such as PubMedBERT or BioLinkBERT is used for fine-tuning and evaluation. The last layer of the relation extraction model is a three-classification model. After a SoftMax function, the probability that the paragraph belongs to the three relations is obtained. This probability can be directly used as the confidence score of the non-Unrelated triple. Assuming the triple Total Source of the article, , and Respectively represent Articles The publication year of the source article, the journal impact factor, and the confidence score given by the relationship extraction model. The comprehensive confidence score of the triple is defined as ,in , and Respectively based on , and The linear segment weighting coefficients are monotonically increasing, they are all less than 1 and the turning points of the linear segments need to be manually defined empirically.
[0085] (3) Construction of intestinal microbial knowledge graph (GMKG-200)
[0086] A multimodal uncertain intestinal microbial knowledge graph was constructed based on the three parts of knowledge: bacteria-bacteria, bacteria-human, and human-human, and named "GMKG-200".
[0087] The knowledge of bacteria-bacteria is provided by the Zhuowei knowledge base, which is structured and extracted from 8 authoritative databases and has passed the manual review of researchers in the early stage, so the confidence scores are directly defined as 1.0.
[0088] The person-to-person knowledge base, namely the "clinical medicine database", is provided by the PMapp database previously constructed by the inventor, which integrates 61 authoritative databases of molecular biology and clinical medicine. The confidence score is also defined as 1.0, which can provide complete knowledge support for the intestinal microbial knowledge graph (GMKG-200).
[0089] The knowledge of bacteria and humans is automatically extracted from PubTator3, and there is a certain extraction error. When extracting relationships, a comprehensive confidence score is calculated by combining all source information, which can measure this uncertainty well. This section contains 3 types of relationships that have been defined during relationship annotation: (gut microorganisms, Symbiotic or Exclusion, gut microorganisms), (gut microorganisms, Secrete or Consume, small molecules) and (gut microorganisms, Activate or Inhibit, diseases). PubTator3 connects the Zhuowei knowledge base and the PMapp knowledge base. It only needs to align the small molecules, drugs and disease entities to finally construct the gut microbial knowledge graph (GMKG-200).
[0090] (4) Knowledge graph multimodal uncertain reasoning system
[0091] A knowledge graph multimodal uncertainty reasoning system is constructed based on the triples of the gut microbial knowledge graph (GMKG-200) to predict potential associated diseases, drugs, genes, etc. for gut microorganisms.
[0092] First, the head entity, tail entity, and relation each have a randomly initialized embedding vector , and In addition, the head entity and the tail entity each have an embedding vector based on the multimodal annotations and ,The multimodal feature vector extraction methods for different entity types are different. Intestinal microorganism entities are vector normalized based on gene expression matrices, gene entities are processed by multi-category vector average pooling based on gene ontology categories, disease entities are encoded based on text descriptions using PubMedBERT or BioLinkBERT, protein entities are encoded based on amino acid sequences using ProteinBERT, and small molecule and drug entities are encoded based on SMILES using ChemBERTa.
[0093] Secondly, the dynamic gating network of the hybrid expert system is used to fuse the two embedding vectors of each entity to obtain its final vector representation, with the head entity vector For example: ,in , is a relation-based weight vector, which is also randomly initialized. The calculation of is the same. For an entity that does not exist in GMKG-200, just replace or Set to zero vector, or is equal to zero, or Will be completely based on or Therefore, as long as multimodal annotations are extracted for any entity, this model can achieve cold-start reasoning.
[0094] Finally, the final vector of the triple is represented as Substitute the basic knowledge graph embedding model and conversion function Generate confidence scores for predictions , and define the loss function by comparing it with the true confidence score. For example, TransE represents the distance after translation in the plane space: ,RotatE represents the distance after rotation in complex space: ,in represents the product of vector elements, and concat represents the concatenation of the real and imaginary parts of a complex number.
[0095] (5) Image database, which can be used for query, provide visual query results, and realize open source sharing of data.
[0096] Furthermore, the Neo4j graph database can be used to implement the above functions.
[0097] In the Gut Microbiome Knowledge Graph (GMKG-200), gut microbes are directly connected to diseases and drugs, and predicting new potential associations is equivalent to performing a simple knowledge graph completion task. However, gut microbes are not directly related to human genes, and can only be indirectly associated through multiple steps of diseases, drugs, and small molecules, which cannot be achieved with traditional knowledge graph link prediction. By achieving homogeneity of existing and new inference relationships based on confidence scores, it is only necessary to comprehensively calculate the probability expectation to effectively achieve confidence score estimation and correlation ranking of long-chain relationships.
Claims
1. A knowledge graph system for intestinal microorganisms, characterized in that: Including intestinal microbial knowledge base, intestinal microbial small molecule drug treatment related knowledge base, intestinal microbial knowledge graph composed of clinical medicine database construction, and knowledge graph multimodal uncertain reasoning system using intestinal microbial knowledge graph; The construction of the intestinal microbial small molecule drug treatment-related knowledge base includes searching relevant databases, indexing literature, obtaining a small database after indexing, dividing it into a training set and a validation set, using a non-generative pre-trained language model for training and evaluation, obtaining a three-classification model and obtaining the probabilities of the three relationships through functions, and obtaining a comprehensive confidence score through calculation; assuming that the triple Total Source of the article, y i 、j i and i They represent the publication year of the 𝑖th source article, the journal impact factor, and the confidence score given by the relationship extraction model. The last layer of the relationship extraction model is a three-classification model. After a SoftMax function, the probability that the paragraph belongs to the three relationships is obtained. The probability is directly used as the confidence score of the triple. The comprehensive confidence score of the triple is , where k n , k yi and k ji Respectively based on ,y i and j i A linear piecewise weighting coefficient of , wherein the weighting coefficient is monotonically increasing; The knowledge graph multimodal uncertain reasoning system is constructed through the following steps: First, the head entity, tail entity, and relation each have a randomly initialized embedding vector , and ; In addition, the head entity and the tail entity each have an embedding vector based on the multimodal annotation and ,The multimodal feature vector extraction methods for different entity types are different; Secondly, the dynamic gating network of the hybrid expert system is used to fuse the two embedding vectors of each entity to obtain its final vector representation, with the head entity vector For example: ,in , is a relation-based weight vector, which is also randomly initialized. The calculation of the tail entity vector is similar. For an entity that does not exist in the intestinal microbial knowledge graph, just replace or Set to zero vector, or is equal to zero, or will be completely based on or generate; Finally, the final vector of the triple is represented as Substitute the basic knowledge graph embedding model and conversion function Generate confidence scores for predictions , and compare it with the true confidence score to define the loss function.
2. According to the intestinal microbial knowledge graph system of claim 1, the construction of the intestinal microbial knowledge base includes obtaining information from a public database, performing data cleaning and de-redundancy processing, and retaining intestinal microbial related knowledge; extracting entity tables and relationship tables, further including extracting attribute tables; and uniformly matching entities in relationship tables and attribute tables with entity tables to construct a complete "intestinal microbial knowledge base".
3. According to the intestinal microbial knowledge graph system of claim 2, the entire database is extracted / or a part of the database is selected to extract the entity table and the relationship table, the entire database is / or a part of the database is selected to be merged and de-redundant to extract the attribute table therefrom, and the association between the intestinal microbial entity and the small molecule entity and the disease entity is extracted according to the entities annotated by the annotation tool platform; according to the position, interval, and frequency of occurrence of the entity in the article, it is judged whether it is a potential entity to extract the relationship.
4. The intestinal microbial knowledge graph system according to claim 1, wherein the training set and the validation set are divided according to a ratio of 6:2, 7:1, 8:1, 9:2, 9:1, or 8:2; and the confidence scores of the intestinal microbial knowledge base and the clinical medicine database are set to 1.
0.
5. According to the intestinal microbial knowledge graph system of claim 1, intestinal microbial entities are vector normalized based on gene expression matrices, gene entities are multi-category vector average pooled based on gene ontology categories, disease entities are encoded based on text descriptions using PubMedBERT or BioLinkBERT, protein entities are encoded based on amino acid sequences using ProteinBERT, and small molecules and drug entities are encoded based on SMILES using ChemBERTa.
6. The intestinal microbial knowledge graph system according to claim 1, further, using TransE to represent the distance after translation in the plane space: ,RotatE represents the distance after rotation in complex space: ,in represents the product of vector elements, and concat represents the concatenation of the real and imaginary parts of a complex number.
7. A construction device for loading the intestinal microbial knowledge graph system according to claims 1-6, comprising: The data acquisition module is used to obtain the original data related to intestinal microorganisms, specifically, to obtain the original database information of "intestinal microorganism knowledge base", "intestinal microorganism small molecule drug treatment related knowledge base" and "clinical medicine database"; Data processing module, including cleaning and removing redundancy from data, retaining target-related knowledge, extracting entity tables, relationship tables, and attribute tables; or using annotation tool platforms, AI technology, and manual indexing to index the acquired data; The data training module is used to divide the labeled data into training sets and validation sets, and use a non-generative pre-trained language model for training and evaluation to obtain a three-classification model and obtain the probabilities of the three relationships through functions, and obtain confidence scores through calculations; The data matching module is used to uniformly match the entities in the relationship table and the attribute table with the entity table; and obtain matching data after data processing through confidence scores; The data reasoning module is used to import the acquired data into the knowledge graph multimodal uncertain reasoning system to predict potential associated diseases, drugs, and genes for intestinal microorganisms.
8. A computer storage medium for loading the intestinal microbial knowledge graph system described in claims 1-6.
Citation Information
Patent Citations
Multi-modal fusion video classification method and system based on brain-like feedback interaction
CN115588148A
Construction method and device of enteric microorganism knowledge graph
CN117292846A