Microorganism and brain disease relation extraction method based on deep learning
Through a deep learning-based method, the relationship between microorganisms and brain diseases is extracted from biomedical unstructured data, solving the problem of low efficiency in relationship extraction in the existing technology, and achieving more efficient and accurate relationship discovery.
Patent Information
- Application Number
- CN202510313012.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-17
- Publication Date
- 2025-06-17
AI Technical Summary
The prior art is difficult to efficiently extract the potential relationship between microorganisms and brain diseases from a large number of unstructured biomedical data, resulting in many undiscovered relationships being missed.
A deep learning-based method is used to retrieve and screen literature data through keywords, preprocess and label, build a training data set, and extract the microbial-brain disease relationship using the Transformer model.
It improves the efficiency of the extraction of the relationship between microorganisms and brain diseases, increases the probability of discovering potential relationships, improves the generalization ability and robustness of the model, and improves the accuracy of relationship extraction.
Smart Images

Figure CN120164633A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of extracting microorganism-disease relationships, and specifically to deep learning methods. Background Art
[0002] Microorganisms: refer to tiny organisms that cannot be observed with the naked eye, including fungi, viruses, protozoa, etc. They exist widely in nature and play important roles. Among them, gut microbiota maintain relatively complex relationships with the host and are involved in aspects such as digestion, absorption, immune regulation, and metabolism of the human body's functions. Gut microbiota mainly include Bacteroides, Bifidobacterium, Lactobacillus, Enterobacteriaceae, etc. Maintaining the balance of these microbial communities helps reduce the risk of disease occurrence, and its mechanism of action is more conducive to the development of new treatment methods.
[0003] Brain diseases: refer to functional or structural damages caused by abnormal brain work, leading to neurological diseases (such as Huntington's disease, dementia, etc.), cerebrovascular diseases (such as cerebral ischemia, thrombosis, etc.), and mental diseases (such as schizophrenia, depression, etc.). The occurrence of these brain diseases may be related to various factors such as environment and immunity, and their manifestations are in aspects such as cognition, emotion, and behavior. Therefore, regulating the microbial flora through methods such as supplementing probiotics and dietary intervention to improve the function of the nervous system, and then directly or indirectly affecting the gut-brain axis mechanism between organisms and brain diseases, can effectively avoid the occurrence of potential diseases.
[0004] Deep learning: is the most important branch of machine learning. It learns and extracts features of multi-layer neural networks by mimicking the human brain, so as to complete prediction and classification tasks. It includes key technologies such as convolutional neural networks (CNNs), recurrent neural networks (RNNs), and Transformers, and is good at processing large amounts of complex unstructured data. Therefore, it is necessary to apply deep learning technology to the relationship between microorganisms and brain diseases, such as automatic feature extraction, multi-modal data fusion, and pattern recognition improvement, which will improve the potential relationship between microorganisms and brain diseases.
[0005] The microbiota and the host have a symbiotic relationship existing in nature. It not only has a profound impact on the host's immune system and metabolic regulation, but also directly or indirectly affects the host's brain diseases. How to extract the relationship between them from a large amount of biomedical unstructured data will improve the accuracy of biomedical information and provide important support for downstream research. However, traditional extraction methods mainly rely on artificial rules and will miss many undiscovered relationships. Therefore, using deep learning methods will increase the efficiency of extracting relationships. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for extracting the relationship between microorganisms and brain diseases based on deep learning, which is convenient for researchers to conduct the latest research and provide new insights, and also provides a favorable reference for downstream tasks to predict the relationship between microorganisms and brain diseases.
[0007] To achieve the above object, the present invention provides a method for extracting the relationship between microorganisms and brain diseases based on deep learning, including:
[0008] S1: Retrieving, screening, and sorting from major databases based on keywords to obtain the data of the collected literature.
[0009] S2: Preprocessing the obtained target literature, including operations such as organizing into structured data, denoising, annotating, and partitioning, to obtain relevant data sets.
[0010] S3: Inputting the training data set in the data set into the relationship model for relationship extraction, and evaluating the performance of the model. Finally, the model outputs the microorganism-brain disease relationship data.
[0011] Among them, the data of the collected literature includes:
[0012] S11: Retrieving relevant literature on microorganisms and brain diseases in the bioinformatics and biomedical literature databases according to keywords.
[0013] S12: Screening the relevant literature according to the requirements of the target data to obtain the preliminary target literature data.
[0014] Among them, preprocessing the target literature to obtain relevant data sets includes:
[0015] S21: Organizing the preliminary target literature sorted from the databases in different biomedical fields into structured data.
[0016] S22: There are various problems such as illegal characters in the structured data, which need to be preprocessed and reduce the adverse effects of text noise on the experiment.
[0017] S23: Marking the preprocessed data to ensure the entity positioning of microorganisms-diseases.
[0018] S24: The marked data does not meet the conditions for inputting into the model and needs to be divided according to rules to obtain three corresponding data sets, including the training set, the validation set, and the test set.
[0019] Among them, performing relationship extraction on the training data set to obtain the final relationship data includes:
[0020] S31: What the model needs is text feature data, and tools must be used to tokenize the three text data sets.
[0021] S32: Use the tokenized test set to test the deep learning model, enabling the deep learning model to extract microorganism-disease relationships;
[0022] S33: Evaluate the deep learning model using three evaluation metrics and obtain the final relationship data.
[0023] A method for extracting the relationship between microorganisms and brain diseases based on deep learning according to the present invention includes: collecting literature data; preprocessing the target literature to obtain a relevant data set; training the data set for relationship extraction to obtain the final relationship data. The beneficial effects of the present invention are: by extracting the text of the unstructured data existing in biomedicine, the potential relationship between microorganisms and brain diseases can be effectively observed; by combining the characterized words and parts of speech of different text data and inputting them into the deep learning model, the optimal relationship between different brain disease types and different microorganism characteristics can be obtained, increasing the probability of their potential relationship; at the same time, the generalization ability and robustness of the model can be improved, and importantly, the accuracy of model named entity recognition and relationship extraction can also be improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0025] Figure 1 is the flowchart of collecting the literature data.
[0026] Figure 2 is the flowchart of preprocessing the target literature to obtain a relevant data set.
[0027] Figure 3 is the flowchart of training the data set for relationship extraction to obtain the final relationship data. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0028] In order to enable researchers in the field to more clearly understand the technical solutions of the present invention and be able to implement them, the present invention will be described in detail below in combination with specific embodiments and the drawings, where the same symbols or labels represent the same content. The following described technical solutions and their specific embodiments are all exemplary, only used to represent the technical solutions of the present invention, and the embodiments all belong to the scope of the present invention application, and should not be construed as a limitation to the present invention.
[0029] Before introducing the technical solutions of the present invention, the noun terms involved in the embodiments will be explained first.
[0030] Pubmed: It mainly collects articles in the fields of life science and biomedicine. Its content is sourced from MedLine and provides various convenient retrieval methods, such as keywords, journals, authors, topics, time, MeSH terms, etc. It can quickly help researchers find relevant literature and provides open resources including author information, literature abstracts, publication dates, and links to full-text literature.
[0031] Embase: It is a global comprehensive biomedical literature database covering multiple fields such as medicine, pharmacy, clinical, and life science. Compared with Pubmed, Embase focuses more on European literature, provides more complete clinical trial and drug literature, and uses a more precise MTREE thesaurus for retrieval functions.
[0032] NLTK (Natural Language Toolkit): It is a Python tool for analyzing text corpora, including functions such as text classification, part-of-speech tagging, and stemming. It is an ideal tool for text processing.
[0033] Gensim: It is a Python library for text modeling and similarity analysis based on unsupervised learning. It is good at processing large-scale text data and mainly includes text processing algorithms such as Word2Vec and Doc2Vec.
[0034] Residual Network (ResNet): It is a deep architecture of neural networks. Its core idea is to introduce a "residual learning" mechanism to solve the degradation problem caused by the increase in the number of network layers in neural networks.
[0035] Transformer: It is a learning architecture of deep neural networks, mainly dealing with sequential data, such as text generation, machine translation, etc. It consists of an encoder and a decoder and is built based on feed-forward neural networks and self-attention mechanisms. Its main advantage is that it can process all elements in parallel and capture the dependencies of distant elements simultaneously, which can improve computing and modeling capabilities.
[0036] Please refer to Figures 1 to 3 , the present invention is a method for extracting the relationship between microorganisms and brain diseases based on deep learning, including the following steps of embodiments:
[0037] S1: Collect literature data;
[0038] S11: According to keywords, retrieve relevant literature on microorganisms and brain diseases in biological information and biomedical literature databases;
[0039] Specifically, search in public biomedical literature databases such as Pubmed and Embase using keywords such as "Disease", "Microorganism", "Microbe", "Brain Disease", etc., so as to obtain a large number of text documents on gut microbiota and brain diseases in the database. The literature data comes from public channels such as hospitals, open journals, health organizations, and research institutions, and is retrieved based on biomedical data such as clinical data, microbiomics, and disease type data, so as to obtain gut microbiota and disease data.
[0040] S12: According to the requirements of the target data, screen the relevant literature to obtain preliminary target literature data;
[0041] Specifically, according to the keyword search, sort and screen the literature data in terms of relevance and currency from aspects such as journals, publication dates, article types, and availability, including literature such as clinical data, microbiomics, and disease type data.
[0042] Data collection refers to integrating multiple collected research data on brain diseases and gut microbiota to provide a comprehensive data set for the experimental model. Repeatedly screen the collected literature to ensure the high quality and accuracy of the data.
[0043] S2: Preprocess the target literature to obtain relevant data sets;
[0044] S21: Organize the preliminary target literature data into structured data;
[0045] Specifically, store the irregular unstructured data extracted from different literature databases using tables, and classify and summarize different types of data, which helps to ensure the consistency, availability, etc. of the data;
[0046] S22: Preprocess the structured data to reduce the adverse effects of noise on the experiment;
[0047] Specifically, clean the repeated, missing, special symbol, and incorrect structured data to ensure the correctness of the data.
[0048] S23: Mark the preprocessed data to obtain marked data;
[0049] Specifically, manually mark different types of the preprocessed data to avoid data non-uniformity in subsequent relation extraction.
[0050] S24: Divide the marked data according to rules to obtain three corresponding data sets;
[0051] Specifically, the research data related to the labeled brain diseases and gut microbiota are divided into a training set, a validation set, and a test set in proportion for training, evaluation, tuning, etc. of the extraction model.
[0052] After preprocessing, labeling, and partitioning the target literature, it is necessary to ensure that each item of the data meets the standard requirements through manual inspection or verification tools, and make optimizations and adjustments in a timely manner.
[0053] S3 Train this dataset for relation extraction to obtain the final relation data;
[0054] S31: Tokenize the text data using a tool;
[0055] Specifically, use NLTK to convert the text sentences in the partitioned dataset into text data, and use Gensim to encode the text data to obtain the representations of the microbiota and diseases e i , which is the feature representation of the i-th word in the text data, and sequentially construct the text vector of the sentence as S = (e1, e2,..., e n ), which includes the sentence structure and text semantic information.
[0056] S32: Train the dataset using the deep learning model method;
[0057] Specifically, use the Transformer model to identify the microbiota and disease entities. Make the constructed sentence text vectors form a sentence bag M = (S1, S2,..., S N ), and use the word vector x and the positional encoding in M as the model input z together;
[0058] Specifically, perform a dot product operation on the z described in S421 and calculate the score value;
[0059] Specifically, calculate the attention weight α through the normalization function SoftMax for the score value described in S422 ij , and finally perform a weighted sum of the attention weight and the value vector;
[0060] Specifically, the output of each layer combines the feed-forward neural network and the residual connection for gradient propagation, so as to stabilize the training and accelerate the convergence to obtain the output z' of the encoder;
[0061] Specifically, perform a residual connection on the output z' and the original feature vector z, and jointly input them into the BERE model to extract entity relations, including the following implementation steps:
[0062] Step 1: For a certain output sentence feature, use the Bi-GRU layer to perform deep feature extraction on the recognized text vector. The purpose is to construct nested phrases from adjacent words and capture the relevant relationships between adjacent words. This method follows the top-down principle.
[0063] Step 2: Encode the sentence structure. Adopt the potential learning method Gumbel Tree-GRU to capture the sentence structure information, and use a scoring mechanism to select the optimal solution among all feasible phrase features.
[0064] Step 3: After encoding the sentence structure for all output sentences, adopt a sentence aggregation strategy to aggregate all sentences in the sentence package.
[0065] Specifically, similar to entity recognition, use Transformer to calculate the relationship category Y of entity time, and use {0, 1} to represent the relationship category.
[0066] S33: Use the test set to test the deep learning model;
[0067] Specifically, after completing the training of the model in S42, use the test set to test the model to check the performance of the model on the unseen dataset and ensure the generalization ability of the model.
[0068] S34: Use the evaluation metrics to evaluate the deep learning model and obtain the final relationship data.
[0069] Specifically, after testing the model, using the evaluation metrics to evaluate the model is beneficial to measuring the model performance, detecting problems such as overfitting, and improving the reliability and effectiveness of the model.
Claims
1. A method for extracting the relationship between microorganisms and brain diseases based on deep learning, characterized in that: include: S1: The data of collected literature were obtained by searching, screening and sorting from various databases based on keywords; S2: Preprocess the target documents, including arranging them into structured data, denoising, labeling, and partitioning, to obtain relevant data sets; S3: The training data set in the dataset is input into the relational model for relation extraction, and the model performance is evaluated. Finally, the model outputs the microbiome-brain disease relationship data.
2. A method for extracting the relationship between microorganisms and brain diseases based on deep learning as described in claim 1, characterized in that: The collected literature data includes: S11: Search for relevant literature on microorganisms and brain diseases in bioinformatics and biomedical literature databases based on keywords; S12: According to the target data requirements, screen relevant literature to obtain preliminary target literature data.
3. A method for extracting the relationship between microorganisms and brain diseases based on deep learning as described in claim 2, characterized in that: The preprocessing of the target document obtains a relevant data set, including: S21: Organize the preliminary target literature from databases in different biomedical fields into structured data; S22: Structured data has various problems such as illegal characters, which require preprocessing and reducing the adverse effects of text noise on the experiment; S23: labeling the preprocessed data to ensure the location of the microorganism-disease entity; S24: The labeled data does not meet the conditions for inputting the model and needs to be divided according to the rules to obtain three corresponding data sets, including training set, validation set and test set.
4. A method for extracting the relationship between microorganisms and brain diseases based on deep learning as described in claim 3, characterized in that: The training data set is subjected to relation extraction to obtain final relation data, including: S31: The model requires text feature data, and tools must be used to lemmatize the three text datasets; S32: Testing the deep learning model using the lexicalized test set, so that the deep learning model can extract the microorganism-disease relationship; S33: Use three evaluation indicators to evaluate the deep learning model and obtain final relationship data.