Method for extracting relationship between diseases and intestinal microorganisms based on transfer learning model

Through the combination of transfer learning model and graph attention network, the technical gap in microbial disease relationship extraction is solved, efficient extraction of the relationship between disease and intestinal microbials is achieved, and the development of personalized medical and treatment strategies is promoted.

CN120337922APending Publication Date: 2025-07-18GUILIN UNIVERSITY OF TECHNOLOGY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410043025.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-01-11
Publication Date
2025-07-18

AI Technical Summary

Technical Problem

There is currently a lack of effective technical means for extracting microbial disease relations, which limits the study of complex relationships between microbial and human diseases and affects the prevention, diagnosis and treatment of diseases.

Method used

Using a transfer learning model-based method, the relevant literature data is preprocessed, entity annotated and relational annotated, combined with NLTK and Gensims tools for word segmentation and embedding representation, data augmentation and text vectorization are performed, training is performed using the BERE model, and graph attention network (GAT) is introduced for relationship extraction.

Benefits of technology

It improves the accuracy and generalization ability of the relationship extraction between disease and intestinal microbials, provides a deeper understanding of disease mechanisms, and provides a powerful tool for personalized medical and therapeutic strategies.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120337922A_ABST
    Figure CN120337922A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, in particular to a method for extracting the relationship between diseases and intestinal microorganisms based on a transfer learning model, which comprises the following steps: preprocessing collected related literature data to obtain a training set, a verification set and a test set; training a pre-training model by using the training set, the verification set and the test set to obtain a relationship extraction model; and performing relation extraction through the relation extraction model to obtain a final relation category. According to the method, the BERE framework of transfer learning is used for performing relation extraction on the diseases and the intestinal microorganisms, so that the correlation between the diseases and the intestinal microorganisms is researched from a comprehensive perspective, and the potential correlation is mined. The application of the method not only is expected to deepen the understanding of disease mechanisms, but also provides a powerful tool for future medical research, and promotes the development of personalized medical treatment and treatment strategies. The problem that no related technology for extracting the relation of the microbial diseases exists at present is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of data processing, and in particular to a method for extracting the relationship between diseases and gut microbiota based on a transfer learning model. Background Art

[0002] Relationship between gut microbiota and human diseases: The gut microbiota of humans constitutes a vast microbial community inhabiting the gut, containing at least 1,000 different microorganisms with a total number of about 10^14, and is considered another organ of the human body. Under normal circumstances, the gut microbiota maintains a dynamic balance of mutualism with the host, affected by factors such as human development, physical condition, and diet. Imbalance of this microbial community may lead to disorders of the gut flora, which is further associated with the occurrence of various diseases, including gastrointestinal diseases, liver cirrhosis, and even mental diseases such as depression, autism, and bipolar disorder. Therefore, in-depth study of the mechanism of action and regulatory relationships of gut microbiota is expected to provide new perspectives and strategies for disease prevention and treatment.

[0003] However, there is currently no relevant technology for extracting the relationship between microorganisms and diseases, which thus limits the study of the complex relationship between microorganisms and human diseases and reduces the profound insights provided for disease prevention, diagnosis, and treatment. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for extracting the relationship between diseases and gut microbiota based on a transfer learning model, aiming to solve the problem that there is currently no relevant technology for extracting the relationship between microorganisms and diseases.

[0005] To achieve the above purpose, the present invention provides a method for extracting the relationship between diseases and gut microbiota based on a transfer learning model, including the following steps:

[0006] Preprocess the collected relevant literature data to obtain a training set, a validation set, and a test set;

[0007] Use the training set, the validation set, and the test set to train a pre-trained model to obtain a relationship extraction model;

[0008] Perform relationship extraction through the relationship extraction model to obtain the final relationship category.

[0009] Among them, the preprocessing of the collected relevant literature data to obtain a training set, a validation set, and a test set includes:

[0010] Screen and organize the data set sorted out by the BGMDB database to obtain relevant literature data;

[0011] Perform text cleaning on the relevant literature data to obtain cleaned data;

[0012] Perform entity annotation and relationship annotation on the cleaned data to obtain annotated data;

[0013] Remove the word segmentation and stop words in the annotated data to obtain the remaining data;

[0014] Perform entity embedding representation on the remaining data to obtain the embedded data;

[0015] Perform data augmentation on the embedded data to obtain augmented data;

[0016] Perform text vectorization on the augmented data to obtain data vectors;

[0017] Partition the data vectors to obtain a training set, a validation set, and a test set.

[0018] Among them, the step of removing the word segmentation and stop words in the annotated data to obtain the remaining data includes:

[0019] Use NLTK to perform word segmentation on the annotated data and remove the stop words in the annotated data at the same time to obtain the remaining data.

[0020] Among them, the step of performing entity embedding representation on the remaining data to obtain the embedded data includes:

[0021] Use Gensims to train a Word2Vec model to obtain the embeddings of disease and microorganism entity pairs to obtain the embedding representation;

[0022] Embed the embedding representation into the entities in the remaining data to obtain the embedded data.

[0023] Among them, the data augmentation method includes synonym replacement, entity swapping, sentence recombination, and generative adversarial networks.

[0024] Among them, the step of performing text vectorization on the augmented data to obtain data vectors includes:

[0025] Combine the outputs of NLTK and Gensims to convert the entire text into a vector to obtain data vectors.

[0026] Among them, the training set, the validation set, and the test set are in the ratio of 8:1:1.

[0027] Among them, the pre-trained model is the BERE model.

[0028] Among them, after obtaining the sentence packets mentioning entities through the relationship extraction model, word embedding, part-of-speech embedding, GAT feature learning, and aggregating sentences are performed in sequence to obtain the final relationship category.

[0029] A method for extracting the relationship between diseases and gut microbiota based on a transfer learning model includes cleaning the relevant literature data to obtain cleaned data; performing entity annotation and relationship annotation on the cleaned data to obtain annotated data; using NLTK to tokenize the annotated data and removing stop words in the annotated data to obtain remaining data; training a Word2Vec model using Gensims to obtain embeddings of disease and microbiota entity pairs, resulting in an embedding representation; embedding the embedding representation into entities in the remaining data to obtain post-embedded data; performing data augmentation on the post-embedded data to obtain augmented data, and the data augmentation methods include synonym replacement, entity swapping, sentence restructuring, and generative adversarial networks; combining the outputs of NLTK and Gensims to transform the entire text into a vector to obtain a data vector; dividing the data vector to obtain a training set, a validation set, and a test set; using the training set, the validation set, and the test set to train a pre-trained model to obtain a relationship extraction model; obtaining a sentence package mentioning entities through the relationship extraction model and then performing word embedding, part-of-speech embedding, GAT feature learning, and aggregating sentences in sequence to obtain the final relationship category. The present invention transfers the BERE model, uses NLTK and Gensims for entity recognition and embedding, and introduces GAT to strengthen the modeling of the relationship between entities in the BERE framework. Finally, experiments are conducted on simulated datasets and real datasets, and the method proposed by us has achieved good performance. The application of the present invention not only is expected to deepen the understanding of disease mechanisms but also provides a powerful tool for future medical research, promoting the development of personalized medicine and treatment strategies. BRIEF DESCRIPTION OF THE DRAWINGS

[0030] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the following-described drawings are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0031] Figure 1 is a data preprocessing flowchart.

[0032] Figure 2 is a flowchart for mining the association information between diseases and gut microbiota using a transferred BERE model.

[0033] Figure 3 is a schematic diagram of the overall process of mining the association information between diseases and gut microbiota from data.

[0034] Figure 4 is a flowchart of a method for extracting the relationship between diseases and gut microbiota based on a transfer learning model provided by the present invention. Detailed implementation manners

[0035] The embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are intended to explain the present invention, and should not be construed as a limitation to the present invention.

[0036] Please refer to Figures 1 to 4 , the present invention provides a method for extracting the relationship between diseases and gut microbiota based on a transfer learning model, including the following steps:

[0037] S1 Preprocess the collected relevant literature data to obtain a training set, a validation set, and a test set;

[0038] The Brain-Gut Microbiota Association Database (BGMDB) was manually collected and organized by us from PubMed. It contains the association relationships between brain diseases and gut microbiota. As a dataset, the basic process of its basic data preprocessing is performed according to Figure 1 shown.

[0039] The specific method is as follows:

[0040] S11 Screen and organize the dataset sorted by the BGMDB database to obtain relevant literature data;

[0041] Specifically, the dataset sorted by the BGMDB database is preliminarily screened and organized to obtain disease, gut microbiota entities, and relevant literature data.

[0042] S12 Clean the relevant literature data to obtain cleaned data;

[0043] Specifically, special characters, HTML tags, and non-English characters in the literature data are removed. We implemented text cleaning to eliminate noise and ensure the purity of the text data.

[0044] S13 Perform entity annotation and relationship annotation on the cleaned data to obtain annotated data;

[0045] Specifically, in this stage, we performed entity annotation and relationship annotation. By manual annotation, named entity recognition tools, or other automatic methods, the positions of diseases and microorganisms mentioned in each literature were determined, and the relationship types between them were annotated.

[0046] S14 Remove word segmentation and stop words from the annotated data to obtain remaining data;

[0047] Specifically, use NLTK to tokenize the labeled data, convert the text into an input format understandable by the model, and at the same time remove the stop words in the labeled data to reduce the feature space and make the model more focused on learning key information. Obtain the remaining data.

[0048] S15 Perform entity embedding representation on the remaining data to obtain the embedded data;

[0049] Specifically, use Gensims to train a Word2Vec model to obtain the embeddings (e1, e2) of disease and microorganism entity pairs, as well as related literature data. e i Represents the encoded feature vector of the i-th word in the input sentence to obtain the embedding representation; perform entity embedding on the remaining data with the embedding representation to obtain the embedded data.

[0050] S16 Perform data augmentation on the embedded data to obtain augmented data;

[0051] Specifically, the data augmentation methods include synonym replacement, entity swapping, sentence restructuring, and adversarial generation networks. It improves the generalization ability and robustness of the dataset and helps the model better adapt to different text samples.

[0052] S17 Perform text vectorization on the augmented data to obtain data vectors;

[0053] Specifically, combine the outputs of NLTK and Gensims to convert the entire text into a vector to obtain the data vector S. This vectorization method can capture the semantic and structural information in the text and provide meaningful input features for the model. S is a sentence composed of n given words, which can be represented as a sequence of vectors: S = (e1, e2, …, e n )

[0054] S18 Divide the data vectors to obtain a training set, a validation set, and a test set.

[0055] Specifically, the training set, the validation set, and the test set are in a ratio of 8:1:1. Ensure sufficient training samples, effective model tuning, and reliable performance evaluation. This step provides an orderly and manageable data basis for subsequent model training, parameter adjustment, and performance evaluation. Through these comprehensive steps, we have constructed a high-quality, diverse disease and gut microbiota relationship dataset suitable for the relation extraction task.

[0056] S2 Use the training set, the validation set, and the test set to train a pre-trained model to obtain a relation extraction model;

[0057] Specifically, the pre-trained model is the BERE model. BERE is a deep learning framework for automatically extracting drug-related relationships from literature. The model uses latent tree learning and self-attention techniques to capture the syntactic information of sentences. Through careful evaluation, the BERE model is selected for transfer training in the present invention. In the pre-trained model for the D-T (Drug-Target) task, parameters are obtained. The parameters obtained from the training are applied to the D-M (Disease-Microbiome) task. This helps to reduce computational resources and time. Secondly, it makes the model suitable for new tasks, namely, the extraction of the relationship between diseases and gut microbiota.

[0058] S3 performs relationship extraction through the relationship extraction model to obtain the final relationship category.

[0059] Specifically, after obtaining the sentence packet mentioning the entities through the relationship extraction model, word embedding, part-of-speech embedding, GAT feature learning, and aggregating sentences are performed in sequence to obtain the final relationship category.

[0060] Apply the pre-trained BERE model to the task of extracting the relationship between diseases and gut microbiota.

[0061] Obtain the sentence packet mentioning the entities. The BERE model collects all the sentences mentioning these two entities from the document to form a sentence packet P = {S1, S2, …, S N}, where N is the number of sentences.

[0062] Word embedding and part-of-speech embedding. For each sentence S i , the BERE model uses word embedding and part-of-speech embedding to represent each word w ij , where j is the index of the word, and then uses the self-attention mechanism to capture long-term dependencies and adds the original word vector back through a residual connection. Finally, x ij is used as the input vector of the word w ij , which is concatenated by the word embedding w e and the part-of-speech embedding w p ; z ij is used as the hidden layer vector of the word w ij , which is calculated by a fully connected layer and an activation function; a ij is used as the self-attention coefficient of the word w ij , which is calculated by a fully connected layer and a softmax function; y ij is used as the output vector of the word w ij , which is obtained by the weighted sum of the self-attention coefficient and the input vector; h ij is used as the final vector of the word w ij , which is obtained by the residual connection between the output vector and the input vector.

[0063] GAT feature learning. Embed the GAT structure into the BERE model to construct a model that combines the advantages of both.

[0064] Word w ij The correlation coefficient between the entity e1 and e2 can be expressed as:

[0065]

[0066] Word w ij The attention coefficient of can be expressed as:

[0067] β ij = Leaky ReLU(e ij )

[0068] w ij The normalized attention coefficient of can be expressed as:

[0069] y ij = softmax(β ij )

[0070] w ij The output vector of can be expressed as:

[0071]

[0072] w ij The final vector of is then expressed as

[0073] f ij = g ij + wh ij .

[0074] Aggregate sentences. Use the Gumbel Tree-GRU in the BERE model to learn a latent syntactic tree, thereby implicitly understanding the sentence structure and propagating information bottom-up to obtain the final vector m of each word w in each sentence S i in, and then aggregate these vectors into a sentence vector u ij , and use the Gumbel Tree-GRU to update it. ij , and then aggregate these vectors into a sentence vector u i , and use the Gumbel Tree-GRU to update it.

[0075] Classification optimization. Finally, use a scoring mechanism to evaluate the importance of each sentence in relation prediction, and use a multi-instance learning framework to aggregate all sentences in the sentence bag, and calculate the final relation category by a fully connected layer and a softmax function

[0076] Define the objective function using cross-entropy:

[0077]

[0078] where y i is the true relationship category vector of the document, is the predicted relationship category vector of the document.

[0079] The present invention provides a method for extracting the relationship between diseases and gut microbiota based on a transfer learning model, including text cleaning of the relevant literature data to obtain cleaned data; entity annotation and relationship annotation of the cleaned data to obtain annotated data; using NLTK to tokenize the annotated data and removing stop words in the annotated data to obtain remaining data; using Gensims to train a Word2Vec model to obtain embeddings of disease and microbiota entity pairs to obtain embedding representations; embedding the embedding representations into entities in the remaining data to obtain post-embedded data; performing data augmentation on the post-embedded data to obtain augmented data, and the data augmentation methods include synonym replacement, entity swapping, sentence restructuring, and generative adversarial networks; combining the outputs of NLTK and Gensims to convert the entire text into a vector to obtain data vectors; partitioning the data vectors to obtain a training set, a validation set, and a test set; using the training set, the validation set, and the test set to train a pre-trained model to obtain a relationship extraction model; obtaining a sentence bag mentioning entities through the relationship extraction model and then performing word embedding, part-of-speech embedding, GAT feature learning, and aggregating sentences in sequence to obtain the final relationship category. The present invention transfers the BERE model, uses NLTK and Gensims for entity recognition and embedding, and introduces GAT to strengthen the modeling of the relationship between entities in the BERE framework. Finally, experiments are conducted on simulated datasets and real datasets, and the method proposed by us has achieved good performance. The application of the present invention not only is expected to deepen the understanding of disease mechanisms, but also provides a powerful tool for future medical research and promotes the development of personalized medicine and treatment strategies.

[0080] 1. Dataset

[0081] PubMed is a free biomedical literature database developed and maintained by the National Library of Medicine (NLM) of the National Institutes of Health (NIH) in the United States. It is part of the information retrieval Entrez system. PubMed aims to provide high-quality and authoritative biomedical literature resources for a wide range of biomedical researchers, doctors, students, and the public, covering all aspects from basic science to clinical medicine.

[0082] The Brain-Gut Microbiome Association Database (BGMDB) was manually collected and organized by us from PubMed. It contains the association between brain diseases and gut microbiota. Through this series of data preprocessing and standardization steps, the research data on the association between brain diseases and gut microbiota was organized into a consistent, standardized, and reliable dataset. This provides a reliable foundation for subsequent data integration, storage, and analysis, enabling researchers to more accurately understand and analyze the interactions and associations of the brain-gut axis, mine potential biomarkers and therapeutic targets, and promote scientific development and medical progress in related fields.

[0083] 2、Methods

[0084] 1) Data preprocessing

[0085] During the data preprocessing process for the task of extracting the relationship between diseases and gut microbiota, we took a series of key steps to ensure the quality and adaptability of the data. First, through text cleaning, we removed special characters, HTML tags, and non-English characters from the text to eliminate noise. Then, through entity annotation and relationship annotation, we determined the positions of diseases and microorganisms mentioned in each literature and labeled the types of relationships between them. This process can be completed through manual annotation, named entity recognition tools, or other automatic methods. Tokenization using NLTK was used to convert the text into an input format understandable by the model, and stop words were removed to reduce the feature space. At the same time, Gensims was used to train a Word2Vec model to obtain the embedded representation of entities. Combining the outputs of NLTK and Gensims, the entire text was transformed into a vector representation. This vectorization method can capture the semantic and structural information in the text and provide meaningful input features for the model.

[0086] To increase the diversity and quantity of the data, we performed data augmentation using a series of common methods, such as synonym replacement, entity swapping, sentence recombination, and generative adversarial networks. Through these technical means, we improved the generalization ability and robustness of the dataset, which helps the model better adapt to different text samples.

[0087] Finally, we performed data partitioning, dividing the entire dataset into a training set, a validation set, and a test set. The data was partitioned at a ratio of 8:1:1, ensuring sufficient training samples, effective model tuning, and reliable performance evaluation. This step provides an orderly and manageable data foundation for subsequent model training, parameter adjustment, and performance evaluation. Through these comprehensive steps, we constructed a high-quality, diverse, and disease-gut microbiota relationship dataset suitable for the relationship extraction task.

[0088] 2) Model transfer and optimization

[0089] In the present invention, we adopted an innovative method. We migrated the BERE model, which had performed excellently in the field of biomedical relation extraction, and introduced the Graph Attention Network (GAT) for model optimization. First, by selecting a BERE pre-trained model suitable for biomedical relations, we achieved transfer learning of the model on a new task (extracting the relationship between diseases and gut microbiota). During this process, we froze the underlying network to retain general features and fine-tuned the upper layer to adapt to the new task. This transfer learning strategy fully utilized the knowledge learned by the BERE model in previous tasks.

[0090] To further improve the performance of the model, we introduced the Graph Attention Network (GAT). We selected a GAT structure suitable for the characteristics of the task and embedded it into the BERE model, enabling the model to better capture the complex relationships between entities, especially in the relationship between diseases and gut microbiota. Through joint training and collaborative learning, we ensured effective information interaction between the BERE model and the GAT network, improving the performance of the overall model. Through this innovative method, we successfully migrated advanced biomedical relation extraction technology to the task of extracting the relationship between diseases and gut microbiota and optimized it by introducing the GAT network, providing a more powerful model for subsequent extraction of the relationship between diseases and gut microbiota.

[0091] Technical effects:

[0092] 1. The present invention uses tools such as NLTK and Gensims to provide rich natural language processing functions for entity recognition and embedding. By combining these tools, it is possible to more accurately identify disease and gut microbiota entities in the text, thereby improving the accuracy of relation extraction.

[0093] 2. The present invention migrates the BERE model and utilizes the model parameters pre-trained in the field of biomedical relation extraction. This transfer learning helps to achieve convergence faster in the task of extracting the relationship between diseases and gut microbiota and achieve better performance on a relatively small dataset. At the same time, the introduction of transfer learning and GAT may help the model discover some meaningful relationships not reported in existing databases. This is expected to provide new insights for biomedical research and promote the exploration of unknown fields.

[0094] 3. The introduction of the GAT network in the present invention helps to more effectively model the complex relationships between entities. The graph attention mechanism of GAT can highlight important nodes, thus better capturing the association between diseases and gut microbiota and improving the model's ability to model relationships. The word embedding technology of Gensims can capture text information, while the GAT network helps to capture relationship information in the text. By fusing these two types of information, the present invention can more comprehensively consider multi-modal features, thereby improving the comprehensive performance of relation extraction.

[0095] 4. The present invention adopts the transfer learning BERE framework, which shows competitiveness on medium and small-scale data sets. On the simulated data set, the accuracy of the model reaches 75%, the recall rate reaches 63%, and the f1-score is 68.7%, achieving good transfer. On the self-constructed data set, the accuracy of the model reaches 74%, the recall rate reaches 70%, and the f1-score is 73%, proving the effectiveness of the present invention in extracting the relationship between diseases and microorganisms.

[0096] The above-disclosed is only a preferred embodiment of a method for extracting the relationship between diseases and intestinal microorganisms based on a transfer learning model of the present invention. Of course, the scope of the rights of the present invention cannot be limited thereby. Those of ordinary skill in the art can understand all or part of the processes of implementing the above embodiments, and the equivalent changes made according to the claims of the present invention still fall within the scope covered by the invention.

Claims

1. A method for extracting the relationship between diseases and gut microbiota based on a transfer learning model, characterized in that, It includes the following steps: Preprocess the collected relevant literature data to obtain a training set, a validation set, and a test set; Use the training set, the validation set, and the test set to train a pre-trained model to obtain a relationship extraction model; Perform relationship extraction through the relationship extraction model to obtain the final relationship category.

2. The method for extracting the relationship between diseases and gut microbiota based on a transfer learning model according to claim 1, wherein The preprocessing of the collected relevant literature data to obtain a training set, a validation set, and a test set includes: Screen and organize the dataset collated from the BGMDB database to obtain relevant literature data; Perform text cleaning on the relevant literature data to obtain cleaned data; Perform entity annotation and relationship annotation on the cleaned data to obtain annotated data; Remove the word segmentation and stop words in the annotated data to obtain remaining data; Perform entity embedding representation on the remaining data to obtain post-embedding data; Perform data augmentation on the post-embedding data to obtain augmented data; Perform text vectorization on the augmented data to obtain data vectors; Divide the data vectors to obtain a training set, a validation set, and a test set.

3. The method for extracting the relationship between diseases and gut microbiota based on a transfer learning model according to claim 2, wherein The removing the word segmentation and stop words in the annotated data to obtain remaining data includes: Use NLTK to perform word segmentation on the annotated data and simultaneously remove the stop words in the annotated data to obtain remaining data.

4. The method for extracting the relationship between diseases and gut microbiota based on a transfer learning model according to claim 3, wherein The performing entity embedding representation on the remaining data to obtain post-embedding data includes: Use Gensims to train a Word2Vec model to obtain the embedding of disease and microorganism entity pairs to obtain an embedding representation; Perform entity embedding of the embedding representation on the remaining data to obtain post-embedding data.

5. The method for extracting the relationship between diseases and gut microbiota based on a transfer learning model according to claim 4, wherein The data augmentation method includes synonym replacement, entity exchange, sentence recombination, and generative adversarial network.

6. The method for extracting the relationship between diseases and gut microbiota based on a transfer learning model according to claim 5, wherein The performing text vectorization on the augmented data to obtain data vectors includes: Combine the outputs of NLTK and Gensims to convert the entire text into a vector to obtain data vectors.

7. The method for extracting the relationship between diseases and gut microbiota based on a transfer learning model according to claim 6, wherein The training set, the validation set, and the test set adopt a ratio of 8:1:

1.

8. The method for extracting the relationship between diseases and gut microbiota based on a transfer learning model according to claim 7, wherein The pre-trained model is a BERE model.

9. The method for extracting the relationship between diseases and gut microbiota based on a transfer learning model according to claim 8, wherein Performing relation extraction through the relation extraction model to obtain the final relation category, including: After obtaining the sentence bag mentioning the entity through the relation extraction model, word embedding, part-of-speech embedding, GAT feature learning, and aggregating sentences are sequentially performed to obtain the final relation category.