Computer device and method for transplanting and matching intestinal flora and computer product
By combining Transformer and GNN's TransGNN model, the intestinal microbiota data were analyzed, and the problem of inaccurate donor and receptor similarity assessment in intestinal microbiota transplantation was solved, precise matching of intestinal microbiota was achieved, the transplant success rate and efficacy were improved, and a personalized treatment plan was provided.
Patent Information
- Application Number
- CN202411893054.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-20
- Publication Date
- 2025-07-22
AI Technical Summary
In the existing intestinal flora transplantation technology, the intestinal flora similarity and complementarity assessment between donor and recipient is inaccurate, resulting in a low transplant success rate, making it difficult to accurately match specific indications, and the traditional methods are time-consuming and expensive.
Using a TransGNN model combined with Transformer and Graph Neural Network (GNN), the microecologic structure characteristics of the intestinal microbiota were constructed by analyzing 16S rRNA and metagenomic data, and combining hierarchical analysis method and intestinal microbiome health index (GHMI), the precise matching of donors and receptors was achieved.
It improves the success rate and efficacy of intestinal microbial transplantation, enhances the analytical accuracy of matching models, reduces time and cost, and provides personalized dietary guidance and auxiliary treatment methods.
Smart Images

Figure SMS_1 
Figure SMS_2 
Figure SMS_3
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of bioinformatics technology, and particularly relates to a computer device, a method and a computer product for intestinal microbiota transplantation matching. Background Art
[0002] The intestinal microbial community (intestinal microbiota) is closely related to human health and participates in many physiological processes, including digestion, absorption, immune system regulation, etc. In April 2011, researchers from the European Molecular Biology Laboratory (EMBL) in Heidelberg, Germany, conducted a global experiment - the International Human Intestinal Metagenome Research. They sequenced the fecal DNA samples of the research subjects using the Sanger sequencing method and found that there are three stable clusters of bacteria that appear in large numbers in the intestine: the Bacteroides type, the Prevotella type, and the Ruminococcus type.
[0003] People with different intestinal types have completely different constitutions and disease characteristics. Each of these three intestinal types has its own advantages. For example:
[0004] Intestinal type I: People of this type tend to have a high-fat and high-protein diet. Bacteroides can effectively decompose carbohydrates, and people of this intestinal type will be more likely to obtain energy from food;
[0005] Intestinal type II: People of this type tend to have a diet rich in plant-based carbohydrates. Prevotella tends to degrade intestinal mucin, and people of this intestinal type are often associated with non-Westernized eating habits;
[0006] Intestinal type III: People of this type tend to have a high-carbohydrate diet. Ruminococcus helps in the absorption of sugars, so people of this intestinal type may be more troubled by weight problems.
[0007] In recent years, fecal microbiota transplantation (FMT) has emerged as a new treatment method and has achieved remarkable results in the treatment of various diseases. However, due to the differences in intestinal microbiota among individuals, the success rate of intestinal microbiota transplantation is affected by the colonization effect of the microbiota between the donor and the recipient and the immune rejection reaction. Therefore, there are significant individual differences in the efficacy of FMT. Finding a suitable donor is crucial for improving the success rate of intestinal microbiota transplantation. How to accurately evaluate the similarity and complementarity of the intestinal microbiota between patients and donors and improve the intestinal microbiota colonization rate to enhance the efficacy of FMT has become an important problem to be solved.
[0008] First, the composition and function of the gut microbiota are influenced by multiple factors, including individual differences, environmental factors, and dietary habits. Therefore, simply comparing the gut enterotypes of donors and recipients cannot fully determine the optimal matching scheme. Second, the changes in the gut microbiota are a dynamic process, and the differences between donors and recipients may change over time, which also increases the difficulty of matching. Although the American Gastroenterological Association has developed some consensus and standardized regulations on FMT donor management, many difficulties still exist in the actual implementation process. For example, how to accurately evaluate the health status of donors and how to ensure that the microbial species in donor feces meet the treatment requirements are all issues that need to be addressed. The existing evaluation indicators for FMT donor gut microbiota mainly rely on parameters such as the abundance and diversity at the genus level related to enterotypes. However, these parameters can only reflect the richness of bacteria and the mechanism of action of single functional bacteria, and cannot fully reflect the functional characteristics of the systematic interaction of the gut microbiota. In the treatment applications of different FMT indications, this method has certain limitations. For example, among donors with the same enterotype, it is impossible to conduct precise matching analysis of the gut microbiota for a specific indication.
[0009] After standard donors are selected, it is necessary to strengthen the trust between donors and donor administrators, and the compliance, daily diet, exercise, and health status of donors all need to be followed up and monitored. Donors are screened intelligently at the levels of healthy diet and psychology. However, different patients may have very different responses to FMT, which may be due to various factors such as the composition of the gut microbiota, disease type, and treatment method of the patients. Therefore, not only should the above criteria be followed for donor health screening, but it is also necessary to analyze the functional gene set of the donor gut microbiota to provide intelligent matching decision support for different indications.
[0010] The characteristic gut microbiota for different indications is inconsistent. The gut microbiota composition of obese people is significantly different from that of normal-weight people. For example, the number of certain bacteria is higher in obese people, while the number of other bacteria is lower. These changes may lead to disorders in energy metabolism, thus increasing the risk of obesity. The gut microbiota composition of diabetic patients is also different from that of normal people. Some studies have found that the characteristic bacteria in the gut of diabetic patients affect the secretion and action of insulin, thus leading to an increase in blood sugar levels. The gut microbiota composition of cardiovascular disease patients is different from that of normal people. For example, changes in certain bacteria in the gut of cardiovascular disease patients may affect cholesterol metabolism and the occurrence of atherosclerosis. The dysregulation of the gut microbiota may be related to the occurrence and development of mental diseases. For example, the gut microbiota composition of depression patients is significantly different from that of healthy people, and it is affected by multiple factors such as gender, genetics, age, region, drugs, and diet, and undergoes dynamic changes continuously.
[0011] Gut microbiota transplantation matching requires simultaneous multi - mode quantitative and qualitative analysis of multi - omics of the microecosystem to capture the complexity of the molecular mechanisms of complex species and the heterogeneity of the microecosystem, which requires a keen understanding of the microbial world. Metagenomics studies microbial communities at the DNA level. In theory, it is possible to reconstruct the genomes of all microorganisms in a sample. However, this is a complex task because DNA sequencing technology can only generate fragments of the whole genome, and due to the incompleteness of the current reference database, the whole genomes of most microorganisms in environmental samples remain unknown. Moreover, due to the complexity, diversity, and heterogeneity of gut microbiomics, traditional gut bacteria research and analysis methods are improved by trained experts through innovative better models, algorithms, and approximations, which can be a time - consuming and expensive process.
[0012] Transformers have achieved great success in natural language processing and computer vision, but due to two important reasons, it is difficult to generalize them to medium - and large - scale graph data: (i) high complexity. (ii) failure to capture complex and entangled structural information. In graph representation learning, graph neural networks (GNNs) can fuse graph structure and node attributes, but the receptive field is limited. Summary of the Invention
[0013] The technical problems to be solved by the present invention are how to match the recipients and donors of gut microbiota transplantation and / or how to accurately evaluate the gut microbiota similarity and complementarity between the recipients and donors of gut microbiota transplantation and / or how to match the recipients and donors of gut microbiota transplantation based on the systematic interaction of gut microbiota microecosystem and / or how to perform specific precise gut microbiota matching analysis for specific indications among donors with the same enterotype and / or how to improve the intestinal bacteria colonization rate of gut microbiota transplantation to improve the efficacy of FMT.
[0014] To solve the above - mentioned technical problems, the present invention first provides a computer device, including a memory, a processor, and a computer program stored on the memory, wherein the processor executes the computer program to implement the following steps:
[0015] S1) Data reception: Receive the intestinal microbiome data and clinical data of the gut microbiota transplantation donor and the gut microbiota transplantation recipient; the microbiome data includes 16S rRNA data and metagenomic data;
[0016] S2) Data processing: After performing quality control on the microbiome data, quality-controlled microbiome data is obtained; the quality-controlled microbiome data includes quality-controlled 16S rRNA data and quality-controlled metagenomic data; data analysis is performed on the quality-controlled 16S rRNA data to obtain 16S rRNA data analysis results; the 16S rRNA data analysis results include the OUT of the intestinal microbiota characteristics table, the classification of the intestinal microbiota, and the species diversity of the intestinal microbiota; data analysis is performed on the quality-controlled metagenomic data to obtain metagenomic data analysis results; the metagenomic data analysis results include the classification of the intestinal metagenomic species, the gene information of the intestinal microbiota sequence annotation, and the metabolic pathways of the intestinal microbiota; based on the 16S rRNA data analysis results and the metagenomic data analysis results, a microecogenomics model is established by combining Transformer and GNN to obtain the microecological structure characteristics of the intestinal microbiota.
[0017] S3) Data output: Based on the analytic hierarchy process scoring algorithm, enterotype, gut microbiome health index (GHMI), and the microecological structure characteristics of the intestinal microbiota, a matching score is given to the intestinal microbiota transplant donor and the intestinal microbiota transplant recipient, and the matching result of the intestinal microbiota transplant donor and the intestinal microbiota transplant recipient is output.
[0018] In the above computer device, the construction of the TransGNN model may include the following steps:
[0019] A3-1) Attention sampling: Using the OUT of the intestinal microbiota characteristics table and the metabolic pathways of the intestinal microbiota as inputs, based on the semantic similarity and graph structure information in the attention sampling module, sampling and extraction of relevant nodes are performed for each central node to obtain representative nodes and features.
[0020] A3-2) Position encoding: Position encoding is performed on the representative nodes and features to obtain a position encoding result.
[0021] A3-3) TransGNN model training: Based on the 16S rRNA data analysis results and the metagenomic data analysis results, representative sequences (species sequences) of the microbial flora are obtained, and functional genes of the microbial flora are obtained based on the metagenomic data analysis results; a species sequence-species functional gene matrix is constructed based on the representative sequences of the microbial flora and the functional genes of the microbial flora; using the species sequence-species functional gene matrix as the input of the TransGNN model, the TransGNN model is trained to obtain the microecological structure characteristics of the intestinal microbiota.
[0022] To solve the above technical problems, the present invention also provides an intestinal microbiota transplant matching decision method, which may include the following steps:
[0023] A1) Data collection: Collect the intestinal microbiome data and clinical data of the intestinal microbiota transplantation donor and the intestinal microbiota transplantation recipient; the microbiome data includes 16S rRNA data and metagenomic data;
[0024] A2) Data analysis: After performing quality control on the microbiome data, the quality-controlled microbiome data is obtained; the quality-controlled microbiome data includes quality-controlled 16S rRNA data and quality-controlled metagenomic data; performing data analysis on the quality-controlled 16S rRNA data to obtain the 16S rRNA data analysis result; the 16S rRNA data analysis result includes the OUT of the intestinal microbiota characteristics table, the classification of the intestinal microbiota, and the species diversity of the intestinal microbiota; performing data analysis on the quality-controlled metagenomic data to obtain the metagenomic data analysis result; the metagenomic data analysis result includes the classification of the intestinal metagenomic species, the gene information of the intestinal microbiota sequence annotation, and the metabolic pathway of the intestinal microbiota;
[0025] A3) TransGNN model construction: Based on the 16S rRNA data analysis result and the metagenomic data analysis result, a microecological genomics model is established by combining Transformer and GNN to obtain the microecological structure characteristics of the intestinal microbiota;
[0026] A4) Based on the analytic hierarchy process scoring algorithm, perform typing scoring on the intestinal microbiota transplantation donor and the intestinal microbiota transplantation recipient based on the enterotype, the gut microbiome health index (GHMI), and the microecological structure characteristics of the intestinal microbiota to determine the typing result of the intestinal microbiota transplantation donor and recipient.
[0027] In the above method, the TransGNN model construction in A3) may include the following steps:
[0028] A3-1) Attention sampling: Using the OUT of the intestinal microbiota characteristics table and the metabolic pathway of the intestinal microbiota as inputs, based on the semantic similarity and graph structure information in the attention sampling module, sample and extract relevant nodes for each central node to obtain representative nodes and features;
[0029] A3-2) Position encoding: Perform position encoding on the representative nodes and features to obtain the position encoding result;
[0030] A3-3) TransGNN model training: Based on the 16S rRNA data analysis results and the metagenomic data analysis results, representative sequences (species sequences) of the microbial flora are obtained, and functional genes of the microbial flora are obtained based on the metagenomic data analysis results; a species sequence-species functional gene matrix is constructed based on the representative sequences of the microbial flora and the functional genes of the microbial flora; using the species sequence-species functional gene matrix as the input of the TransGNN model, the TransGNN model is trained to obtain the microecological structure characteristics of the intestinal flora.
[0031] To solve the above technical problems, the present invention also provides a device for intestinal flora transplantation matching, and the device may include the following modules:
[0032] B1) Data collection module: used to collect intestinal microbiome data and clinical data of the intestinal flora transplantation donor and the intestinal flora transplantation recipient; the microbiome data includes 16S rRNA data and metagenomic data;
[0033] B2) Data analysis module: used to perform quality control on the microbiome data to obtain quality-controlled microbiome data; the quality-controlled microbiome data includes quality-controlled 16S rRNA data and quality-controlled metagenomic data; data analysis is performed on the quality-controlled 16S rRNA data to obtain 16S rRNA data analysis results; the 16S rRNA data analysis results include the OUT of the intestinal microbial flora characteristics table, the classification of the intestinal microbial flora, and the species diversity of the intestinal microorganisms; data analysis is performed on the quality-controlled metagenomic data to obtain metagenomic data analysis results; the metagenomic data analysis results include the classification of intestinal metagenomic species, the gene information of intestinal flora sequence annotation, and the metabolic pathways of intestinal flora.
[0034] B3) TransGNN model construction module: used to establish a microecogenomics model using the combination of Transformer and GNN based on the 16S rRNA data analysis results and the metagenomic data analysis results to obtain the microecological structure characteristics of the intestinal flora;
[0035] B4) Matching scoring module: used to perform matching scoring on the intestinal flora transplantation donor and the intestinal flora transplantation recipient based on the analytic hierarchy process scoring algorithm, based on the intestinal type, the gut microbiome health index (GHMI), and the microecological structure characteristics of the intestinal flora, and determine the matching results of the intestinal flora transplantation donor and recipient.
[0036] In the above device, the construction of the TransGNN model in B3) may include the following steps:
[0037] B3-1) Attention sampling: Taking the OUT of the intestinal microbiota feature table and the intestinal flora metabolic pathway as inputs, based on the semantic similarity and graph structure information in the attention sampling module, sample and extract relevant nodes for each central node to obtain representative nodes and features;
[0038] B3-2) Position encoding: Perform position encoding on the representative nodes and features to obtain a position encoding result;
[0039] B3-3) TransGNN model training: Obtain representative sequences (species sequences) of the microbiota based on the 16S rRNA data analysis result and the metagenomic data analysis result, and obtain functional genes of the microbiota based on the metagenomic data analysis result; Construct a species sequence-species functional gene matrix based on the representative sequences of the microbiota and the functional genes of the microbiota; Use the species sequence-species functional gene matrix as the input of the TransGNN model, and train the TransGNN model to obtain the microecological structure characteristics of the intestinal flora.
[0040] To solve the above technical problems, the present invention also provides a computer-readable storage medium, on which a computer program / instructions are stored, and when the computer program / instructions are executed by a processor, the steps of the above-mentioned method are implemented.
[0041] To solve the above technical problems, the present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the following steps can be implemented:
[0042] C1) Data collection: Collect intestinal microbiome data and clinical data of intestinal flora transplantation donors and intestinal flora transplantation recipients; the microbiome data includes 16S rRNA data and metagenomic data;
[0043] C2) Data analysis: After performing quality control on the microbiome data, obtain quality-controlled microbiome data; the quality-controlled microbiome data includes quality-controlled 16S rRNA data and quality-controlled metagenomic data; Perform data analysis on the quality-controlled 16S rRNA data to obtain a 16S rRNA data analysis result; the 16S rRNA data analysis result includes the OUT of the intestinal microbiota, the classification of the intestinal microbiota, and the species diversity of the intestinal microbiota; Perform data analysis on the quality-controlled metagenomic data to obtain a metagenomic data analysis result; the metagenomic data analysis result includes the classification of intestinal metagenomic species, the information of annotated genes in the intestinal flora sequences, and the intestinal flora metabolic pathway;
[0044] C3) Construction of the TransGNN model: Based on the analysis results of the 16S rRNA data and the metagenomic data, a microecological genomics model is established by combining Transformer and GNN to obtain the microecological structure characteristics of the gut microbiota;
[0045] C4) Based on the analytic hierarchy process scoring algorithm, the gut microbiota transplant donor and the gut microbiota transplant recipient are typed and scored based on the enterotype, the gut microbiome health index (GHMI), and the microecological structure characteristics of the gut microbiota to determine the typing results of the gut microbiota transplant donor and recipient.
[0046] The purpose of the present invention is to provide a gut microbiota transplant typing decision method based on an artificial intelligence framework, which not only refers to the enterotype results of the recipient, but also increases the in-depth learning analysis of the gut microbiota and functional genes of the donor and recipient, clarifies the interaction relationship of the gut microbiota, realizes the precise typing of the gut microbiota in FMT indications, so as to make a better selection of FMT donors in combination with FMT indications, and improve the success rate of intestinal microbiota transplantation in FMT indications.
[0047] The present invention combines Transformer and GNN and adopts a new model of TransGNN, in which the Transformer layer and the GNN layer are alternately used to improve each other. Specifically, in order to expand the receptive field and untangle the edge information aggregation, Transformer is used to aggregate the information of more relevant gut microbiota gene nodes to improve the message passing of GNN. The TransGNN model framework consists of three important components: (1) an attention sampling module, (2) a positional encoding module, and (3) a TransGNN module.
[0048] The attention mechanism in this TransGNN model framework topological feature network diagram ( Figure 6 ) can combine the relative abundance of k-mers frequencies to estimate the importance of genes for specific species, extract representative nodes and features, which can be used to distinguish gene contributions and improve biological interpretability. This model has no assumptions and does not depend on the constraints of gene expression, so it can infer new gene expressions and regulatory relationships between species.
[0049] Here, the present invention proposes a method for making decisions on intestinal microbiota transplantation matching based on an artificial intelligence framework. It models the intestinal microbiota microecology in a GNN heterogeneous graph, constructs a method for matching the intestinal microbiota of healthy donors and indication recipients based on the Transformer deep learning model architecture, and uses the analytic hierarchy process scoring algorithm by combining intestinal microbiota indicators such as enterotype and gut microbiome health index (GHMI), which can significantly improve the analytical accuracy of the intestinal microbiota matching model and provide strong technical support for many downstream tasks such as the analysis of intestinal microbiota relationships, intestinal similarity matching, and microecological differences in different diseases.
[0050] The method for making decisions on intestinal microbiota transplantation matching based on an artificial intelligence framework proposed by the present invention uses the pyDeepInsight software (https: / / github.com / alok-ai-lab / pyDeepInsight), which allows non-image data to be converted into images through a mapping process without the need to manually construct meta-paths, establishes a microecological genomics model with a heterogeneous graph, models heterogeneous nodes and edges, and learns the dependence relationship features between species locally and globally to improve the robustness of the model. This method is more comprehensive and stable in the evaluation of intestinal microbiota transplantation matching decisions compared to the evaluation based on representative species enterotypes.
[0051] Therefore, inventing a method for making decisions on intestinal microbiota transplantation matching based on an artificial intelligence framework has important practical significance to improve the efficacy and safety of FMT indications.
[0052] Currently, the specific method for donor matching is not clearly defined in the consensus on intestinal microbiota transplantation. The method for making decisions on intestinal microbiota transplantation matching based on an artificial intelligence framework is proposed as a technological innovation. It is internationally recognized that both non-related standard donors and related donors can provide feces for patients, and standard donors should be preferentially used when conditions permit. The method for making decisions on intestinal microbiota transplantation matching based on an artificial intelligence framework in the present invention intelligently collects and quantifies the evaluation to establish a health data scale for FMT donor evaluation collection and identification, and conducts cross-combination intelligent scoring. Based on scientific scoring and grading, it automatically identifies the input information and dynamically plans the next-level data collection. A more refined scoring mechanism.
[0053] Not only taking the above-mentioned donor physical and mental health indicators as parameters, but also comprehensively considering the donor intestinal microbiota species, genes, metabolic pathways, and intestinal health indicators, training an artificial intelligence matching model based on indication intestinal microbiota data, and adding indication-related bacteria and related species gene co-evaluation indicators compared to the evaluation of representative enterotype bacteria.
[0054] The gut microbiota transplantation matching decision-making method based on the artificial intelligence framework in the present invention is an innovative application of multidisciplinary intersection, which combines the latest research results in multiple technical fields such as bioinformatics, microbiomics, clinical medicine, artificial intelligence and machine learning, medical data processing, and immunology. It has important clinical significance for the treatment of some diseases related to gut microbiota, such as digestive system diseases, metabolic diseases, mental system diseases, nervous system diseases, and enhancing the efficacy of tumor immunotherapy. In clinical nutrition, combined with the matching characteristics of the gut microbiome, it provides personalized dietary guidance for patients undergoing fecal microbiota transplantation to promote the recovery and reconstruction of gut microbiota. In drug research and development, aiming at the functional research of the effective gut microbiome after matching, new live bacterial drugs, probiotic preparations, etc. are developed to provide adjuvant treatment means for fecal microbiota transplantation to improve the safety and effectiveness of gut microbiota transplantation.
[0055] The gut microbiota transplantation matching decision-making method based on the artificial intelligence framework in the present invention is a method that uses artificial intelligence technology to assist doctors in making decisions on fecal microbiota transplantation matching. This method can improve the matching efficiency and comprehensively consider various factors, including the types, quantities, species antagonism, synergistic relationships, gene functional pathways, metabolites, etc. of gut microbiota, as well as clinical information such as the age, gender, and medical history of patients, so as to create more value for improving the clinical efficacy of fecal microbiota transplantation matching in China and the scientific research on the indications of fecal microbiota transplantation. BRIEF DESCRIPTION OF THE DRAWINGS
[0056] Figure 1 It is the key step flow chart of data processing in the present invention.
[0057] Figure 2 It is the scoring criteria of part of the screening questionnaire in clinical data.
[0058] Figure 3 It is an example of the scoring of part of the health screening in clinical data.
[0059] Figure 4 It is an example of the species sequence gene heterograph.
[0060] Figure 5 It is the training process of the TransGNN model.
[0061] Figure 6 It is the topological feature network graph.
[0062] Figure 7 It is the logic diagram of the analytic hierarchy process. DETAILED DESCRIPTION OF THE INVENTION
[0063] The present invention will be further described in detail below in conjunction with specific embodiments. The provided embodiments are only for clarifying the present invention and not for limiting the scope of the present invention. The following embodiments can be used as a guide for those of ordinary skill in the art to make further improvements, and do not limit the present invention in any way.
[0064] For the experimental methods in the following embodiments, unless otherwise specified, they are all conventional methods, carried out according to the techniques or conditions described in the literature in this field or according to the product instructions. The materials, reagents, etc. used in the following embodiments, unless otherwise specified, can all be obtained from commercial channels.
[0065] Example 1: Establishment of an Intestinal Microbiota Transplantation Matching Decision Method Based on an Artificial Intelligence Framework
[0066] The present invention adopts a heterogeneous graph transformer architecture for modeling heterogeneous graphs. This process uses an artificial intelligence framework, combining advanced algorithms of heterogeneous graph neural networks (GNNs) and Transformer models, to optimize the donor and recipient matching process for fecal microbiota transplantation (FMT).
[0067] The method of the present invention is based on multi-dimensional data quality control: collecting microbiome data and clinical-related data including clinical patients, i.e., recipients of fecal microbiota transplantation, and healthy donors. The microbiome data includes: 16S rRNA gene sequence and metagenomic sequencing data, which are used to describe the composition and structure of the microbiota, and these data can be obtained through high-throughput sequencing technology; the clinical data includes: the medical history of the patient, drug use situation, laboratory test results, etc., in order to better understand the relationship between the intestinal microbiota and host health. By analyzing the microbiome data of the intestinal microbiota, data quality control is carried out through screening questionnaire scoring criteria, and a basic database of fecal microbiota transplantation recipients (hereinafter referred to as recipients) and fecal microbiota transplantation donors (hereinafter referred to as donors) is established.
[0068] 1. Data collection of fecal microbiota transplantation recipients and donors
[0069] 1.1 Collection of microbiome data
[0070] First, collect fecal samples from donors and subjects (recipients), and perform 16S rRNA gene sequencing and metagenome sequencing by extracting the DNA of the fecal samples to obtain the original sequencing data of the intestinal microbiota data (including 16S rRNA gene sequencing data and metagenome sequencing data) of fecal microbiota transplantation recipients and donors.
[0071] 1.2 Collection of clinical data
[0072] Collect clinical information such as the medical history, drug use, and laboratory test results of the recipients and donors of fecal microbiota transplantation. The scoring criteria for some of the questionnaires for collecting information are shown in Figure 2 .
[0073] Through the analysis and score calculation of data such as the personal information, medical history, and lifestyle habits of the donors, evaluate and screen the health status and self-safety of the donors, and generate a detailed donor screening score table (example shown in Figure 3 ).
[0074] 2. Data Quality Control and Analysis
[0075] 2.1 Data Quality Control
[0076] Perform quality control on the raw sequencing data of the recipients and donors obtained in step 1.1 based on high-throughput sequencing, filter out low-quality sequences, and then perform preprocessing steps such as sequence splicing and chimera removal to obtain the 16S rRNA gene sequencing data after quality control for the recipients and donors and the Metagenome sequencing data after quality control.
[0077] 2.2 Data Integration and Analysis
[0078] Analyze the 16S rRNA gene sequencing data after quality control and the Metagenome sequencing data after quality control for the recipients and donors respectively to obtain the analysis results of the composition and structure of the intestinal microbiota of the recipients and donors.
[0079] Among them, the 16S rRNA gene sequencing data after quality control for the recipients and donors is analyzed using the qiime2 software (software official website https: / / qiime2.org / ), and the OTU (Operational Taxonomic Units) of the intestinal microbiota characteristics table of the recipients, the classification of the intestinal microbiota of the recipients, and the species diversity of the intestinal microbiota of the recipients are obtained respectively; the OTU (Operational Taxonomic Units) of the intestinal microbiota characteristics table of the donors, the classification of the intestinal microbiota of the donors, and the species diversity of the intestinal microbiota of the donors; the main steps of the analysis include: input entry, duplicate removal and quality control, generation of the OTU of the microbiota characteristics table, and annotation and classification to obtain the microbiota classification and species diversity.
[0080] After quality control, the Metagenome sequencing data was annotated and statistically analyzed using the Kraken 2 software (software official website: https: / / ccb.jhu.edu / software / kraken2 / ) based on sequence reads with k-mer (which refers to dividing a sequence into substrings containing k bases. If the read length is L and the k-mer length is set to k, the number of generated k-mers is: L - k + 1) and the LCA algorithm to obtain the taxonomic classification of the recipient and donor intestinal metagenomic species; the prokka software (official website: https: / / github.com / tseemann / prokka) was used to quickly annotate the genomic sequences of the recipient and donor intestinal flora to obtain the gene information of the annotated flora sequences; the HUMAnN3 software (software official website: https: / / huttenhower.sph.harvard.edu / humann / ) was used for genomic function annotation of the metagenomic sequences to obtain the metabolic pathways of the recipient and donor intestinal flora, and to obtain the species distance feature data between samples related to the pathways.
[0081] Extract the unique microbial species diversity in the donor intestine (Table 2, obtained based on the comparison of the microbial species diversity in the donor intestine and the recipient intestine), the unique taxonomic units at the phylum and genus levels of the flora in the donor intestine (obtained based on the comparison of the microbial flora classification in the donor intestine and the recipient intestine), the unique taxonomic units of the flora metabolic pathways in the donor intestine (Table 3 and Table 4, obtained based on the comparison of the metabolic pathways of the recipient and donor intestinal flora), and the compositional feature of the donor-recipient distance data used to compare the compositional differences between the donor and the recipient (obtained based on the intestinal microbial flora feature table OTU) (mainly referring to the species abundance (presence or absence) and evenness (relative abundance) of the donor and recipient. Community comparison methods based on OTU include: Euclidean distance, bray curtis distance, Jaccard distance). Based on these data, the following information in the recipient and donor samples was extracted: the representative sequences of the microbial flora (as species sequences) obtained from the analysis results of the 16S rRNA gene sequencing data and the metagenomic sequencing data, and the relative abundances at the species and genus levels of the microbial flora obtained from the analysis results of the 16S rRNA gene sequencing data; the presence of the functional genes of the microbial flora (as species functional genes) obtained from the metagenomic data to obtain the analysis results of the composition and structure of the recipient and donor intestinal flora.
[0082] Table 1. Results of Intestinal Microbial Health Index
[0083]
[0084] Note: According to the evaluation of the GMHI index, the human intestinal health status is divided into five categories: healthy (degree 2), healthy tendency (degree 1), sub-healthy (degree 0), unhealthy tendency (degree -1), and unhealthy (degree -2).
[0085] Table 2. Examples of the number of intestinal microorganism species, species diversity, enterotype, and ratio in an individual of a donor sample
[0086]
[0087] Note: Among them, the F / B ratio represents the ratio of Firmicutes to Bacteroidetes obtained based on the OTU of the intestinal microbiota characteristic table.
[0088] Table 3. Examples of evaluating the short-chain fatty acid synthesis ability of the microbial metabolic pathway in an individual of a donor sample
[0089] Short-chain fatty acid Synthetic potential value Synthetic ability assessment Formic acid 7.441 Normal synthetic ability Acetic acid 111.679 Slightly weak synthetic ability Propionic acid 55.202 Normal synthetic ability Butyric acid 71.066 Normal synthetic ability Valeric acid 19.732 Normal synthetic ability
[0090] Note: Among them, the synthesis potential value is the synthesis and metabolic ability calculated according to the composition of the intestinal flora and the abundance of functional genes of the flora through the KEGG PATHWAY database (https: / / www.genome.jp / kegg / ). The synthesis ability of essential amino acids in the intestine is evaluated by comparing the levels of relevant metabolic pathways in the population.
[0091] Table 4. Examples of evaluating the metabolic potential of nutrients by the microbial metabolic pathway in an individual of a donor sample
[0092]
[0093]
[0094] In the above analysis results of the composition and structure of the intestinal flora, the taxonomic units at the genus level of the flora are illustrated as follows:
[0095] Macro-grouping of intestinal microorganisms (distribution at the phylum level): According to the detection of the intestinal flora, the macroscopic distribution of microorganisms in the intestine is as follows: Firmicutes (40.085%), Bacteroidetes (50.967%), Actinobacteria (2.838%), Proteobacteria (5.885%), Verrucomicrobia (0.000%), Fusobacteria (0.012%), Others (0.212%);
[0096] Microscopic distribution of gut microbiota (core genus-level distribution): According to the detection of gut microbiota, the microscopic distribution of microorganisms in the gut is as follows: Bacteroides (44.950%), Ruminococcus (10.240%), Sutterella (5.206%), Roseburia (5.081%), Dialister (4.832%), Alistipes (3.859%), Lachnospira (3.608%), Bifidobacterium (2.660%), Parabacteroides (2.007%), Blautia (1.889%), Anaerostipes (1.258%), Coprococcus (1.244%), Streptococcus (1.030%).
[0097] 3. Construct the model framework
[0098] Based on the species sequences (classified species sequences) obtained in step 2 and the gene information (species functional genes) annotated by the microbiota sequence, a software pyDeepInsight (website: https: / / github.com / alok-ai-lab / pyDeepInsight) that converts non-image data into images is used to generate an integrated matrix, named the species sequence-species functional gene matrix, to represent the combined activity of each species functional gene in each microbiota species sequence; the species sequence-species functional gene matrix is used to construct a heterogeneous graph, where the species sequences and species functional genes are used as nodes, and the relationship between the classified species sequences and species functional genes is used as the edge, indicating the existence of the relationship between the species sequences and species functional genes.
[0099] Obtain the richness, microbiota structure, beneficial bacteria, and harmful bacteria of gut bacteria from the OTU of the gut microbiota characteristics table in step 2; construct a complex network structure model of gut microbiota from multiple dimensions such as beneficial pathways and harmful pathways obtained from the microbiota metabolic pathway taxonomic units in step 2; learn the low-dimensional representation of the global microecological map of a single microecology by spreading the characteristics of adjacent species and constructing the relationships between species, autonomously learn the unique patterns in the genomic sequences, and construct a heterogeneous graph with different types of nodes and edges; use the attention mechanism to deeply learn the model, improve the interpretability of the model, and infer the specific biological network of the gut microecological type.
[0100] The present invention combines Transformer and GNN, and adopts a new model of TransGNN, in which the Transformer layer and the GNN layer are alternately used to improve each other. Specifically, in order to expand the receptive field and untangle the edge information aggregation, Transformer is used to aggregate the information of more relevant microbial community gene nodes to improve the message passing of GNN. The TransGNN model framework consists of three important components: (1) an attention sampling module ( Figure 5 represented by "Attention SamplingModule" in Figure 5 ), (2) a positional encoding module ( Figure 5 represented by "Positional Encoding Module" in Figure 6 ), and (3) a TransGNN module (
[0101] represented by "TransGNN Module" in
[0102] ). The attention mechanism in the attention sampling module (taking the richness, flora structure, beneficial bacteria, and harmful bacteria of intestinal bacteria obtained from the OTU of the intestinal microbiota characteristics table in step 2; and the beneficial pathways and harmful pathways obtained from the microbiota metabolic pathway taxonomic unit in step 2 as inputs) can combine the relative abundance of k-mers frequencies to estimate the importance of genes for specific species, extract representative nodes and features, which can be used to distinguish gene contributions and improve biological interpretability. This model has no assumptions and does not rely on the constraints of gene expression, so new gene expressions and regulatory relationships between species can be inferred. The topological feature network diagram of this TransGNN framework is as shown in
[0103] . Using the Kraken2 and prokka software to quickly annotate the flora classification and flora genome, an integrated heterogeneous graph is constructed with species sequences and species functional genes as nodes and the relationships between species sequences and species functional genes as edges ( Figure 4 ).
[0104] Read the DNA sequence from the sample sequencing data and convert the original signal into one of the four bases. When finding the best alignment, the read data is decomposed into k-mers (k = 3 in this example, but usually much larger), that is, the data to be aligned is aligned in units of k-mers (nodes), and the matching k-mer units are regarded as successful alignments. The overlapping matching k-mers are aggregated and organized in the de Brujin graph to obtain the relative abundance of k-mer frequencies. Here, each alignment path from the root node (the k-mer unit at the 5' end of the sequence: C G A) to the end node (the k-mer unit at the 3' end of the sequence: (T T A), (C T A) or (G T A)) generates a contig. For example, the sequence C G AT T G T A is the alignment standard. The integers reported in the de Brujin graph correspond to the number of k-mer comparisons of different sequences. Finally, each contig is associated with a node in the assembly graph. The weight of the edge is the proportion of the sequences overlapping at the intersection of the contig pair. In this case, two read data are aligned to two edges.
[0105] 3.2 Modeling the gut microbiota microecology in the heterogeneous graph of graph neural network (GNN)
[0106] Graph neural network (GNN) has shown great strength in learning the dimensionality reduction representation of species genes by propagating the features of similar species and constructing species-species relationships in the global species graph, inferring the gut species type-specific biological network from the microecological data. Using the non-image data to image conversion method based on the pyDeepInsight software, a microecological genomics model is established with a heterogeneous graph, without manually constructing meta-paths, modeling heterogeneous nodes and edges, learning the dependence relationship features between species and species locally and globally, and improving the robustness of the model.
[0107] Arrange the species and gene features in adjacent regions of a two-dimensional feature map (d genes × N species) to facilitate learning the complex relationships and interactions between them, and convert the non-image data into a feature map image, that is, the GNN heterogeneous graph. Specifically, first, apply a non-linear dimensionality reduction technique, such as t-SNE or kernel principal component analysis (Kernel PCA), to convert the original features into a 2D embedded feature space. Second, use the convex hull algorithm to find the smallest rectangle containing all the features and perform rotation to align the feature map frame in a horizontal or vertical form. Finally, map the original feature values to the pixel coordinate positions of the feature map image.
[0108] 3.3 TransGNN model
[0109] Combining Transformer and GNN to establish a microecological genomics model. The TransGNN model captures the neighbor message passing (local relationship) and global topological features (global relationship) between species and functional genes. The integrated species-gene matrix is used to construct a heterogeneous graph, including annotated species (yellow) and functional genes (red) as nodes, and related species and genes (other colors) are predicted ( Figure 5 ). The TransGNN model architecture consists of three important components: (1) an attention sampling module, (2) a positional encoding module, and (3) a TransGNN module.
[0110] First, by considering the semantic similarity and graph structure information in the attention sampling module, the most relevant nodes are sampled for each central node. Then, in the positional encoding module, the positional encoding representing the nodes and features is calculated to help Transformer capture the graph topology information. After these two modules, the present invention uses the TransGNN module, which sequentially includes three sub-modules: (i) a Transformer layer, (ii) a GNN layer, and (iii) a sample update sub-module. Among them, the Transformer layer is used to expand the receptive field of the GNN layer and efficiently aggregate the attention sample information, while the GNN layer helps the Transformer layer perceive the graph structure information and obtain more relevant information of neighboring nodes. The sample update sub-module is used to effectively update the attention samples during new representations. The TransGNN model is trained on multiple subgraphs that cover as many nodes in the entire graph as possible. The trained model is applied to the entire graph to learn and update the intestinal microbiota microecological structure features by constructing species-species relationships and gene-gene relationships ( Figure 5 ).
[0111] The present invention combines the Transformer and GNN in the TransGNN module to make their advantages complementary. The specific steps of the three sub-modules ((i) Transformer layer, (ii) GNN layer, and (iii) samples update sub-module) included are as follows:
[0112] (i) Transformer layer: Use the Transformer layer to improve the GNN layer and expand the receptive field to more relevant nodes that may be far from the neighborhood.
[0113] q = h i W q
[0114]
[0115] h i = softmax(a t )V
[0116] Multi-head attention calculation:
[0117] MultiHead(h i )=Concat(head1,…,head m )W m
[0118] (ii) GNN layer: The GNN layer is used to fuse representation and graph structure to help the Transformer layer better utilize the graph structure.
[0119]
[0120] h i =Combine(h i ,h M ( i ))
[0121] (iii) Sample update submodule: After the Transformer layer and the GNN layer, the attention samples should be updated according to the new representation. Considering the locality of graph data, similar nodes are more likely to be contained in the neighborhood. The attention samples are updated by exploring the neighborhood of the attention samples to find possible new related nodes. A random walk strategy is used to explore the local neighborhood of each sampled node. The transition probability of the random walk is calculated based on the similarity:
[0122]
[0123] 4. Hierarchical Analysis
[0124] By combining enterobacteria indicators such as enterotype and gut microbiome health index (GHMI) to perform a hierarchical analysis method scoring algorithm, the analytical accuracy of the enterobacteria matching model was significantly improved.
[0125] The microecological structural characteristics of the donor recipients are analyzed and calculated. The characteristic hierarchy is decomposed by the hierarchical analysis method, the donors in the donor library are scored and ranked, the global weight is calculated, and the most suitable donor is recommended to the patient.
[0126] The analysis steps of enterotype are as follows:
[0127] The first step in enterotype analysis is to calculate the Jensen-Shannon divergence (JSD) between samples. JSD is a method to measure the difference between two probability distributions, which is suitable for comparing species composition differences between different samples. By calculating the JSD distance of relative abundance at the genus level, the similarities and differences between samples can be quantified.
[0128] The Calinski-Harabasz (CH) index is used to evaluate the clustering effect of the dataset. By calculating the CH index under different K values (i.e., the number of clusters), the K value corresponding to the highest CH value is selected as the optimal number of clusters to ensure that the samples are reasonably assigned to different enterotype categories.
[0129] The Partitioning Around Medoids (PAM) algorithm is used for sample clustering. Clustering is based on the central points, which can effectively classify the samples according to the predetermined number of clusters. After clustering, the quality of the clustering is evaluated by calculating the silhouette value of the samples. The higher the silhouette value, the more the samples conform to the assigned clustering results, ensuring the accuracy and reliability of the enterotype classification.
[0130] The identification of enterotypes relies on the analysis of specific indicator species or taxa. For example, Bacteroides is generally considered an indicator taxon for enterotype 1 (ET B); Prevotella drives enterotype 2 (ET P); while enterotype 3 (ETF) is mainly distinguished by the proportion of Firmicutes, among which Ruminococcus is the main taxon. By statistically analyzing the abundances of these marker species in different enterotypes, the enterotype classification of the samples can be further confirmed.
[0131] Example 2: Implementation of the decision-making method for intestinal microbiota transplantation matching based on the artificial intelligence framework
[0132] 1. Establish functions in the computer program to collect and calculate the clinical information of intestinal microbiota transplantation donors:
[0133] The first function of collecting donor information is to read the worksheet named "Donor ID" from a donor health questionnaire screening file, then convert the two columns of data (question column and target option column) in this worksheet into dictionary forms respectively, merge these two dictionaries into a new DataFrame object, transpose it so that the original questions are used as the index, and then convert the transposed DataFrame object into a dictionary form and print it out.
[0134] A function named title() is defined to read the basic information data in the donor health questionnaire screening file using the pandas library in Python. Specifically, it reads the first 20 rows of data (basic information) from the Excel file named 'Questionnaire Screening.xlsx'.
[0135] A function named livingHabit() is defined to calculate the score of living habits. Inside the function, a series of dictionaries are defined, representing the scoring criteria for different living habits respectively. Then, by traversing the questionnaire answers, the scores are calculated according to the corresponding scoring criteria, and all the scores are accumulated to get the total score. Finally, the total score is rounded to two decimal places and returned.
[0136] A function named dietaryHabit() is defined to calculate the score of eating habits. Inside the function, a series of dictionaries are defined, representing the scoring criteria for different eating habits respectively. Then, by traversing the questionnaire answers, the scores are calculated according to the corresponding scoring criteria, and all the scores are accumulated to get the total score. Finally, the total score is rounded to two decimal places and returned.
[0137] A function named gastrointestinalQualityOfLife() is defined to calculate the score of gastrointestinal quality of life. Inside the function, a series of dictionaries are defined, representing the scoring criteria for different gastrointestinal living habits respectively. Then, by traversing the questionnaire answers, the scores are calculated according to the corresponding scoring criteria, and all the scores are accumulated to get the total score. Finally, the total score is rounded to two decimal places and returned.
[0138] A function named mentalHealth() is defined to calculate the mental health score. Inside the function, a series of dictionaries are defined, representing the scoring criteria for different mental health problems respectively. Then, by traversing the questionnaire answers, the scores are calculated according to the corresponding scoring criteria, and all the scores are accumulated to get the total score. Finally, the total score is rounded to two decimal places and returned.
[0139] The last step of scoring is to write some health assessment results into a file named "Score.txt". First, it calls the title() function and writes the results into the file. Then, it calls the livingHabit(), dietaryHabit(), gastrointestinalQualityOfLife() and mentalHealth() functions and displays their results in the file. Finally, it calculates the total score and writes it into the file.
[0140] The above scores are obtained by the full-time donor management staff for scoring. It is preferred to select donors with scores above 90 for the next step of matching.
[0141] 2. Establish functions in the computer program to implement the analysis of intestinal flora sequencing data
[0142] For the intestinal microbiota 16S rRNA gene data, the QIIME 2.0 software package was used to cluster the sequencing fragments into ASVs (amplicon sequence variants). The representative sequences of the ASVs were aligned with the Silva database (or the Greengenes database) to obtain the taxonomic units (including phylum, class, order, family, genus, and species) corresponding to the ASVs and their corresponding abundance information.
[0143] The QIIME 2 core-diversity plugin was used to calculate the diversity matrix. Alpha diversity indices at the feature sequence level, including observed OTUs, Chao1, Shannon index, and Faith's phylogenetic diversity, were used to evaluate the diversity of the samples themselves. Beta diversity indices, including Bray Curtis, unweighted UniFrac, and weighted UniFrac, were used to evaluate the differences in the microbial community structures between samples, and then PCoA and NMDS plots were used for visualization.
[0144] Methods such as ANOVA, Kruskal Wallis, LEfSe, and DESeq2 were comprehensively applied to screen for species with different abundances between groups and samples.
[0145] RDA, correlation heatmap, etc. were used to analyze the compositional structure of the associated microbial communities and environmental factor data to find the microbial communities that are positively or negatively correlated with the physiological characteristics of the samples. Based on the relative abundances of the main microbial species in the samples, the Spearman rank correlation coefficient was calculated using co-occurrence analysis to understand the associations between species.
[0146] The PICRUSt software was used to predict the possible functional composition of the microbial community, and the components that differed between the transplant recipients and donors were identified, laying a foundation for further exploring the functions of the microbial community and analyzing the mechanism of action.
[0147] The intestinal microbiota species abundance data was read from a file named "silvaAbundance.txt", the species data in the first column was processed, and then the processed data and the corresponding species abundance data in the second column were written into a file named "silvanew.txt".
[0148] Each function is used to process different bacterial genera. These functions calculate the sum of the values of different bacterial genera in the text file and return a list containing the bacterial genus name and the "sum of bacterial values". If a certain bacterial genus is not detected, "bacterial genus name" and "not detected" will be returned.
[0149] Call a series of functions related to fecal microbiota transplantation (such as bacteroides(), prevotella(), etc.), and write the return results of these functions into a file named "Fecal Microbiota Transplantation - Related Bacterial Values.txt".
[0150] For the generation of the metagenome data species - gene heterogeneous graph, read the DNA sequences from the samples and convert the original signals into one of the four bases. When finding the best alignment, break the reads into k - mers (in this case k = 3, but usually larger), and align the matching k - mers. The overlapping k - mers are aggregated and organized in a de Bruijn graph. Here, each path generates a contig from the root (C G A) to an end node ((T T A), (C T A), (G T A)). For example, the sequence C G A T T G T A is a contig. The numbers reported in the de Bruijn graph correspond to the number of times the k - mers of different reads overlap. Finally, each contig is associated with a node in the assembly graph. The edge weight is the fraction of the read data that overlaps at the intersection of the contig pair. In this case, two reads align to both sides.
[0151] 3. Combining Transformer and GNN to build a microecological genomics model
[0152] Adopt the non - image data to image conversion method of the pyDeepInsight software (https: / / github.com / alok - ai - lab / pyDeepInsight) to build a microecological genomics model with a heterogeneous graph. Without manually constructing metapaths, model the heterogeneous nodes and edges, and learn the dependency relationship features between species locally and globally to improve the robustness of the model.
[0153] Combine Transformer and GNN to make their advantages complementary, which includes three sub - modules: (i) Transformer layer, (ii) GNN layer, (iii) samples update sub - module.
[0154] (i) Transformer layer: Use the Transformer layer to improve the GNN layer and expand the receptive field to more relevant nodes that may be far from the neighborhood.
[0155] q = h i W q
[0156]
[0157] h i = softmax(at )V
[0158] (ii) Multi - head attention calculation:
[0159] MultiHead(h i ) = Concat(head1, …, head m )W m
[0160] (iii) GNN layer: Use the GNN layer to fuse the representation and the graph structure to help the Transformer layer better utilize the graph structure.
[0161]
[0162] h i = Combine(h i , h M (υ i ))
[0163] Sample update sub - module: After the Transformer layer and the GNN layer, the attention samples should be updated according to the new representation. Considering the locality of graph data, similar nodes are more likely to be included in the neighborhood. Update the attention samples by exploring the neighborhood of the attention samples to find possible new relevant nodes. Use the random - walk strategy to explore the local neighborhood of each sampled node. The transition probability of the random walk is calculated according to the similarity as:
[0164]
[0165] 4. Calculate the gut microbiome health index GMHI:
[0166] Read the input files: the species relative abundance table species_relative_abundances.csv, the database of healthy dominant species table MH_species.txt, and the healthy scarce species table MN_species.txt.
[0167] Step 1: Run MetaPhlAn2 on the fecal metagenome using the '--tax_lev s' parameter.
[0168] Step 2: Use the'merge_metaphlan_tables.py' script provided in the MetaPhlAn2 pipeline to merge the outputs (see the online tutorial of MetaPhlAn2).
[0169] Step 3: Ensure the following: The merged species relative abundance profile should be in accordance with
[0170] Arrange as shown in'species_relative_abundances.csv'. Accordingly, the first column should contain the names of species-level taxa (i.e., taxonomic names with the's__' flag). Subsequent columns should contain the relative abundances of species corresponding to each metagenomic sample.
[0171] Step 4: Save the input data from Step 3 as a '.csv' file and run the following script to calculate the GMHI for each fecal metagenome. The GMHI values for each sample in'species_relative_abundances.csv' (https: / / www.nature.com / articles / s41467-020-18476-8) are shown in 'GMHI_output.csv'.
[0172] Preprocess the data matrix of species relative abundances:
[0174] species_profile_1: Results after removing unclassified and viral species
[0175] species_profile_2: Results after transposing species_profile_1
[0176] species_profile_3: Re-normalize the transposed data matrix of species relative abundances after removing unclassified and viral species, healthy dominant species (7 in total), healthy rare species (43 in total), extract healthy dominant species from fecal metagenomes, extract healthy rare species from fecal metagenomes, diversity among healthy dominant species, diversity among healthy rare species, richness of healthy dominant species, richness of healthy rare species, median RMH from the top 1% samples, median RMN from the bottom 1% samples, collective abundance of healthy dominant species, collective abundance of healthy rare species, save the GMHI results as 'GMHI_output.csv'.
[0177] Specific implementation of the Analytic Hierarchy Process: Used for comprehensive decision-making in typing calculations, diversity of the microbial community, number and abundance of characteristic bacteria predicted by the TransGNN model, assign weights to the gut microbiome health index and calculate. Four functions were defined first: distance() calculates the distance of the microbial community; diversity() calculates the diversity of the microbial community; Bacteria() calculates the number and abundance of positively and negatively correlated bacteria; GMHI() calculates the gut microbiome health index. Finally, these four functions were called respectively and the results of gut microbiota transplantation typing decisions were output.
[0178] The present invention has been described in detail above. For those skilled in the art, without departing from the gist and scope of the present invention and without the need for unnecessary experiments, the present invention can be implemented within a relatively wide range under equivalent parameters, concentrations and conditions. Although specific embodiments of the present invention are given, it should be understood that the present invention can be further improved. In short, according to the principle of the present invention, this application intends to cover any modifications, uses or improvements of the present invention, including those that depart from the scope disclosed in this application and are made with conventional techniques known in the art.
Claims
1. A computer device, comprising a memory, a processor, and a computer program stored on the memory, characterized in that, The processor executes the computer program to implement the following steps: S1) Data reception: Receive the intestinal microbiome data and clinical data of the intestinal microbiota transplantation donor and the intestinal microbiota transplantation recipient; the microbiome data includes 16S rRNA data and metagenomic data; S2) Data processing: After performing quality control on the microbiome data, obtain the quality-controlled microbiome data; the quality-controlled microbiome data includes quality-controlled 16S rRNA data and quality-controlled metagenomic data; perform data analysis on the quality-controlled 16S rRNA data to obtain the 16S rRNA data analysis results; the 16S rRNA data analysis results include the OUT of the intestinal microbiota characteristics table, the classification of the intestinal microbiota, and the species diversity of the intestinal microbiota; perform data analysis on the quality-controlled metagenomic data to obtain the metagenomic data analysis results; the metagenomic data analysis results include the classification of the intestinal metagenomic species, the gene information of the intestinal microbiota sequence annotation, and the metabolic pathway of the intestinal microbiota; Based on the 16S rRNA data analysis results and the metagenomic data analysis results, use the combination of Transformer and GNN to establish a microecological genomics model to obtain the microecological structure characteristics of the intestinal microbiota; S3) Data output: Based on the analytic hierarchy process scoring algorithm, the enterotype, the intestinal microbiome health index, and the microecological structure characteristics of the intestinal microbiota, perform a matching score on the intestinal microbiota transplantation donor and the intestinal microbiota transplantation recipient, and output the matching results of the intestinal microbiota transplantation donor and the intestinal microbiota transplantation recipient.
2. The computer device according to claim 1, wherein: The construction of the TransGNN model includes the following steps: A3-1) Attention sampling: Using the OUT of the intestinal microbiota characteristics table and the metabolic pathway of the intestinal microbiota as inputs, based on the semantic similarity and graph structure information in the attention sampling module, sample and extract relevant nodes for each central node to obtain representative nodes and features; A3-2) Position encoding: Perform position encoding on the representative nodes and features to obtain the position encoding results; A3-3) TransGNN model training: Based on the 16S rRNA data analysis results and the metagenomic data analysis results, obtain the representative sequences of the microbial flora, and obtain the functional genes of the microbial flora based on the metagenomic data analysis results; Construct a species sequence-species functional gene matrix based on the representative sequences of the microbial flora and the functional genes of the microbial flora; use the species sequence-species functional gene matrix as the input of the TransGNN model, and train the TransGNN model to obtain the microecological structure characteristics of the intestinal microbiota.
3. A decision-making method for intestinal microbiota transplantation typing, characterized in that: The method includes the following steps: A1) Data collection: Collect the intestinal microbiome data and clinical data of the intestinal microbiota transplantation donor and the intestinal microbiota transplantation recipient; the microbiome data includes 16S rRNA data and metagenomic data; A2) Data analysis: After quality control of the microbiome data, the quality-controlled microbiome data is obtained; the quality-controlled microbiome data includes quality-controlled 16S rRNA data and quality-controlled metagenomic data; data analysis is performed on the quality-controlled 16S rRNA data to obtain 16S rRNA data analysis results; the 16S rRNA data analysis results include the OUT of the intestinal microbiota characteristics table, the classification of the intestinal microbiota, and the species diversity of the intestinal microbiota; data analysis is performed on the quality-controlled metagenomic data to obtain metagenomic data analysis results; the metagenomic data analysis results include the classification of intestinal metagenomic species, the gene information of intestinal microbiota sequence annotation, and the metabolic pathways of intestinal microbiota. A3) TransGNN model construction: Based on the 16S rRNA data analysis results and the metagenomic data analysis results, a microecogenomics model is established by combining Transformer and GNN to obtain the microecological structure characteristics of the intestinal microbiota. A4) Based on the analytic hierarchy process scoring algorithm, the intestinal microbiota transplant donor and the intestinal microbiota transplant recipient are scored for matching based on the enterotype, the intestinal microbiome health index, and the microecological structure characteristics of the intestinal microbiota to determine the matching results of the intestinal microbiota transplant donor and recipient.
4. The method according to claim 1, wherein: A3) The construction of the TransGNN model includes the following steps: A3-1) Attention sampling: Using the OUT of the intestinal microbiota characteristics table and the metabolic pathways of the intestinal microbiota as inputs, based on the semantic similarity and graph structure information in the attention sampling module, sampling and extraction of relevant nodes are performed for each central node to obtain representative nodes and features. A3-2) Position encoding: Position encoding is performed on the representative nodes and features to obtain position encoding results. A3-3) TransGNN model training: Representative sequences of the microbial flora are obtained based on the 16S rRNA data analysis results and the metagenomic data analysis results, and functional genes of the microbial flora are obtained based on the metagenomic data analysis results. Based on the representative sequences of the microbial flora and the functional genes of the microbial flora, a species sequence-species functional gene matrix is constructed; using the species sequence-species functional gene matrix as the input of the TransGNN model, the TransGNN model is trained to obtain the microecological structure characteristics of the intestinal microbiota.
5. Device for intestinal microbiota transplantation typing, characterized in that: The device includes the following modules: B1) Data collection module: Used to collect the intestinal microbiome data and clinical data of the intestinal microbiota transplant donor and the intestinal microbiota transplant recipient; the microbiome data includes 16S rRNA data and metagenomic data. B2) Data analysis module: used to obtain quality-controlled microbiome data after quality control of the microbiome data; the quality-controlled microbiome data includes quality-controlled 16S rRNA data and quality-controlled metagenomic data; perform data analysis on the quality-controlled 16S rRNA data to obtain 16S rRNA data analysis results; the 16S rRNA data analysis results include the OUT of the intestinal microbiota characteristics table, the classification of the intestinal microbiota, and the species diversity of the intestinal microbiota; perform data analysis on the quality-controlled metagenomic data to obtain metagenomic data analysis results; the metagenomic data analysis results include the classification of intestinal metagenomic species, the gene information of intestinal microbiota sequence annotation, and the intestinal microbiota metabolic pathway. B3) TransGNN model construction module: used to establish a microecogenomics model by combining Transformer and GNN based on the 16S rRNA data analysis results and the metagenomic data analysis results to obtain the microecological structure characteristics of the intestinal microbiota. B4) Matching scoring module: used to perform matching scoring on the intestinal microbiota transplantation donor and the intestinal microbiota transplantation recipient based on the analytic hierarchy process scoring algorithm, based on the intestinal type, the intestinal microbiome health index, and the microecological structure characteristics of the intestinal microbiota, and determine the matching result of the intestinal microbiota transplantation donor and recipient.
6. The device according to claim 5, characterized in that: B3) The construction of the TransGNN model includes the following steps: B3-1) Attention sampling: using the OUT of the intestinal microbiota characteristics table and the intestinal microbiota metabolic pathway as inputs, based on the semantic similarity and graph structure information in the attention sampling module, sample and extract relevant nodes for each central node to obtain representative nodes and features. B3-2) Position encoding: perform position encoding on the representative nodes and features to obtain a position encoding result. B3-3) TransGNN model training: obtain representative sequences of the microbial flora based on the 16S rRNA data analysis results and the metagenomic data analysis results, and obtain functional genes of the microbial flora based on the metagenomic data analysis results. Construct a species sequence-species functional gene matrix based on the representative sequences of the microbial flora and the functional genes of the microbial flora; use the species sequence-species functional gene matrix as the input of the TransGNN model, and train the TransGNN model to obtain the microecological structure characteristics of the intestinal microbiota.
7. A computer-readable storage medium having computer programs / instructions stored thereon, characterized in that, When the computer program / instructions are executed by a processor, the steps of the method described in claim 3 or 4 are implemented.
8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, the following steps are implemented: C1) Data collection: collect the intestinal microbiome data and clinical data of the intestinal microbiota transplantation donor and the intestinal microbiota transplantation recipient; the microbiome data includes 16S rRNA data and metagenomic data. C2) Data analysis: After performing quality control on the microbiome data, quality-controlled microbiome data is obtained; the quality-controlled microbiome data includes quality-controlled 16S rRNA data and quality-controlled metagenomic data; data analysis is performed on the quality-controlled 16S rRNA data to obtain 16S rRNA data analysis results; the 16S rRNA data analysis results include the OUT of the intestinal microbiota characteristics table, the classification of the intestinal microbiota, and the species diversity of the intestinal microbiota; data analysis is performed on the quality-controlled metagenomic data to obtain metagenomic data analysis results; the metagenomic data analysis results include the classification of the intestinal metagenomic species, the gene information of the intestinal microbiota sequence annotation, and the metabolic pathway of the intestinal microbiota. C3) TransGNN model construction: Based on the 16S rRNA data analysis results and the metagenomic data analysis results, a microecogenomics model is established by combining Transformer and GNN to obtain the microecological structure characteristics of the intestinal microbiota. C4) Based on the analytic hierarchy process scoring algorithm, the intestinal microbiota transplant donor and the intestinal microbiota transplant recipient are typed and scored based on the enterotype, the intestinal microbiome health index, and the microecological structure characteristics of the intestinal microbiota to determine the typing results of the intestinal microbiota transplant donor and recipient.