A donor-recipient matching model for intestinal bacteria transplantation and its construction method

By constructing a donor receptor matching model, using metagenomic sequencing and neural network technology to screen out significantly different bacterial genus, it solves the problem of donor matching in intestinal microbial transplantation of gynecological diseases, improves the effectiveness and accuracy of treatment, and improves the patient's symptoms and hormone levels.

CN118824368BActive Publication Date: 2025-08-26SHANGHAI CHANGSHOU MEDICAL TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411040047.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-07-31
Publication Date
2025-08-26
Estimated Expiration
2044-07-31

AI Technical Summary

Technical Problem

In the prior art, it is difficult to determine the appropriate intestinal flora donor for gynecological diseases, resulting in insufficient effectiveness and accuracy of intestinal flora transplant treatment.

Method used

A donor receptor matching model was constructed. By collecting and analyzing the intestinal microbial data of patients and healthy donors, using metagenomic sequencing, metata analysis and neural network modeling, significantly different bacterial characteristics were screened to predict donor-receptor matching.

Benefits of technology

The successful construction of donor receptor matching model has improved the effectiveness and accuracy of intestinal microbiota transplantation in the treatment of gynecological diseases, and significantly improved the clinical symptoms and hormone levels of patients.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure SMS_5
    Figure SMS_5
  • Figure SMS_6
    Figure SMS_6
  • Figure SMS_7
    Figure SMS_7
Patent Text Reader

Abstract

The present invention provides a donor-recipient matching model for enterobacteria transplantation and a method for constructing the same. By combining donor and recipient information, the donor-recipient matching model is successfully constructed through metagenomic sequencing, meta-analysis, and neural network modeling. The model is applied to clinical experiments for performance comparison and evaluation of the reliability of the model, verifying the feasibility of the model prepared by the present invention and can be successfully applied to the treatment of gynecological diseases.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of automatic medical diagnosis and treatment, and in particular to a donor-recipient matching model for intestinal bacteria transplantation to treat gynecological diseases. Background Art

[0002] Benign gynecological diseases include uterine fibroids, endometriosis, polycystic ovary syndrome, pelvic inflammatory disease, etc. At present, common gynecological diseases are generally treated with GnRH analogs, oral contraceptives, steroids (such as prednisone) and / or anti-androgen drugs, and even require interventional treatment, which brings certain discomfort to patients. Intestinal bacteria play an important role in the physiological and immune functions of the human gastrointestinal tract. They can regulate intestinal inflammation and affect intestinal motility, absorption, secretion and permeability. Intestinal flora can change intestinal function and intestinal immunity, leading to functional gastrointestinal symptoms through the hypothalamic-pituitary-adrenal axis (HPA), autonomic nervous system (ANS), and central nervous system (CNS). Intestinal bacterial translocation can penetrate the intestinal epithelial barrier through transcellular and intercellular pathways. Intestinal flora transplantation is an emerging solution for the treatment of gynecological diseases (Fecal Microbiota Transplantation A Potential Tool for Treatment of Human Female Reproductive Tract Diseases, Quaranta et al. Frontiers in Immunology Volume 10). Intestinal flora transplantation (Fecal Microbiota Transplantation, referred to as FMT) is a medical therapy that transfers the flora in the feces of healthy people to the intestines of patients to adjust and restore the patient's intestinal microbial community. The so-called fecal microbiota transplantation is now generally believed to be the transplantation of feces from healthy individuals into the intestines of patients to achieve the purpose of treating diseases by improving and rebuilding the intestinal flora. However, matching suitable intestinal flora donors for gynecological diseases is a major problem in the existing technology. Ensuring the effectiveness and accuracy of flora transplantation treatment and improving the matching degree are technical problems that intestinal flora transplantation medical technology urgently needs to solve. Summary of the Invention

[0003] Based on the existing technology, the present invention is dedicated to providing a donor-recipient matching model and its construction method

[0004] The model is constructed as follows:

[0005] Step 1: Collect data and establish a database on the use of FMT to treat gynecological diseases

[0006] Collect intestinal flora data from clinical patient recipients and healthy donors and establish a database.

[0007] The collected data included clinical data: patient information, basic medical information, clinical symptoms, imaging examination results, follow-up data, and intestinal flora data of patients and healthy donors.

[0008] The clinical data includes the patient's basic information (such as name, gender, age, height, weight);

[0009] The basic medical information includes blood pressure, blood lipids, blood sugar, liver function test, kidney function test, etc.;

[0010] The clinical symptoms include menstrual cycle, hormone levels, and insulin resistance, wherein hormone levels include: follicle-stimulating hormone (FSH), luteinizing hormone (LH), estradiol (E2), progesterone (P), testosterone (T), prolactin (PRL), and androgen levels;

[0011] The imaging examination results include ovarian ultrasound images),

[0012] Furthermore, the collected data also includes follow-up data before and after treatment, including hormone levels, metabolic parameters, menstrual cycle changes, etc. before and after intestinal flora transplantation (FMT) treatment.

[0013] Among them, the intestinal flora data of patients and healthy donors were obtained by collecting fecal samples for metagenomic sequencing to obtain the complete genomic information of the intestinal flora.

[0014] Furthermore, information on the basic health status, dietary habits, medication use, and other relevant information of healthy donors should be collected, and the donor's metagenomic sequencing data should be recorded. All collected data should be stored in a structured database for subsequent analysis and model building.

[0015] To store and manage the above data, the following specific methods and database types are used:

[0016] 1. Relational database: MySQL, used to store structured data, such as clinical data, basic patient information, basic medical information, clinical symptoms, imaging examination results, and follow-up data. This data can be stored using a standardized database table structure, with each table dedicated to different types of data. For example:

[0017] `patients` table: stores basic information of patients (such as name, gender, age, etc.).

[0018] `clinical_data` table: stores the patient's basic medical information (such as blood pressure, blood lipids, blood sugar, etc.).

[0019] `symptoms` table: stores the patient's clinical symptoms (such as menstrual cycle, hormone levels, etc.).

[0020] `imaging_results` table: stores imaging examination results (such as ovarian ultrasound images).

[0021] `follow_up_data` table: stores follow-up data before and after treatment.

[0022] 2. NoSQL database: MongoDB, used to store unstructured or semi-structured data, such as metagenomic sequencing data, dietary habits and medication usage of healthy donors. MongoDB allows the storage of documents in JSON format, which is very suitable for storing complex metagenomic data. For example:

[0023] `metagenomics_data` collection: stores metagenomic sequencing data of patients and healthy donors.

[0024] `donor_health_data` collection: stores the health status, eating habits, medication usage, etc. of healthy donors.

[0025] 3. Data Warehouse: Google BigQuery, used to store and analyze large amounts of historical data, supports complex queries and data analysis. Data from relational and NoSQL databases can be regularly imported into the data warehouse for large-scale data analysis.

[0026] 4. Data backup and security: Back up the database regularly to ensure data integrity and security. Use SSL / TLS to encrypt data transmission and use access control and permission management to protect sensitive data.

[0027] Step 2: Metagenomic Sequencing Analysis

[0028] Metagenomic sequencing analysis is performed based on the same parameters to obtain detailed intestinal flora composition and functional gene information. This includes genome extraction, metagenomic sequencing, quality control processing, sequence reading, functional annotation, diversity analysis, and data analysis. Furthermore, the metagenomic sequencing is performed using a high-throughput sequencing platform;

[0029] Furthermore, the quality control processing includes removing low-quality sequences and adapter sequences, and splicing high-quality sequencing reads using software tools;

[0030] Furthermore, the functional annotation is used to identify gene functions and metabolic pathways, and the classification annotation is used to identify the types and relative abundances of microorganisms in the sample;

[0031] Furthermore, the diversity analysis includes species diversity analysis, further including Alpha diversity, Beta diversity, and Gamma diversity;

[0032] Furthermore, the data analysis is used to analyze the relative abundance of different bacterial genera in the samples, identify significantly different bacterial genera between patients and healthy donors, and use visualization tools to display the bacterial genus composition characteristics.

[0033] Among them, the genome extraction step preferably extracts the genome from the collected fecal sample, specifically the total DNA and / or RNA sample, and further determines the concentration and quality of the DNA and / or RNA. For example, gel electrophoresis can be used to confirm the size and purity of the RNA / DNA, preferably DNA as the genome sample; the extraction method can be completed using an extraction kit known in the prior art; ensure that the sample quality and concentration meet the requirements of metagenomic sequencing.

[0034] Among them, metagenomic sequencing uses a high-throughput sequencing platform to perform metagenomic sequencing on the sample. The high-throughput sequencing platform can be a platform known in the prior art, such as Illumina, PacBio, Ion Torrent sequencing platform, 454FLX pyrophosphate sequencing platform, DNBSEQ-T7, MGISEQ-2000, Solexa genome analysis platform, etc., preferably Illumina and PacBio platforms; during the library construction process, it is necessary to ensure the uniformity of fragmented DNA and the high quality of the library.

[0035] The raw sequencing data undergoes quality control to remove low-quality sequences and adapter sequences to ensure data quality for subsequent analysis. Software tools are then used to splice high-quality sequencing reads together to reconstruct the microbial genome fragments in the sample; for example, MEGAHIT is used to splice sequencing data into contigs.

[0036] The contigs are functionally annotated using a database to identify gene functions and metabolic pathways, while the sequences are classified and annotated using a classification database to identify the microbial species and relative abundance in the sample. The annotation database can be KEGG, and the classification database can be SILVA.

[0037] Among them, diversity analysis includes calculating the species diversity indicators of the sample, including but not limited to Chao1 index, Ace index, Shannon index, Simpson index, Richness index, goods_coverage index, Pielou index, invsimpson index, preferably Shannon index and Chao1 index, for evaluating the richness and evenness of species in the sample (Alpha diversity); and using methods such as PCoA and NMDS to compare the diversity between samples and evaluate the differences in microbial communities between different samples (Beta diversity); and evaluating the species diversity (Gamma diversity) of multiple samples (communities) at a regional or terrestrial scale; preferably using Alpha diversity and Beta diversity for evaluation.

[0038] Among them, the relative abundance of different bacterial genera in the sample is analyzed, the significantly different bacterial genera between patients and healthy donors are identified, and the bacterial genus composition characteristics are displayed using visualization tools to facilitate subsequent analysis and model building; the visualization tools can be stacked bar charts, line charts, heat maps, etc.

[0039] Through the above-mentioned metagenomic sequencing and analysis steps, detailed intestinal flora composition, functional gene information and diversity characteristics can be obtained, providing basic data for subsequent neural network modeling and donor-recipient matching model construction.

[0040] Step 3: Meta-analysis

[0041] Meta-analysis was performed on the data in the database through metagenomic analysis to obtain the alpha diversity, beta diversity, genus composition characteristics and confidence scores of the intestinal flora.

[0042] That is, organize and standardize the data, and evaluate the species richness and evenness within each sample through different indicators; calculate alpha diversity, and use indicators such as Shannon index, Chao1 index and Simpson index to evaluate the species richness and evenness within each sample; calculate beta diversity, and evaluate the microbial differences between different samples through methods such as Bray-Curtis distance, Jaccard index or weighted UniFrac distance; then, perform functional gene annotation and differential function analysis, use PICRUSt2 to predict functional genes and metabolic pathways, and compare the differences in functional gene abundance between different groups; optionally, the results can be visualized.

[0043] Among them, data is organized and standardized to ensure that metagenomic data from different samples are comparable.

[0044] The species richness and evenness within each sample were evaluated using different indicators. Among them, alpha diversity was evaluated using indicators such as the Shannon index, Chao1 index, and Simpson index. These indicators were calculated using the QIIME2 tool.

[0045] The Shannon diversity index calculation formula is:

[0046] Among them, p i represents the relative abundance of the i-th microorganism, and S represents the total number of species in the sample.

[0047] The Chao1 richness estimation formula is:

[0048] Among them, S obs represents the number of species observed, F1 represents the number of species that appear only once in a single sample, and F2 represents the number of species that appear twice in a single sample.

[0049] The Simpson index calculation formula is:

[0050] Among them, p i represents the relative abundance of the i-th microorganism, and S represents the total number of species in the sample.

[0051] Among them, beta diversity was calculated, and the differences in bacterial communities between different samples were evaluated by methods such as Bray-Curtis distance, Jaccard index, or weighted UniFrac distance. The calculation method is:

[0052] The Bray-Curtis distance formula is:

[0053] Among them, X ik and X jk represents the abundance of the kth microorganism in the i-th and j-th samples, and n represents the total number of species in the sample. These distance matrices were generated using QIIME2 and the vegan package in R language and visualized using principal coordinate analysis (PCoA) or non-metric multidimensional scaling (NMDS). In addition, differential microbiome analysis was performed, using the LEfSe (linear discriminant analysis effect size) method to identify bacterial genera that were significantly different between patients and healthy donors. Through these tools, differences in microbiome composition between different groups can be statistically analyzed and compared.

[0054] Among them, functional gene annotation and differential function analysis were performed, PICRUSt2 was used to predict functional genes and metabolic pathways, and the differences in functional gene abundance between different groups were compared.

[0055] Furthermore, a confidence score is calculated to determine the reliability of the data. The confidence score is determined by evaluating the relative abundance and distribution of the target bacterial genus in each sample. The confidence score calculation formula is:

[0056]

[0057] Among them, p i,observed represents the observed relative abundance of the i-th microorganism, p i,expected The above meta-analysis steps can systematically evaluate the differences in intestinal flora between patients and healthy donors, providing key data for subsequent neural network modeling and donor-recipient matching model construction.

[0058] Step 4: Modeling via Neural Network

[0059] Using the NetMoss network model and the metagenomic data obtained from meta-analysis, a network diagram was constructed based on the co-occurrence frequency and correlation between microorganisms to screen out the significantly different genus characteristics between patients and healthy donors. The NetMoss network model is a network analysis method based on the co-occurrence relationship of microorganisms; specifically, it includes constructing a microbial collinearity matrix, constructing a collinear network, screening differential genus characteristics, and constructing a genus characteristic model.

[0060] Construct a microbial co-occurrence matrix: Based on the metagenomic data, calculate the co-occurrence relationship between each pair of bacterial genera. Use the Spearman rank correlation coefficient to calculate the co-occurrence relationship matrix between bacterial genera:

[0061]

[0062] Among them, p ij represents the Spearman rank correlation coefficient between bacterial genera i and j, cov(X i , X j ) represents the covariance of bacterial genera i and j, σx i and σx j represent the standard deviation of bacterial genus i and j, respectively.

[0063] Construct a co-occurrence network: Convert the co-occurrence matrix into a network graph, where nodes represent bacterial genera and edges represent co-occurrence relationships between genera. By setting a threshold, significant co-occurrence relationships (e.g., relationships with an absolute value of the Spearman rank correlation coefficient greater than 0.6) were screened.

[0064] Differential bacterial genera screening: The NetMoss method was used to identify significantly different bacterial genera between patients and healthy donors. NetMoss screened for bacterial genera that were significantly altered in patients by comparing the network structure differences between the two groups.

[0065] The specific steps are as follows:

[0066] Calculate the weighted degree centrality of two sets of networks:

[0067] C w (v)=Σ u∈N(v) w(v,u)

[0068] Among them, C w (v) represents the weighted degree centrality of node v, N(v) represents the set of nodes adjacent to node v, and w(v, u) represents the edge weight between nodes v and u.

[0069] The weighted degree centrality differences of bacterial genera in the patient and healthy donor networks were compared to screen out bacterial genera that were significantly upregulated or downregulated in patients.

[0070] Construct a genus signature model: Incorporate the identified differential genus signatures into the model using machine learning algorithms (such as support vector machines and random forests) to predict donor-recipient compatibility. During model training and validation, use the confidence scores and distance matrix calculated in step 3 as input features to optimize model parameters and evaluate model performance.

[0071] Confidence score calculation: Calculate the confidence score for each donor and recipient sample to assess the degree of match with the target genus. The confidence score calculation formula is as follows:

[0072]

[0073] Among them, p i,pbserved represents the observed relative abundance of the i-th microorganism, p i,expected Indicates the expected relative abundance.

[0074] Through the above steps, the significantly different bacterial genus characteristics between patients and healthy donors can be screened out, and a prediction model can be constructed based on these characteristics to provide a scientific basis for donor-recipient matching.

[0075] Step 5: Construct a donor-recipient matching model

[0076] Organize the collected data and select a suitable neural network model. Use these feature data as input layer nodes to train and verify the model, and finally build a donor-recipient matching model.

[0077] Among them, the data collected and organized are the data obtained in steps 1 to 4, including alpha diversity, beta diversity, bacterial genus composition characteristics, functional gene annotations, confidence scores, etc. Then, a suitable neural network model (such as multilayer perceptron, convolutional neural network, etc.) is selected, and these feature data are used as input layer nodes for model training and verification. During the training process, the cross-validation method is used to evaluate the performance of the model and select the optimal network structure and parameters. The trained model is used to predict the matching degree of the donor and the recipient. The distance between the donor and the recipient can be calculated by the weighted UniFrac distance:

[0078] Among them, p i,donor and p i,recipient represents the relative abundance of the i-th microorganism in the donor and recipient, respectively, w i represents the weight of the i-th microorganism.

[0079] Based on the distance calculation results and confidence scores, a neural network model is used to predict the optimal donor-recipient pairing. The model's output layer provides a matching score for each donor-recipient pair, which is used to assess their compatibility.

[0080] Donor-recipient matching model

[0081] The intestinal flora transplantation donor-recipient matching model prepared by the construction method of steps 1-5.

[0082] The invention further provides the use of the donor-recipient matching model in preparing a drug for treating a disease that can be treated by intestinal flora transplantation.

[0083] The "disease that can be treated by intestinal flora transplantation" is preferably a gynecological disease.

[0084] Gynecological diseases are further gynecological diseases that can be treated by intestinal bacteria transplantation; preferably, they are benign gynecological diseases.

[0085] Gynecological conditions include, for example, ovarian cysts, polycystic ovary syndrome (PCOS), intrauterine adhesion syndrome, premature ovarian failure, hyperprolactinemia, pelvic inflammatory disease, menopausal syndrome, early-stage cancer, low-grade cancer, etc.

[0086] The early stage cancer and low-grade cancer include, for example, low-grade ovarian cancer and early endometrial cancer.

[0087] Furthermore, the present invention provides a donor-recipient matching model for intestinal bacteria transplantation to treat gynecological diseases.

[0088] Beneficial effects

[0089] The intestinal bacteria transplant donor-recipient matching model provided by the present invention combines the donor and recipient information and successfully constructs the donor-recipient matching model through metagenomic sequencing, meta-analysis, and neural network modeling. The model is applied to clinical experiments for performance comparison and evaluation of the reliability of the model, verifying the feasibility of the model prepared by the present invention and can be successfully applied to the treatment of various gynecological diseases. DETAILED DESCRIPTION

[0090] Example 1 Method for constructing a donor-recipient matching model for enterobacteria transplantation

[0091] Step 1: Collect data and establish a database on the use of FMT to treat gynecological diseases

[0092] First, clinical data of specific patients with gynecological diseases are collected, including basic information of patients (such as age, height, weight), clinical symptoms (such as menstrual cycle, androgen levels, insulin resistance), imaging examination results (such as ovarian ultrasound images), and follow-up data before and after treatment. These follow-up data include hormone levels, metabolic parameters, menstrual cycle changes before and after intestinal microbiota transplantation (FMT, Fecalmicrobiota transplantation), etc. Secondly, intestinal microbiota data of patients and healthy donors are collected, and metagenomic sequencing is performed by collecting fecal samples to obtain complete genomic information of the intestinal microbiota. In addition, relevant information such as basic health status, dietary habits, and drug use of healthy donors is collected, and the donor's metagenomic sequencing data is recorded. All collected data should be stored in a structured database for subsequent analysis and model construction.

[0093] Step 2: Metagenomic Sequencing Analysis

[0094] Metagenomic sequencing analysis is performed based on the same parameters to obtain detailed intestinal flora composition and functional gene information. The specific steps are as follows: First, total DNA is extracted from the collected fecal samples to ensure that the sample quality and concentration meet the requirements of metagenomic sequencing. Then, a high-throughput sequencing platform (such as Illumina or PacBio) is used to perform metagenomic sequencing on the samples. During the library construction process, the uniformity of the fragmented DNA and the high quality of the library need to be ensured. Next, the raw data obtained by sequencing is quality controlled to remove low-quality sequences and adapter sequences to ensure the data quality of subsequent analysis. Afterwards, the software tool (MEGAHIT) is used to splice high-quality sequencing reads into contigs to reconstruct the microbial genome fragments in the sample. Next, the database (KEGG) is used to functionally annotate the spliced ​​contigs to identify gene functions and metabolic pathways. At the same time, the classification database (SILVA) is used to classify and annotate the sequences to identify the microbial species and relative abundance in the sample. Then, diversity analysis is performed, including calculating the species diversity indicators of the samples (such as the Shannon index and Chao1 index) to evaluate the richness and evenness of species within the samples (Alpha diversity), and using methods such as PCoA and NMDS to compare the diversity between samples and evaluate the differences in the microbial communities between different samples (Beta diversity). Finally, the relative abundance of different bacterial genera in the samples is analyzed, and the significantly different bacterial genera between patients with gynecological diseases and healthy donors are identified. Visualization tools such as stacked bar charts or heat maps are used to display the genus composition characteristics to facilitate subsequent analysis and model construction. Through the above-mentioned metagenomic sequencing analysis steps, detailed intestinal microbial community composition, functional gene information, and diversity characteristics can be obtained, providing basic data for subsequent neural network modeling and donor-recipient matching model construction.

[0095] Step 3: Meta-analysis

[0096] In step 3, we performed a meta-analysis on the data in the database using metagenomic analysis to obtain the alpha diversity, beta diversity, genus composition characteristics, and confidence scores of the intestinal flora. The specific steps are as follows:

[0097] First, the data was collated and standardized to ensure comparability between metagenomic data from different samples. Alpha diversity was then calculated, and metrics such as the Shannon index, Chao1 index, and Simpson index were used to assess species richness and evenness within each sample. These metrics were calculated using the QIIME2 tool.

[0098] The Shannon diversity index calculation formula is:

[0099] Among them, p irepresents the relative abundance of the i-th microorganism, and S represents the total number of species in the sample.

[0100] The Chao1 richness estimation formula is:

[0101] Among them, S obs represents the number of species observed, F1 represents the number of species that appear only once in a single sample, and F2 represents the number of species that appear twice in a single sample.

[0102] The Simpson index calculation formula is:

[0103] Among them, p i represents the relative abundance of the i-th microorganism, and S represents the total number of species in the sample.

[0104] Next, beta diversity was calculated, and the differences in bacterial communities between different samples were assessed using methods such as Bray-Curtis distance, Jaccard index, or weighted UniFrac distance.

[0105] The Bray-Curtis distance formula is:

[0106] Among them, X ik and X jk represents the abundance of the kth microorganism in the i-th and j-th samples, and n represents the total number of species in the samples.

[0107] These distance matrices were generated using QIIME2 and the R language's vegan package and visualized using principal coordinate analysis (PCoA) or nonmetric multidimensional scaling (NMDS). Furthermore, differential microbiome analysis was performed, using the LEfSe (Linear Discriminant Analysis Effect Size) method to identify bacterial genera that were significantly different between patients and healthy donors. These tools allow statistical analysis and comparison of differences in microbiome composition between different groups.

[0108] Then, functional gene annotation and differential function analysis were performed, and functional genes and metabolic pathways were predicted using PICRUSt2. The functional gene abundance differences between different groups were compared. In addition, confidence scores were calculated. The confidence score was determined by evaluating the relative abundance and distribution of the target bacterial genus in each sample. The confidence score calculation formula is:

[0109]

[0110] Among them, p i,observed represents the observed relative abundance of the i-th microorganism, p i,expectedThe expected relative abundance is represented by the expression "α" in the table. The analysis results were visualized using heatmaps, stacked bar charts, or scatter plots to display alpha diversity, beta diversity, and differential microbial communities. These plots were generated using the R package ggplot2. Through the above meta-analysis steps, differences in the gut microbiota between patients and healthy donors can be systematically assessed, providing key data for subsequent neural network modeling and donor-recipient matching model construction.

[0111] Step 4: Modeling via Neural Network

[0112] In step 4, we used the NetMoss network to model and screen out the significantly different bacterial genus characteristics between patients and healthy donors. The specific steps are as follows:

[0113] First, we constructed a NetMoss network model. NetMoss is a network analysis method based on microbial co-occurrence relationships. It constructs a network graph based on the co-occurrence frequency and correlation between microorganisms, from which we screen for significantly different genus characteristics. We used the metagenomic data obtained in step 3 to construct a co-occurrence network for patients and healthy donors.

[0114] Constructing a microbial co-occurrence matrix: Based on metagenomic data, the co-occurrence relationship between each pair of bacterial genera was calculated. We used the Spearman rank correlation coefficient to calculate the co-occurrence relationship matrix between bacterial genera:

[0115]

[0116] Among them, p ij represents the Spearman rank correlation coefficient between bacterial genera i and j, cov(X i , X j ) represents the covariance of bacterial genera i and j, σx i and σx j represent the standard deviation of bacterial genus i and j, respectively.

[0117] Construct a co-occurrence network: Convert the co-occurrence matrix into a network graph, where nodes represent bacterial genera and edges represent co-occurrence relationships between genera. By setting a threshold, significant co-occurrence relationships (e.g., relationships with an absolute value of the Spearman rank correlation coefficient greater than 0.6) were screened.

[0118] Differential bacterial genera screening: The NetMoss method was used to identify significantly different bacterial genera between patients and healthy donors. NetMoss screened for bacterial genera that were significantly altered in patients by comparing the network structure differences between the two groups. The specific steps are as follows:

[0119] Calculate the weighted degree centrality of two sets of networks:

[0120] C w (v)=∑ u∈N(v) w(v,u)

[0121] Among them, C w (v) represents the weighted degree centrality of node v, N(v) represents the set of nodes adjacent to node v, and w(v, u) represents the edge weight between nodes v and u.

[0122] Compare the weighted degree centrality differences of bacterial genera in the patient and healthy donor networks, and screen out bacterial genera that are significantly upregulated or downregulated in patients. The recommended genera for model construction include:

[0123] Significantly upregulated bacterial genera: Prevotella, Clostridium, Bacteroides;

[0124] Significantly downregulated bacterial genera: Lactobacillus, Bifidobacterium, Faecalibacterium, Akkermansia;

[0125] Construct a genus signature model: Incorporate the identified differential genus signatures into the model using machine learning algorithms (such as support vector machines and random forests) to predict donor-recipient compatibility. During model training and validation, use the confidence scores and distance matrix calculated in step 3 as input features to optimize model parameters and evaluate model performance.

[0126] Confidence score calculation: Calculate the confidence score for each donor and recipient sample to assess the degree of match with the target genus. The confidence score calculation formula is as follows:

[0127]

[0128] Among them, p i,pbserved represents the observed relative abundance of the i-th microorganism, p i,expected Indicates the expected relative abundance.

[0129] Through the above steps, we can screen out the bacterial genus characteristics that are significantly different between patients and healthy donors, and build a predictive model based on these characteristics to provide a scientific basis for donor-recipient matching.

[0130] Step 5: Construct a donor-recipient matching model. Using the data and features obtained from steps 1 to 4, construct a donor-recipient matching model suitable for gynecological diseases. The specific steps are as follows:

[0131] First, collect and organize the data obtained in steps 1 to 4, including alpha diversity, beta diversity, bacterial genus composition characteristics, functional gene annotations, confidence scores, etc. Then, select an appropriate neural network model (such as a multilayer perceptron, convolutional neural network, etc.) and use this feature data as input layer nodes for model training and verification.

[0132] During the training process, cross-validation is used to evaluate the model's performance and select the optimal network structure and parameters. The trained model is used to predict the matching degree between the donor and the acceptor. The distance between the donor and the acceptor can be calculated using the weighted UniFrac distance:

[0133] Among them, p i,donor and p i,recipient represents the relative abundance of the i-th microorganism in the donor and recipient, respectively, w i represents the weight of the i-th microorganism.

[0134] Based on the distance calculation results and confidence scores, a neural network model is used to predict the optimal donor-recipient pairing. The model's output layer provides a matching score for each donor-recipient pair, which is used to assess their compatibility.

[0135] Example 2 Verification of model application results

[0136] Verification of model application results

[0137] To verify the effectiveness of the model, the inventors selected several patients with gynecological diseases in actual clinical applications, matched donors with recipients based on the best match predicted by the model, and performed intestinal bacteria transplantation treatment. The specific steps are as follows:

[0138] 1. Select clinical samples: Based on the model prediction results, select appropriate donor-recipient pairs for intestinal bacterial transplantation treatment. Record the patient's clinical data before treatment, including symptoms, hormone levels, metabolic parameters, BMI index, follicle development, menstrual cycle status, etc. The specific situation is as follows:

[0139] PCOS patient A:

[0140] Age: 28

[0141] Symptoms: Irregular menstruation, weight gain, hirsutism

[0142] Hormone levels: FSH 5.2 IU / L, LH 10.8 IU / L, E2 45 pg / mL, T 65 ng / dL

[0143] Metabolic parameters: fasting blood glucose 5.6 mmol / L, HOMA-IR 3.5

[0144] BMI: 29

[0145] Follicular development: multiple small follicles, increased ovarian size

[0146] Menstrual cycle: irregular, with intervals ranging from 35 to 60 days

[0147] Ovarian cyst patient B:

[0148] Age: 34

[0149] Symptoms: Lower abdominal pain, irregular menstruation Hormone levels: FSH 6.5 IU / L, LH 12.0 IU / L, E2 50 pg / mL, T 60 ng / dL Metabolic parameters: Fasting blood glucose 5.7 mmol / L BMI: 26

[0150] Follicular development: There is a 5cm cyst on the right ovary

[0151] Menstrual cycle: irregular, with an interval of 40-55 days

[0152] Patient C with intrauterine adhesion syndrome:

[0153] Age: 30

[0154] Symptoms: Oligomenorrhea, infertility Hormone levels: FSH 7.0 IU / L, LH 11.5 IU / L, E2 48 pg / mL, T 55 ng / dL Metabolic parameters: Fasting blood glucose 5.8 mmol / L BMI: 25

[0155] Uterine development: uterine cavity adhesions

[0156] Menstrual cycle status: Reduced menstruation

[0157] Patients with premature ovarian failure D:

[0158] Age: 29

[0159] Symptoms: Irregular periods, hot flashes, night sweats Hormone levels: FSH 40.0 IU / L, LH 25.0 IU / L, E2 20 pg / mL, T 50 ng / dL Metabolic parameters: Fasting blood glucose 5.5 mmol / L BMI: 24

[0160] Ovarian development: Ovarian size decreases

[0161] Menstrual cycle status: Oligomenorrhea

[0162] Hyperprolactinemia patients E:

[0163] Age: 33

[0164] Symptoms: Lactation, infertility, irregular menstruation Hormone levels: FSH 5.0IU / L, LH 10.0IU / L, E2 55pg / mL, PRL 80ng / mL Metabolic parameters: Fasting blood glucose 5.9mmol / L

[0165] BMI: 27

[0166] Menstrual cycle: irregular, with an interval of 35-50 days

[0167] Pelvic inflammatory disease patient F:

[0168] Age: 31

[0169] Symptoms: Lower abdominal pain, fever, abnormal vaginal discharge

[0170] Hormone levels: FSH 6.0 IU / L, LH 10.5 IU / L, E2 60 pg / mL, T 55 ng / dL

[0171] Metabolic parameters: fasting blood glucose 5.6mmol / L

[0172] BMI: 26

[0173] Menstrual cycle: Normal

[0174] Menopausal syndrome patients G:

[0175] Age: 50

[0176] Symptoms: Hot flashes, night sweats, palpitations

[0177] Hormone levels: FSH 50.0 IU / L, LH 30.0 IU / L, E2 15 pg / mL, T 40 ng / dL

[0178] Metabolic parameters: fasting blood glucose 5.5mmol / L

[0179] BMI: 23

[0180] Menstrual cycle status: Amenorrhea

[0181] Patient H of early gynecological cancer:

[0182] Age: 45

[0183] Symptoms: irregular bleeding, lower abdominal pain

[0184] Hormone levels: FSH 7.5 IU / L, LH 12.5 IU / L, E2 35 pg / mL, T 50 ng / dL

[0185] Metabolic parameters: fasting blood glucose 5.7mmol / L

[0186] BMI: 24

[0187] Diagnosis: Early-stage cervical cancer

[0188] Menstrual cycle: irregular cycle

[0189] 2. Intestinal bacteria transplantation therapy: Intestinal bacteria transplantation is performed according to the donor-recipient pairing recommended by the model. The specific operation process is as follows:

[0190] a. Donor screening: A comprehensive health check is conducted on healthy donors to ensure that their intestinal flora is healthy and free of potential pathogens.

[0191] b. Microbial flora extraction: Fecal samples are collected from the donor and a high-concentration intestinal flora suspension is obtained through microbial flora extraction techniques (such as centrifugation and filtration).

[0192] c. Microbiota transplantation: The extracted intestinal microbiota suspension is transplanted into the recipient through oral capsules, gastrointestinal endoscopic injection or enema.

[0193] 3. Regular follow-up: Follow up patients 1 month, 3 months, 6 months, and 12 months after intestinal bacteria transplantation treatment to record changes in clinical data after treatment. Follow-up content includes:

[0194] a. Symptom assessment: Record changes in the patient's subjective symptoms, such as abdominal pain, abdominal distension, irregular menstruation, hirsutism, acne, etc.

[0195] b. Hormone level determination: Detect changes in sex hormones (such as testosterone, estradiol, luteinizing hormone (LH), and follicle-stimulating hormone (FSH)) in serum.

[0196] c. Metabolic parameter assessment: Measure metabolic indicators such as fasting blood glucose, insulin, and blood lipid levels to evaluate improvements in metabolic status.

[0197] d. BMI index monitoring: record weight and height, calculate BMI index, and evaluate weight changes.

[0198] e. Follicular development and menstrual cycle monitoring: Ultrasound examination and menstrual cycle records are used to assess follicular development and the regularity of the menstrual cycle.

[0199] The specific changes are as follows:

[0200] PCOS patient A: Menstrual cycle returns to 28-32 days

[0201] Weight loss of 3 kg

[0202] Improved hormone levels: FSH 5.0 IU / L, LH 8.5 IU / L, E2 55 pg / mL, T 50 ng / dL

[0203] Fasting blood glucose dropped to 5.2mmol / L and HOMA-IR dropped to 2.8

[0204] Improved follicular development: decreased number of follicles and smaller ovarian size

[0205] Ovarian cyst patient B: menstrual cycle returned to 30-35 days

[0206] The cyst volume decreased to 3 cm

[0207] Improved hormone levels: FSH 6.0 IU / L, LH 10.5 IU / L, E2 60 pg / mL, T 50 ng / dL

[0208] Patient C with intrauterine adhesion syndrome: menstruation returned to normal flow

[0209] Improved hormone levels: FSH 6.5 IU / L, LH 10.0 IU / L, E2 55 pg / mL, T 45 ng / dL

[0210] Patient D of premature ovarian failure: hot flashes were reduced and sleep quality improved

[0211] Improved hormone levels: FSH 30.0 IU / L, LH 20.0 IU / L, E2 35 pg / mL, T 45 ng / dL

[0212] Hyperprolactinemia patient E: menstrual cycle returns to 28-30 days

[0213] PRL levels decreased to 40 ng / mL

[0214] Improved hormone levels: FSH 5.0 IU / L, LH 9.0 IU / L, E2 60 pg / mL, T 45 ng / dL

[0215] Pelvic inflammatory disease patient F: Lower abdominal pain relieved, leucorrhea normal

[0216] Hormone levels remain normal: FSH 6.0 IU / L, LH 10.0 IU / L, E2 60 pg / mL, T 50 ng / dL

[0217] Menopausal syndrome patients G: Hot flashes and night sweats reduced

[0218] Hormone levels: FSH 45.0 IU / L, LH 25.0 IU / L, E2 20 pg / mL, T 35 ng / dL

[0219] Patient H of early-stage gynecological cancer: irregular bleeding decreased and lower abdominal pain was relieved

[0220] Hormone levels: FSH 7.0 IU / L, LH 12.0 IU / L, E2 40 pg / mL, T 45 ng / dL

[0221] 4. Efficacy evaluation: Compare the clinical data of patients before and after treatment to comprehensively evaluate the treatment effect. The following statistical analysis methods are used:

[0222] a. Paired sample t-test or Wilcoxon signed-rank test: evaluate the significant difference in data before and after treatment and verify the statistical significance of the treatment effect.

[0223] b. Clinical symptom score: Based on the patient's subjective symptoms and the doctor's objective examination, a comprehensive symptom score is given to evaluate symptom improvement.

[0224] c. Analysis of hormone level changes: Compare the changes in sex hormone levels before and after treatment to evaluate the endocrine regulation effect.

[0225] d. Metabolic parameter improvement analysis: Evaluate changes in metabolic indicators before and after treatment to verify the improvement in metabolic status.

[0226] 5. Side effect monitoring: Throughout the treatment and follow-up process, closely monitor the patient's side effects, including but not limited to:

[0227] a. Immune rejection reaction: Observe whether the patient experiences adverse reactions such as allergies, fever, abdominal pain, diarrhea, etc.

[0228] b. Intestinal discomfort: Record the patient's gastrointestinal reactions, such as abdominal distension, abdominal pain, constipation, diarrhea, etc.

[0229] c. Other potential side effects: Monitor patients for other abnormal symptoms or signs, and promptly address and record them.

[0230] 6. Statistical analysis and results presentation: Conduct detailed statistical analysis of all collected data to verify the efficacy of the donor-recipient pairings calculated using the model in actual applications. Present the results in the following ways:

[0231] a. Data visualization: Use the R language ggplot2 package to generate data visualization charts, such as heat maps, stacked bar charts, and scatter plots, to intuitively display the results of alpha diversity, beta diversity, and differential bacterial communities.

[0232] b. Clinical case presentation: Select representative clinical cases, describe in detail the changes before and after treatment, and demonstrate actual examples of the model application effect.

[0233] The above validation steps demonstrate that the donor-recipient distance and matching degree calculated using this model have indeed yielded effective treatment results in practice, without significant rejection or other side effects. This further validates the scientific nature and clinical practicality of the donor-recipient matching model, providing strong support for enterobacterial transplantation treatments for gynecological diseases.

[0234] Clinical results:

[0235] Patient A's matching score: 0.85

[0236] Before treatment: HOMA-IR 3.5, E2 45 pg / mL

[0237] After treatment: HOMA-IR 2.8, E2 40 pg / mL

[0238] Patient B's matching score: 0.86

[0239] Before treatment: ovarian cyst 5cm

[0240] After treatment: ovarian cyst 3cm

[0241] Patient C's matching score: 0.84

[0242] Before treatment: decreased menstruation

[0243] After treatment: normal menstrual flow

[0244] Patient D's matching score: 0.78 Before treatment: hot flashes, night sweats After treatment: symptom relief, improved hormone levels Patient E's matching score: 0.83 Before treatment: high prolactin, irregular menstruation After treatment: decreased PRL levels, restoration of menstrual cycle Patient F's matching score: 0.81 Before treatment: lower abdominal pain, abnormal leucorrhea After treatment: symptom relief, normal leucorrhea

[0245] Patient G's matching score: 0.77 Before treatment: hot flashes, night sweats, palpitations After treatment: symptoms reduced, hormone levels improved Patient H's matching score: 0.79 Before treatment: irregular bleeding, lower abdominal pain After treatment: symptoms relieved, hormone levels maintained normal.

Claims

1. A method for constructing a donor-recipient matching model for intestinal flora transplantation to treat gynecological diseases, characterized in that The method includes the following steps: collecting intestinal microbial data from clinical patient recipients and healthy donors to establish a database; performing metagenomic sequencing analysis based on the same parameters to obtain detailed intestinal microbial composition and functional gene information; performing meta-analysis, organizing and standardizing the data, and evaluating the species richness and uniformity within each sample by obtaining the alpha diversity, beta diversity, genus composition characteristics, and confidence scores of the intestinal microbial community; using the NetMoss network model, using the metagenomic data obtained from the meta-analysis, constructing a network diagram based on the co-occurrence frequency and correlation between microorganisms, and screening for significantly different genus characteristics between patients and healthy donors, including constructing a microbial collinearity matrix, constructing a collinearity network, screening for differential genus characteristics, and constructing a genus characteristic model; organizing the collected data to select an appropriate neural network model, using these characteristic data as input layer nodes for model training and verification, and finally constructing a donor-recipient matching model; the trained model is used to predict the matching degree of the donor and recipient, and the distance between the donor and recipient is calculated using the weighted UniFrac distance: ; Among them, p i,donor and p i,recipient represents the relative abundance of the i-th microorganism in the donor and recipient, respectively, w i represents the weight of the i-th microorganism; The confidence score of each donor and recipient sample is calculated to evaluate the degree of matching with the target bacterial genus. The confidence score is calculated as follows: ; Among them, p i,pbserved represents the observed relative abundance of the i-th microorganism, p i,expected represents the expected relative abundance; Based on the distance calculation results and combined with the confidence score, the optimal donor-acceptor pairing is predicted by a neural network model.

2. The construction method according to claim 1, wherein constructing a microbial co-occurrence matrix refers to calculating the co-occurrence relationship between each pair of bacterial genera based on metagenomic data, and using the Spearman rank correlation coefficient to calculate the co-occurrence relationship matrix between bacterial genera: ; in, p ij represents the Spearman rank correlation coefficient between bacterial genera i and j, cov(X i , X j ) represents the covariance of bacterial genera i and j, σx i and σx j represent the standard deviation of bacterial genus i and j, respectively.

3. The construction method according to claim 1, wherein the specific steps of screening for differential bacterial genera are as follows: Calculate the weighted degree centrality of two sets of networks: C w (v)=∑ u∈N(v )w(v,u); in, C w (v) represents the weighted degree centrality of node v, N(v) represents the set of nodes adjacent to node v, and w(v, u) represents the edge weight between nodes v and u; The weighted degree centrality differences of bacterial genera in the patient and healthy donor networks were compared to screen out bacterial genera that were significantly upregulated or downregulated in patients.

4. The construction method according to claim 1, wherein constructing a co-occurrence network refers to converting a co-occurrence relationship matrix into a network graph, wherein nodes represent bacterial genera and edges represent co-occurrence relationships between bacterial genera; by setting a threshold, significant co-occurrence relationships are screened out; constructing a genus feature model refers to incorporating the screened differential genus features into the model, using a machine learning algorithm to perform modeling, and predicting the donor-recipient matching degree.

5. The construction method according to claim 1, wherein the meta-analysis includes calculating alpha diversity, using indicators such as the Shannon index, Chao1 index, and Simpson index to evaluate species richness and evenness within each sample; calculating beta diversity, and evaluating microbial community differences between different samples using methods such as Bray-Curtis distance, Jaccard index, or weighted UniFrac distance; then, performing functional gene annotation and differential function analysis, using PICRUSt2 to predict functional genes and metabolic pathways, and comparing the differences in functional gene abundance between different groups.

6. The construction method according to claim 5, wherein the beta diversity calculation method is: Bray-Curtis distance formula is: ; in, X ik and X jk represents the abundance of the kth microorganism in the i-th and j-th samples, and n represents the total number of species in the samples.

7. The construction method according to claim 5, wherein the alpha diversity calculation method is: The Shannon diversity index calculation formula is: ; in, p i represents the relative abundance of the i-th microorganism, and S represents the total number of species in the sample; The Chao1 richness estimation formula is: ; Among them, S obs represents the number of species observed, F1 represents the number of species that appear only once in a single sample, and F2 represents the number of species that appear twice in a single sample; The Simpson index calculation formula is: ; Among them, p i represents the relative abundance of the i-th microorganism, and S represents the total number of species in the sample.

8. The construction method according to claim 1, wherein the confidence score of the meta-analysis is determined by evaluating the relative abundance and distribution of the target bacterial genus in each sample, and the confidence score calculation formula is: ; in, p i,observed represents the observed relative abundance of the i-th microorganism, p i,expected Indicates the expected relative abundance.

9. The intestinal flora transplantation donor-recipient matching model prepared by the construction method according to claims 1-8.

10. Use of the donor-recipient matching model according to claim 9 in the preparation of a drug for treating gynecological diseases that can be treated by intestinal flora transplantation.

11. The use according to claim 10, wherein the gynecological disease is ovarian cyst, polycystic ovary syndrome, intrauterine adhesion syndrome, premature ovarian failure, hyperprolactinemia, pelvic inflammatory disease, menopausal syndrome, early cancer, low-grade cancer; The early-stage cancer and low-grade cancer include low-grade ovarian cancer and early-stage endometrial cancer.

Citation Information

Patent Citations

  • Metagenome 16S rRNA high-throughput sequencing data processing and analysis process control method

    CN105279391A

  • Method for constructing donor-receptor matching model for treating irritable bowel syndrome through intestinal flora transplantation

    CN117393170A