IBS micro-ecological transplantation intelligent prediction method and system based on multi-omics driving

Through the joint modeling of multi-omics data fusion and transfer learning framework, the problem of low accuracy in predicting the efficacy of IBS microecological transplantation was solved, personalized and dynamic efficacy prediction was achieved, the prediction efficiency and accuracy were improved, and data security was guaranteed.

CN120636552AActive Publication Date: 2025-09-12THE FIRST MEDICAL CENT CHINESE PLA GENERAL HOSPITAL

Patent Information

Application Number
CN202510576783.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-06
Publication Date
2025-09-12
Estimated Expiration
2045-05-06

AI Technical Summary

Technical Problem

Existing IBS diagnosis and treatment evaluation methods lack in-depth analysis of intestinal microecology and host intrinsic factors, resulting in low accuracy in predicting the efficacy of microecological transplantation, making it difficult to accurately judge patient suitability and efficacy, and limiting the application of microecological transplantation in the treatment of IBS.

Method used

By establishing a multi-omics data fusion subsystem, collecting metagenomic, metabolomic, host genome and clinical phenotype data, constructing a microbiome-metabolite joint network, generating a host-microbiome interaction matrix, and combining it with a dynamic response algorithm, using a transfer learning framework for joint modeling, and outputting the prediction results of the efficacy of microecological transplantation.

Benefits of technology

It achieves personalized, dynamic and accurate prediction of the efficacy of microecological transplantation, improves prediction efficiency and accuracy, reduces dependence on sample size, enhances the generalization ability of the model, and ensures data security and privacy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636552A_ABST
    Figure CN120636552A_ABST
Patent Text Reader

Abstract

The invention provides an IBS micro-ecological transplantation intelligent prediction method and system based on multi-omics driving, and relates to the technical field of biomedicine. The method comprises the following steps: establishing a multi-omics data fusion subsystem to collect metagenome, metabolome, host genome and clinical phenotype group data of a target patient; inputting the data into a flora-metabolite combined network analysis model to construct an interaction network and extracting features; generating an incidence matrix based on the features and the host genome data and calculating indexes; generating indexes through a dynamic response algorithm in combination with the clinical phenotypic data and the indexes; and outputting a curative effect prediction result by using a transfer learning framework combined with modeling. The system comprises a data acquisition module, a network analysis module, a correlation calculation module, a dynamic response module and a joint modeling module. According to the method, multiple omics data are integrated, the flora and host relation is accurately mined, intelligent prediction of the micro-ecological transplantation curative effect is achieved, powerful support is provided for IBS personalized treatment, and meanwhile data processing and safety guarantee measures are taken.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of biomedical technology, and in particular to a multi-omics-driven intelligent prediction method and system for IBS microecological transplantation. Background Art

[0002] Irritable bowel syndrome (IBS) is a common functional gastrointestinal disorder with a high prevalence worldwide, severely impacting patients' quality of life and socioeconomic burden. While the pathogenesis of IBS remains unclear, it is generally believed to be closely related to multiple factors, including intestinal microecological imbalance, immune dysfunction, impaired intestinal barrier function, and host genetic factors.

[0003] Among the treatments for IBS, microbiome transplantation (such as fecal microbiota transplantation) has shown certain therapeutic potential. By transplanting the intestinal microbial community of a healthy donor into the patient's body, it is expected to regulate the patient's intestinal microecological balance, improve intestinal function, and relieve IBS symptoms. However, the efficacy of microbiome transplantation varies significantly between different patients, and not all patients can benefit from it. This uncertainty in efficacy poses a huge challenge to clinicians when choosing treatment options. It is difficult to accurately judge which patients are more suitable for microbiome transplantation treatment, and it is impossible to predict the treatment effect in advance, which limits the widespread application of microbiome transplantation in the treatment of IBS.

[0004] Existing methods for the diagnosis and treatment assessment of IBS have many limitations. Traditional diagnosis relies primarily on patient descriptions of symptoms and simple clinical examinations, lacking in-depth analysis of the intestinal microbiome and intrinsic host factors. In evaluating the efficacy of microbiome transplantation, subjective indicators such as clinical symptom scores are currently used, lacking objective and accurate predictive methods. Although some studies have attempted to predict efficacy by detecting a single microbial indicator or host factor, due to the complexity of the pathogenesis of IBS, a single indicator is unlikely to fully reflect the interaction between the microbiome and the host, resulting in low predictive accuracy.

[0005] With the rapid development of biotechnology, multi-omics technologies, including metagenomics, metabolomics, host genomics, and clinical phenomics, have provided new insights into the pathogenesis and treatment strategies of IBS. Metagenomics can comprehensively analyze the composition and function of the intestinal microbiome; metabolomics can detect changes in metabolites within an organism, reflecting the body's metabolic state; host genomics helps to reveal the influence of host genetic factors on disease susceptibility and treatment response; and clinical phenomics integrates information such as a patient's clinical symptoms, signs, and treatment history. However, how to effectively integrate these multi-omics data and apply them to the intelligent prediction of the efficacy of IBS microecological transplantation remains an urgent challenge in this field.

[0006] In terms of data processing and analysis, multi-omics data are characterized by large volumes, high dimensionality, and high complexity. Traditional data processing methods struggle to uncover the key information hidden within them. Furthermore, the complex interactions between different omics data sets make constructing rational models to reveal these relationships and leverage them for accurate predictions a major challenge in current research. Summary of the Invention

[0007] The purpose of the present invention is to provide a multi-omics-driven IBS microecological transplantation intelligent prediction method and system to solve the problems raised in the above background technology.

[0008] To achieve the above objectives, the present invention provides the following technical solution: a multi-omics-driven intelligent prediction method for IBS microecological transplantation, the method comprising:

[0009] Establishing a multi-omics data fusion subsystem to collect four-dimensional data of target patients, wherein the four-dimensional data includes metagenomic data, metabolomic data, host genome data and clinical phenotype data;

[0010] Inputting the four-dimensional data into a preset microbiome-metabolite joint network analysis model to construct an interaction network between microbiome and metabolites and extract key network topological features;

[0011] Generate a host-microbiome interaction matrix based on the key network topology features and host genome data, and calculate the functional gene synergy index;

[0012] Combining the clinical phenotype group data with the functional gene synergy index, generating a symptom-microbe dynamic response index through a dynamic response algorithm;

[0013] The interaction network, the host-microbiome interaction matrix and the symptom-microbe dynamic response index are jointly modeled using a preset transfer learning framework to output the prediction results of the microecological transplantation efficacy.

[0014] Optionally, the steps of establishing a multi-omics data fusion subsystem to collect four-dimensional data of a target patient include:

[0015] Obtain metagenomic data of target patients and annotate bacterial flora composition and functional gene expression levels;

[0016] Collect metabolomic data to extract concentrations of various short-chain fatty acids, bile acids, tryptophan, and metabolic pathway activity indicators;

[0017] Analyze host genome data and mark single nucleotide polymorphism sites and epigenetic modification regions;

[0018] The clinical phenotype group data, including symptom rating scale results and historical treatment records, are integrated, and the metagenomic data, the metabolomic data, and the host genome data are aligned in time series.

[0019] Optionally, the microbiome-metabolite joint network analysis model includes:

[0020] A microbial co-occurrence network was constructed based on the bacterial abundance in metagenomic data, where nodes represent bacterial taxa and edge weights represent the strength of symbiotic or competitive relationships between species.

[0021] Mapping the metabolite concentrations in the metabolomics data to the microbial co-occurrence network to generate a metabolite-microbial community association subnetwork;

[0022] The associated subnetwork is traversed by a random walk algorithm to identify the microbiome-metabolite core module with bidirectional regulatory effects.

[0023] Optionally, the step of generating a host-microbiome interaction matrix includes:

[0024] Extract immune regulatory genes and metabolism-related genes from host genome data as candidate interacting gene sets;

[0025] Calculate the Pearson correlation coefficient between the expression level of each gene in the candidate interaction gene set and the abundance of the microbial community-metabolite core module;

[0026] Gene-microbial community pairs whose absolute value of correlation coefficient exceeds the first threshold are screened, and a multidimensional correlation matrix with genes as rows and microbial communities as columns is constructed.

[0027] Optionally, the dynamic response algorithm is implemented as follows:

[0028] Dynamic response periods were divided according to symptom severity in clinical phenotype group data;

[0029] In each period, the lagged correlation between the change rate of the microbial abundance of the microbial community-metabolite core module and the symptom score was calculated;

[0030] The optimal hysteresis window was determined based on the maximum information coefficient, and the dynamic response weights of the microbiome to symptoms in each time period were calculated.

[0031] Optionally, the steps of constructing the transfer learning framework include:

[0032] Pre-training basic prediction models, whose input is the multi-omics data of healthy donors and the corresponding microecological transplantation efficacy labels;

[0033] Freezing the feature extraction layer of the basic prediction model and adding an adaptive parameter adjustment layer;

[0034] The joint modeling data of the target patients are input into the adjusted model, and the prediction results are optimized by the domain adaptation loss function.

[0035] Optionally, the method further includes:

[0036] Performing a credibility assessment on the efficacy prediction results by calculating the confidence distribution of the prediction results in the feature space in the transfer learning framework;

[0037] If the confidence level is lower than the second threshold, the incremental learning mechanism is triggered, the current prediction data is added to the training set and the model parameters are re-optimized.

[0038] Optionally, the method further comprises a data preprocessing step:

[0039] The metagenomic data were filtered for low-abundance bacterial groups, and bacterial genera with relative abundance exceeding the third threshold were retained;

[0040] Metabolomics data were corrected for batch effects, and a linear regression model based on quality control samples was used to eliminate instrument detection bias;

[0041] Linkage disequilibrium analysis was performed on the host genome data to remove redundant single nucleotide polymorphism sites.

[0042] Optionally, the method further includes:

[0043] Establish a secure storage protocol for multi-omics data, where data of different dimensions are encrypted with different levels of privacy;

[0044] During the data transmission process, dynamic key sharding technology is used to transmit encrypted data in segments and perform real-time verification.

[0045] The present invention also provides an IBS microecological transplantation intelligent prediction system based on multi-omics drive, the system comprising:

[0046] A data acquisition module is used to collect four-dimensional data of the target patient through a multi-omics data fusion subsystem, wherein the four-dimensional data includes metagenomic data, metabolomic data, host genome data and clinical phenotype group data;

[0047] A network analysis module is used to input the four-dimensional data into a preset microbiome-metabolite joint network analysis model to construct an interaction network between microbiome and metabolites and extract key network topological features;

[0048] An association calculation module is used to generate a host-microbiome interaction association matrix based on the key network topology features and host genome data, and calculate the functional gene synergy index;

[0049] A dynamic response module, configured to combine the clinical phenotype group data with the functional gene synergy index and generate a symptom-microbe dynamic response index through a dynamic response algorithm;

[0050] The joint modeling module is used to jointly model the interaction network, the host-microbiome interaction matrix and the symptom-microorganism dynamic response index using a preset transfer learning framework, and output the prediction results of the microecological transplantation efficacy.

[0051] Compared with the prior art, the present invention has the following beneficial effects:

[0052] From the perspective of data collection and integration, by establishing a multi-omics data fusion subsystem, we comprehensively collect metagenomic data, metabolomic data, host genome data, and clinical phenotype data from target patients and align them in time series. This multi-dimensional data integration method can comprehensively and dynamically reflect the interaction between the patient's intestinal microecology system and the host. Compared with traditional analysis methods that rely only on a single data type, this greatly enriches the source of information, avoids analytical bias caused by missing information, and lays a solid data foundation for subsequent accurate predictions.

[0053] In terms of analytical model construction, the pre-defined microbiome-metabolite joint network analysis model, by constructing microbial co-occurrence networks, generating metabolite-microbiome association subnetworks, and identifying core modules, can deeply explore the complex interactions between microbiota and metabolites and extract key network topological features. This helps to reveal the intrinsic connection between microbiota and metabolism in the pathogenesis of IBS, providing a deeper perspective for understanding the development and progression of the disease and providing more biologically meaningful indicators for predicting the efficacy of microbiome transplantation.

[0054] Generating a host-microbiome interaction matrix and calculating functional gene synergy indices further integrates host genomic data with microbiome information, analyzing the interactions between host and microbiome at the genetic level. This analytical approach considers the influence of host genetic factors on the intestinal microbiome, enabling a more precise assessment of the potential impact of individual differences on the efficacy of microbiome transplantation, making predictions more personalized and targeted.

[0055] Combining clinical phenotype data with functional gene synergy indicators, a dynamic response algorithm was used to generate a symptom-microbe dynamic response index, fully accounting for the dynamic relationship between disease symptoms and microbes. This index can reflect the impact of microbes on symptoms in real time over different time periods, capturing subtle changes between the microbiome and clinical symptoms. Compared to traditional static analysis methods, it better reflects the dynamic development of the disease, providing a more timely and accurate basis for efficacy prediction.

[0056] Using a pre-defined transfer learning framework to jointly model multi-omics data not only fully leverages the multi-omics data of healthy donors and the corresponding microbiome transplant efficacy labels for pre-training, but also, through an adaptive parameter adjustment layer and domain adaptation loss function, better adapts the model to the data characteristics of the target patients and optimizes prediction results. This transfer learning approach improves the model's generalization ability, reduces its reliance on a large number of target patient samples, and enables accurate predictions even with limited sample sizes, significantly improving prediction efficiency and accuracy.

[0057] Furthermore, the credibility of efficacy predictions is assessed, and when confidence is low, an incremental learning mechanism is triggered to continuously optimize the model. By adding new prediction data to the training set and re-optimizing the model parameters, the model can learn more diverse data features, gradually improving its predictive capabilities for different patients and further enhancing the reliability of predictions.

[0058] The present invention implements a series of effective measures for data processing and security. Data preprocessing steps, such as filtering low-abundance bacterial communities in metagenomic data, correcting for batch effects in metabolomics data, and performing linkage disequilibrium analysis on host genome data, improve data quality and reduce the interference of noise and redundant information on prediction results. Furthermore, the established multi-omics data security storage protocol and the dynamic key sharding technology employed ensure data security during storage and transmission, protecting patient privacy and providing a safe and reliable environment for the clinical application of multi-omics data. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 This is a working principle diagram of the multi-omics-driven IBS microecological transplantation intelligent prediction method of the present invention;

[0060] Figure 2 This is the workflow diagram of the microbiome-metabolite joint network analysis model;

[0061] Figure 3 Workflow diagram for efficacy prediction outcome evaluation and model optimization;

[0062] Figure 4 Workflow diagram for data preprocessing. DETAILED DESCRIPTION

[0063] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0064] See also Figures 1-4 The present invention relates to a multi-omics-driven intelligent prediction method and system for IBS microecological transplantation, and its specific implementation methods will be described in detail below.

[0065] A multi-omics data fusion subsystem was established to collect four-dimensional data from target patients, including metagenomic data, metabolomics data, host genome data, and clinical phenotype data. This subsystem integrates data acquired from different dimensions to provide a comprehensive information foundation for subsequent analysis.

[0066] The four-dimensional data is input into a pre-defined microbiome-metabolite joint network analysis model to construct a microbiome-metabolite interaction network and extract key network topological features. This step uses a specific model to explore the intrinsic connections between microbiome and metabolites, identifying network features that are important for subsequent analysis.

[0067] Based on these key network topology features and host genomic data, a host-microbiome interaction matrix was generated, and functional gene synergy indices were calculated. This approach allows for analysis of the interactions between host genes and microbiome, providing a deeper understanding of the microbial ecosystem.

[0068] Combining the clinical phenotype data with the functional gene synergy index, a dynamic response algorithm is used to generate a symptom-microbe dynamic response index. This index reflects the dynamic association between microbes and patient symptoms, providing a more accurate basis for prediction.

[0069] Using a pre-defined transfer learning framework, the interaction network, the host-microbiome interaction matrix, and the symptom-microbe dynamic response index are jointly modeled to output predictions of the efficacy of microecological transplantation. This transfer learning framework leverages existing data and models to improve the accuracy and reliability of predictions.

[0070] Other technical features of the present invention are described in detail below through specific embodiments.

[0071] Example 1:

[0072] Establishing a multi-omics data fusion subsystem to collect four-dimensional data from target patients specifically involves obtaining metagenomic data from the target patient and annotating the microbial composition and functional gene expression levels. Specifically, specialized gene sequencing equipment is used to perform metagenomic sequencing on the patient's stool sample to obtain raw sequencing data. Bioinformatics analysis tools such as MetaPhlAn and HUMAnN are used to process the sequencing data to determine the species and relative abundance of various microbial communities in the sample, thereby annotating the microbial composition. These tools are also used to quantitatively analyze the expression of functional genes, identify genes associated with microbial function, and record their expression levels.

[0073] Collect metabolomic data to extract concentrations of various short-chain fatty acids, bile acids, tryptophan, and metabolic pathway activity indicators. Collect metabolomic mass spectrometry data from the patient's blood, feces, or other appropriate biological samples. Use metabolomics data processing software such as XCMS and MZmine to perform peak identification, alignment, and quantitative analysis on the mass spectrometry data. Various short-chain fatty acids, bile acids, or tryptophan are important products of intestinal microbial metabolism and are closely related to the pathogenesis of IBS. Through specific detection methods and data analysis methods, accurately measure the concentrations of various short-chain fatty acids, bile acids, or tryptophan in the sample. For metabolic pathway activity indicators, by mapping metabolites to known metabolic pathway databases such as KEGG and Reactome, pathway enrichment analysis algorithms are used to calculate the activity levels of various metabolic pathways to reflect changes in metabolic function.

[0074] Analyze host genome data and mark single nucleotide polymorphism sites and epigenetic modification regions. Use high-throughput sequencing technology to sequence the host genome and obtain massive amounts of gene sequence information. Use SNP calling tools such as GATK and SAMtools to analyze sequencing data and identify single nucleotide polymorphism (SNP) sites. Changes in these sites may affect the host's susceptibility to disease and its interaction with microorganisms. At the same time, use epigenetic research techniques such as chromatin immunoprecipitation sequencing (ChIP-seq) and whole-genome bisulfite sequencing (WGBS) to determine epigenetic modification regions, including DNA methylation and histone modification. These modifications can regulate gene expression, thereby affecting host-microorganism interactions and the occurrence and development of diseases.

[0075] Integrate clinical phenotype group data, including symptom score scale results and historical treatment records, and align the metagenomic data, metabolomic data, and host genome data in time series. Collect clinical phenotypic information of patients, such as using the IBS Symptom Severity Rating Scale (IBS-SSS) to quantitatively score patients' symptoms such as abdominal pain, bloating, and changes in bowel habits. At the same time, organize patients' historical treatment records, including used medications, treatment time, treatment effects, etc. In order to facilitate comprehensive analysis, the metagenomic data, metabolomic data, and host genome data collected at different time points are aligned in time order to ensure the temporal consistency of the data, so that subsequent analysis can accurately reflect the dynamic changes of the disease.

[0076] Example 2:

[0077] The construction process of the microbial community-metabolite joint network analysis model is as follows: a microbial co-occurrence network is constructed based on the microbial community abundance in the metagenomic data. In this network, each bacterial genus taxonomic unit is regarded as a node, and the edge weight is determined by calculating the correlation between different bacterial genera. If two bacterial genera often appear together in the sample, it means that there is a symbiotic relationship between them, and the edge weight is positive; conversely, if two bacterial genera rarely appear together, there may be a competitive relationship, and the edge weight is negative. The specific calculation method can use the Spearman correlation coefficient or the Pearson correlation coefficient, the formula is:

[0078]

[0079] Among them, r represents the correlation coefficient, x i and y i Respectively represent the abundance of the two bacterial genera in the i-th sample, and are the average abundances of the two genera in all samples, and n is the number of samples. When |r| exceeds a certain threshold (such as 0.6), it is considered that there is a significant co-occurrence relationship between the two genera, and thus a connection is established in the network.

[0080] The metabolite concentrations in the metabolome data are mapped to the microbial co-occurrence network to generate a metabolite-microbial community association subnetwork. Metabolites are treated as new nodes, and connections between metabolite nodes and genus nodes are added to the microbial co-occurrence network based on the known biochemical relationship between metabolites and bacterial genera or the correlation obtained through data analysis. For example, if a certain genus of bacteria can produce a specific metabolite, or if the concentration changes of the two show a significant correlation, a corresponding edge is established in the network. In this way, an association subnetwork containing microbial community and metabolite information is constructed, which more comprehensively displays the interaction relationship in the microbial ecosystem.

[0081] The random walk algorithm is used to traverse the associated subnetwork to identify the microbiome-metabolite core modules with bidirectional regulatory effects. The random walk algorithm is an algorithm that performs random exploration on a graph structure. In the associated subnetwork, starting from a randomly selected node, the edges connected to it are selected for movement according to a certain probability, and this process is repeated continuously. During each movement, the nodes and edges passed through are recorded. After a large number of random walk steps, which nodes and edges are frequently visited are analyzed. The microbiome-metabolite core modules with bidirectional regulatory effects are usually those subnetwork parts that are frequently visited during the random walk process. In this way, the core modules that play a key role in the interaction between microbiome and metabolites can be found, providing an important basis for subsequent analysis.

[0082] Example 3:

[0083] When generating the host-microbiome interaction matrix, immune-regulatory and metabolic-related genes from the host genomic data were extracted as candidate interacting gene sets. Gene information related to immune regulation and metabolism was obtained from existing gene databases such as NCBI and Ensembl. Based on the pathogenesis of IBS and existing research findings, genes that may interact with the gut microbiome were screened and formed into a candidate interacting gene set. These genes may be involved in regulating the host immune response, influencing intestinal barrier function, or participating in the processing of microbial metabolites.

[0084] The Pearson correlation coefficient between the expression level of each gene in the candidate interaction gene set and the abundance of the microbial community-metabolite core module was calculated. The gene expression profile data and the abundance data of the microbial community-metabolite core module were used to calculate the Pearson correlation coefficient between the expression level of each gene and the abundance of each microbial community-metabolite core module. The calculation formula of the Pearson correlation coefficient is the same as the formula for calculating the correlation coefficient between bacterial genera in Example 2 (x in the formula is 0. i Indicates the expression level of the gene in the i-th sample, y i represents the abundance of the microbial core module in the i-th sample. This coefficient reflects the degree of linear correlation between gene expression and microbial abundance, ranging from -1 to 1. A positive value indicates a positive correlation, i.e., as gene expression levels increase, microbial abundance tends to increase; a negative value indicates a negative correlation, i.e., as gene expression levels increase, microbial abundance tends to decrease.

[0085] Gene-microbiome pairs whose absolute value of the correlation coefficient exceeds a first threshold (e.g., 0.5) are screened, and a multidimensional correlation matrix is ​​constructed with genes as rows and microbiome as columns. Based on the calculated Pearson correlation coefficient, gene-microbiome pairs with strong correlation are screened. These gene-microbiome pairs are organized into a multidimensional matrix, with the rows representing different genes and the columns representing different microbiome core modules. The elements in the matrix are the correlation coefficients of the corresponding gene-microbiome pairs, thus constructing a host-microbiome interaction correlation matrix, which intuitively shows the interaction relationship between host genes and microbiome.

[0086] Example 4:

[0087] The implementation method of the dynamic response algorithm is as follows: Divide the dynamic response period according to the symptom severity in the clinical phenome data. First, quantify the symptom severity. For example, according to the IBS-SSS score, the symptoms are divided into mild (75 < IBS-SSS score ≤ 175), moderate (175 < IBS-SSS score ≤ 300), and severe (IBS-SSS score > 300). Then, in chronological order, divide the patient's disease course into different periods, and each period corresponds to a different stage of symptom severity change. For example, when the symptoms change from mild to moderate, a new period is divided; after the symptoms remain stable at the moderate level for a period of time, if there are obvious changes, a new period is divided again. In this way, the relationship between microorganisms and symptoms can be analyzed more accurately according to the dynamic changes of symptoms.

[0088] Within each period, statistically analyze the lag correlation between the change rate of the microbiota-metabolite core module microbiota abundance and the symptom score. For each divided period, calculate the change rate of the core module microbiota abundance at different time points. Let t1 and t2 be two time points within the period, and the calculation formula for the change rate of the microbiota abundance is:

[0089]

[0090] At the same time, record the symptom scores at the corresponding time points. Then, calculate the lag correlation between the change rate of the microbiota abundance and the symptom score. The lag correlation refers to the correlation between the change of the microbiota abundance lagging behind the symptom change in time. For example, calculate the correlation between the change rate of the microbiota abundance at time t1 and the symptom score at time t1 + k (k is the lag time, k = 1, 2,...). In this way, it can be found whether there is a time delay in the response of microorganisms to symptom changes.

[0091] Determine the optimal lag window based on the maximum information coefficient, and calculate the dynamic response weight of the microbiota to the symptoms within each period. The maximum information coefficient (MIC) is an index to measure the complex correlation between two variables. After calculating the correlations between the change rate of the microbiota abundance and the symptom score at different lag times, use the MIC to find the lag time corresponding to the strongest correlation, and this lag time is the optimal lag window. After determining the optimal lag window, calculate the dynamic response weight of the microbiota to the symptoms within each period according to the correlation strength between the change rate of the microbiota abundance and the symptom score under the optimal lag window. The stronger the correlation, the greater the dynamic response weight, indicating that the impact of the microbiota on the symptoms is more significant. In this way, an index that can accurately reflect the symptom-microbiota dynamic response relationship is generated.

[0092] Example 5:

[0093] The steps for constructing the transfer learning framework are as follows: Pretrain a basic prediction model, whose input is the multi-omics data of healthy donors and the corresponding microbiome transplantation efficacy labels. Collect a large amount of metagenomic data, metabolomics data, and host genomic data from healthy donors, as well as the efficacy data after microbiome transplantation of these donors (the efficacy label can be expressed as indicators such as cure rate and symptom improvement). Select an appropriate machine learning model, such as a deep neural network (DNN) or random forest (RF), and train the model using the multi-omics data of these healthy donors as input and the efficacy label as output. During the training process, by adjusting the model parameters, the model can learn the relationship between the multi-omics data of healthy donors and the efficacy of microbiome transplantation, thereby obtaining a pretrained basic prediction model.

[0094] The feature extraction layer of the base prediction model is frozen, and an adaptive parameter adjustment layer is added. After pre-training, the base prediction model's feature extraction layer has already learned important features from the multi-omics data. To adapt the model to the target patient's data characteristics, the feature extraction layer is frozen, and its parameters are no longer updated. Then, an adaptive parameter adjustment layer is added to the model. This adjustment layer can fine-tune the model based on the target patient's data, for example by adding fully connected layers and convolutional layers. The parameters of the adjustment layer are then trained using the target patient's data, enabling the model to better adapt to the target patient's microbiome and clinical characteristics.

[0095] The joint modeling data of the target patient is input into the adjusted model, and the prediction results are optimized by the domain adaptation loss function. The joint modeling data of the target patient, such as the interaction network, association matrix and dynamic response index, are input into the model with the addition of the adaptive parameter adjustment layer. In order to make the model more accurate in predicting the target patient data, the model is optimized using the domain adaptation loss function. The role of the domain adaptation loss function is to minimize the difference between the source domain (healthy donor data) and the target domain (target patient data), so that the model can better predict the target patient data. For example, the maximum mean difference (MMD) can be used as the domain adaptation loss function, and its formula is:

[0096]

[0097] Among them, x and y represent the data of the source domain and target domain respectively, x i and y j are samples in the source and target domains, n and m are the number of samples in the source and target domains, φ is the function that maps samples to the Reproducing Kernel Hilbert Space (RKHS), By continuously adjusting the parameters of the adaptive parameter adjustment layer and minimizing the domain adaptation loss function, the model can optimize the prediction results of the microecological transplantation efficacy of the target patient.

[0098] Example 6:

[0099] The method of the present invention also includes a credibility assessment of the efficacy prediction results and related data processing and safety measures. The credibility assessment of the efficacy prediction results is performed by calculating the confidence distribution of the prediction results in the feature space in the transfer learning framework. In the transfer learning framework, the distribution of the prediction results in the feature space can be obtained by predicting the target patient data through the model. Some statistical methods, such as calculating the probability density function or confidence interval of the prediction results, are used to evaluate the credibility of the prediction results. For example, a Gaussian mixture model (GMM) can be used to model the prediction results and estimate the probability that the prediction results belong to different categories, thereby obtaining the confidence distribution of the prediction results.

[0100] If the confidence level is lower than the second threshold (such as 0.6), the incremental learning mechanism is triggered, the current prediction data is added to the training set and the model parameters are re-optimized. When the confidence level of the prediction result is lower than the set threshold, it means that the model has low reliability for the prediction result. In order to improve the accuracy of the model, the incremental learning mechanism is triggered. The multi-omics data of the current target patient and the corresponding prediction results (even if the prediction results may not be accurate) are added to the training set, and then the model is retrained and optimized. During the retraining process, the model can learn more sample information, especially data features similar to the target patient, thereby adjusting the parameters of the model and improving the model's predictive ability for similar patients.

[0101] In terms of data preprocessing, the metagenomic data is filtered for low-abundance bacterial communities, retaining bacterial genera with relative abundances exceeding a third threshold (e.g., 0.01%). Metagenomic data contain a large number of bacterial communities with extremely low relative abundances. These communities may be caused by experimental errors or environmental pollutants, have little impact on the analysis results, and increase the computational burden. By setting a relative abundance threshold, low-abundance bacterial communities are filtered out, retaining only biologically significant dominant bacterial communities, and improving the accuracy and efficiency of data analysis.

[0102] The metabolomics data were corrected for batch effects, and a linear regression model based on quality control samples was used to eliminate instrument detection bias. During the metabolomics data collection process, batch effects may occur due to differences in experimental conditions, instrument status, and other factors between different batches, affecting the accuracy of the data. Using quality control samples (the same standard samples were added to each batch of experiments), the metabolomics data were corrected using a linear regression model. Assume y ij is the metabolite measurement value of the i-th sample in the j-th batch, x ij is the corresponding covariate (such as batch number), the linear regression model can be expressed as:

[0103] y ij =β0+β1xij +∈ ij

[0104] Among them, β0 and β1 are regression coefficients, ∈ ij is the error term. By analyzing the quality control samples, the regression coefficient is estimated, and then the data of all samples are corrected to eliminate the instrument detection bias caused by batch effect.

[0105] Linkage disequilibrium analysis was performed on the host genome data to remove redundant single nucleotide polymorphism sites. Linkage disequilibrium refers to the non-random association phenomenon between different sites in the genome. In the host genome data, there are a large number of single nucleotide polymorphism (SNP) sites, some of which are highly disequilibrium-linked and carry similar information. Through linkage disequilibrium analysis, the linkage disequilibrium coefficient (such as r) between different SNP sites was calculated. 2 ), when r 2 When the ratio exceeds a certain threshold (e.g., 0.8), it indicates that the two loci are in strong linkage disequilibrium. Selecting one locus to retain and removing the other redundant loci can reduce data dimensionality, lower computational complexity, and avoid overfitting caused by redundant information.

[0106] In terms of data security, a secure storage protocol for multi-omics data is established, in which encryption strength is divided according to the privacy level of data of different dimensions. Metagenomic data, metabolomics data, host genome data, and clinical phenotype group data are divided into different levels according to their privacy sensitivity. For example, host genome data and clinical phenotype group data involve the personal privacy of patients and are set to a high privacy level; metagenomic data and metabolomics data are relatively less sensitive to privacy and are set to a medium privacy level. For data with a high privacy level, a high-strength encryption algorithm, such as AES-256, is used for encrypted storage; for data with a medium privacy level, a relatively weaker encryption algorithm that still guarantees a certain degree of security, such as AES-128, is used for encrypted storage. This ensures data security while also reasonably allocating encryption resources based on the importance and sensitivity of the data.

[0107] During data transmission, dynamic key sharding technology is used to segment encrypted data for transmission and real-time verification. Dynamic key sharding divides the encryption key into multiple segments, dynamically generating and updating these segments during data transmission. Specifically, at the sending end, the encrypted data is divided into multiple data segments according to a specific rule, and a corresponding key fragment is generated for each data segment. These key fragments are transmitted along with the data segments, but the transmission path can be different, enhancing data transmission security. At the receiving end, after receiving the data segments and key fragments, a real-time verification mechanism is used to ensure data integrity and accuracy. This verification process can use methods such as hashing to calculate a hash of the received data segment and compare it with a hash value provided by the sending end. If the hash value matches, the data has not been tampered with during transmission, and verification has passed. If verification fails, a retransmission mechanism is triggered, requiring the sending end to resend the data segment and the corresponding key fragment, thus ensuring the security and reliability of data transmission. This dynamic key sharding technology and real-time verification mechanism effectively prevents data theft or tampering during transmission, ensuring the secure transmission of multi-omics data.

[0108] Accordingly, an embodiment of the present invention further provides an IBS microecological transplantation intelligent prediction system based on multi-omics, comprising:

[0109] A data acquisition module is used to collect four-dimensional data of the target patient through a multi-omics data fusion subsystem, wherein the four-dimensional data includes metagenomic data, metabolomic data, host genome data and clinical phenotype group data;

[0110] A network analysis module is used to input the four-dimensional data into a preset microbiome-metabolite joint network analysis model to construct an interaction network between microbiome and metabolites and extract key network topological features;

[0111] An association calculation module is used to generate a host-microbiome interaction association matrix based on the key network topology features and host genome data, and calculate the functional gene synergy index;

[0112] A dynamic response module, configured to combine the clinical phenotype group data with the functional gene synergy index and generate a symptom-microbe dynamic response index through a dynamic response algorithm;

[0113] A joint modeling module is used to jointly model the interaction network, the host-microbiome interaction matrix, and the symptom-microbe dynamic response index using a preset transfer learning framework, and output prediction results for the efficacy of microecological transplantation.

[0114] The system of this embodiment can be used to perform Figure 1The technical solution of the method embodiment shown has similar implementation principles and technical effects, which will not be repeated here.

[0115] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "includes," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0116] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A multi-omics-driven intelligent prediction method for IBS microecological transplantation, characterized by: include: Establishing a multi-omics data fusion subsystem to collect four-dimensional data of target patients, wherein the four-dimensional data includes metagenomic data, metabolomic data, host genome data and clinical phenotype data; Inputting the four-dimensional data into a preset microbiome-metabolite joint network analysis model to construct an interaction network between microbiome and metabolites and extract key network topological features; Generate a host-microbiome interaction matrix based on the key network topology features and host genome data, and calculate the functional gene synergy index; Combining the clinical phenotype group data with the functional gene synergy index, generating a symptom-microbe dynamic response index through a dynamic response algorithm; The interaction network, the host-microbiome interaction matrix and the symptom-microbe dynamic response index are jointly modeled using a preset transfer learning framework to output the prediction results of the microecological transplantation efficacy.

2. The multi-omics-driven IBS microecological transplantation intelligent prediction method according to claim 1, characterized in that: The steps to establish a multi-omics data fusion subsystem to collect four-dimensional data of target patients include: Obtain metagenomic data of target patients and annotate bacterial flora composition and functional gene expression levels; Collect metabolomic data to extract concentrations of various short-chain fatty acids, bile acids, tryptophan, and metabolic pathway activity indicators; Analyze host genome data and mark single nucleotide polymorphism sites and epigenetic modification regions; The clinical phenotype group data, including symptom rating scale results and historical treatment records, are integrated, and the metagenomic data, the metabolomic data, and the host genome data are aligned in time series.

3. The multi-omics-driven IBS microecological transplantation intelligent prediction method according to claim 1, characterized in that: The microbiome-metabolite joint network analysis model includes: A microbial co-occurrence network was constructed based on the bacterial abundance in metagenomic data, where nodes represent bacterial taxa and edge weights represent the strength of symbiotic or competitive relationships between species. Mapping the metabolite concentrations in the metabolomics data to the microbial co-occurrence network to generate a metabolite-microbial community association subnetwork; The associated subnetwork is traversed by a random walk algorithm to identify the microbiome-metabolite core module with bidirectional regulatory effects.

4. The multi-omics-driven IBS microecological transplantation intelligent prediction method according to claim 3, characterized in that: The steps to generate the host-microbiome interaction matrix include: Extract immune regulatory genes and metabolism-related genes from host genome data as candidate interacting gene sets; Calculate the Pearson correlation coefficient between the expression level of each gene in the candidate interaction gene set and the abundance of the microbial community-metabolite core module; Gene-microbial community pairs whose absolute value of correlation coefficient exceeds the first threshold are screened, and a multidimensional correlation matrix with genes as rows and microbial communities as columns is constructed.

5. The multi-omics-driven IBS microecological transplantation intelligent prediction method according to claim 4 is characterized in that: The implementation of the dynamic response algorithm includes: Dynamic response periods were divided according to symptom severity in clinical phenotype group data; In each period, the lagged correlation between the change rate of the microbial abundance of the microbial community-metabolite core module and the symptom score was calculated; The optimal hysteresis window was determined based on the maximum information coefficient, and the dynamic response weights of the microbiome to symptoms in each time period were calculated.

6. The multi-omics-driven IBS microecological transplantation intelligent prediction method according to claim 5, characterized in that: The steps for building the transfer learning framework include: Pre-training basic prediction models, whose input is the multi-omics data of healthy donors and the corresponding microecological transplantation efficacy labels; Freezing the feature extraction layer of the basic prediction model and adding an adaptive parameter adjustment layer; The joint modeling data of the target patients are input into the adjusted model, and the prediction results are optimized by the domain adaptation loss function.

7. The multi-omics-driven IBS microecological transplantation intelligent prediction method according to claim 6, characterized in that: The method further comprises: Performing a credibility assessment on the efficacy prediction results by calculating the confidence distribution of the prediction results in the feature space in the transfer learning framework; If the confidence level is lower than the second threshold, the incremental learning mechanism is triggered, the current prediction data is added to the training set and the model parameters are re-optimized.

8. The multi-omics-driven IBS microecological transplantation intelligent prediction method according to claim 1, characterized in that: The method further comprises a data preprocessing step: The metagenomic data were filtered for low-abundance bacterial groups, and bacterial genera with relative abundance exceeding the third threshold were retained; Metabolomics data were corrected for batch effects, and a linear regression model based on quality control samples was used to eliminate instrument detection bias; Linkage disequilibrium analysis was performed on the host genome data to remove redundant single nucleotide polymorphism sites.

9. The multi-omics-driven IBS microecological transplantation intelligent prediction method according to claim 8, characterized in that: The method further comprises: Establish a secure storage protocol for multi-omics data, where data of different dimensions are encrypted with different levels of privacy; During the data transmission process, dynamic key sharding technology is used to transmit encrypted data in segments and perform real-time verification.

10. A multi-omics-driven IBS microecological transplantation intelligent prediction system, the system is used to implement the method according to any one of claims 1 to 9, characterized in that: include: A data acquisition module is used to collect four-dimensional data of the target patient through a multi-omics data fusion subsystem, wherein the four-dimensional data includes metagenomic data, metabolomic data, host genome data and clinical phenotype group data; A network analysis module is used to input the four-dimensional data into a preset microbiome-metabolite joint network analysis model to construct an interaction network between microbiome and metabolites and extract key network topological features; An association calculation module is used to generate a host-microbiome interaction association matrix based on the key network topology features and host genome data, and calculate the functional gene synergy index; A dynamic response module, configured to combine the clinical phenotype group data with the functional gene synergy index and generate a symptom-microbe dynamic response index through a dynamic response algorithm; The joint modeling module is used to jointly model the interaction network, the host-microbiome interaction matrix and the symptom-microorganism dynamic response index using a preset transfer learning framework, and output the prediction results of the microecological transplantation efficacy.

Citation Information

Patent Citations

  • Marker combination for diagnosing Basal type pancreatic ductal adenocarcinoma (PDAC) and application thereof

    CN115125312A

  • Alzheimer's disease marker identification method based on microorganism and host interaction

    CN116052767A

  • Methods of diagnosing and treating microbiome-associated disease using interaction network parameters

    WO2011022660A1

  • System and method for drug target and biomarker discovery and diagnosis using a multidimensional multiscale module map

    WO2016118771A1

Cited By

  • Clinical microbial infection intelligent diagnosis system and method based on multi-omics data fusion

    CN120913824A

  • Targeted nutrient source mining method based on beef cattle intestine type-host gene interaction

    CN121884961A

  • Intestinal bacteria transplantation donor and receptor intelligent matching method and system based on multi-modal data

    CN122112659A