Intelligent prediction method and system for ibs micro-ecological transplantation based on multi-omics driving

By integrating multi-omics data and using a transfer learning framework, an IBS microecological transplantation efficacy prediction system was constructed. This system addresses the accuracy issues in IBS diagnosis and treatment assessment, enabling personalized and dynamic efficacy prediction and improving prediction accuracy and data security.

CN120636552BActive Publication Date: 2026-04-10THE FIRST MEDICAL CENT CHINESE PLA GENERAL HOSPITAL
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
THE FIRST MEDICAL CENT CHINESE PLA GENERAL HOSPITAL
Filing Date
2025-05-06
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

Existing diagnostic and treatment assessment methods for IBS lack in-depth analysis of gut microbiota and host intrinsic factors, resulting in low accuracy in predicting the efficacy of microbiota transplantation and making it difficult for clinicians to select appropriate treatment options.

Method used

By establishing a multi-omics data fusion subsystem, metagenomic, metabolomics, host genome, and clinical phenotype data of target patients are collected. A microbiota-metabolite joint network is constructed, a host-microbiota interaction correlation matrix is ​​generated, and combined with a dynamic response algorithm, a transfer learning framework is used for joint modeling to output the efficacy prediction results of microecological transplantation.

Benefits of technology

It achieves accurate prediction of the efficacy of microecological transplantation, takes into account individual differences and the dynamic development of the disease, improves the accuracy and reliability of prediction, reduces dependence on sample size, and ensures data security.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120636552B_ABST
    Figure CN120636552B_ABST
Patent Text Reader

Abstract

The application provides an IBS micro-ecological transplantation intelligent prediction method and system based on multi-omics driving, and relates to the field of biomedical technology. The method comprises the following steps: establishing a multi-omics data fusion subsystem to collect target patient metagenome, metabolome, host genome and clinical phenotype group data; inputting the data into a bacterial flora-metabolite joint network analysis model to construct an interaction network and extract features; generating a correlation matrix based on the features and host genome data and calculating an index; combining the clinical phenotype data and the index to generate an index through a dynamic response algorithm; and using a transfer learning framework to jointly model and output a therapeutic effect prediction result. The system comprises data acquisition, network analysis, correlation calculation, dynamic response and joint modeling modules. The application integrates multi-omics data, accurately mines the relationship between flora and host, realizes intelligent prediction of micro-ecological transplantation efficacy, provides strong support for IBS personalized treatment, and has data processing and security guarantee measures.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the field of biomedical technology, in particular to an IBS micro-ecological transplantation intelligent prediction method and system based on multi-omics driving. BACKGROUND

[0002] Irritable bowel syndrome (IBS) is a common functional gastrointestinal disease with a high incidence worldwide, which seriously affects the quality of life of patients and the social and economic burden. At present, the pathogenesis of IBS has not been fully elucidated, and it is generally believed to be closely related to intestinal micro-ecological imbalance, abnormal immune function, impaired intestinal barrier function, and host genetic factors.

[0003] Among the treatment methods for IBS, micro-ecological transplantation (such as fecal transplantation) shows certain therapeutic potential. By transplanting the intestinal microbial community of healthy donors into patients, it is expected to regulate the intestinal micro-ecological balance of patients, improve the intestinal function, and relieve the symptoms of IBS. However, the efficacy of micro-ecological transplantation varies significantly among different patients, and not all patients can benefit from it. This uncertainty of efficacy makes it a great challenge for clinicians to choose treatment options, as it is difficult to accurately determine which patients are more suitable for micro-ecological transplantation treatment and to predict the treatment effect in advance, thereby limiting the widespread application of micro-ecological transplantation in the treatment of IBS.

[0004] The existing IBS diagnosis and treatment evaluation methods have many limitations. Traditional diagnosis mainly relies on the description of patients' symptoms and simple clinical examination, lacking in-depth analysis of intestinal micro-ecology and host intrinsic factors. In terms of evaluating the efficacy of micro-ecological transplantation, subjective indicators such as clinical symptom scores are currently used, lacking objective and accurate prediction methods. Although some studies attempt to predict efficacy by detecting a single microbial indicator or host factor, due to the complexity of the pathogenesis of IBS, a single indicator is difficult to fully reflect the interaction between the micro-ecological system and the host, resulting in low prediction accuracy.

[0005] With the rapid development of biotechnology, multi-omics technologies, including metagenomics, metabolomics, host genomics, and clinical phenomics, provide new perspectives for in-depth study of the pathogenesis and treatment strategies of IBS. Metagenomics can comprehensively analyze the composition and function of intestinal microbial communities; metabolomics can detect changes in metabolic products in the body, reflecting the metabolic state of the body; host genomics helps to reveal the influence of host genetic factors on disease susceptibility and treatment response; clinical phenomics integrates information such as patients' clinical symptoms, signs, and treatment history. However, how to effectively integrate these multi-omics data and apply them to the intelligent prediction of the efficacy of IBS micro-ecological transplantation remains a problem to be solved in this field.

[0006] In terms of data processing and analysis, multi-omics data has the characteristics of large data volume, high dimensionality and strong complexity, and traditional data processing methods are difficult to mine the key information hidden therein. In addition, there are complex interaction relationships between different omics data, how to build a reasonable model to reveal these relationships and use it for accurate prediction is also an important challenge currently faced by research. SUMMARY

[0007] The purpose of the present application is to provide an IBS micro-ecological transplantation intelligent prediction method and system based on multi-omics driving, to solve the problems raised in the background art.

[0008] To achieve the above-mentioned purpose, the present application provides the following technical solution: an IBS micro-ecological transplantation intelligent prediction method based on multi-omics driving, the method comprising:

[0009] A multi-omics data fusion subsystem is established to collect four-dimensional data of a target patient, wherein the four-dimensional data includes metagenomic data, metabolomic data, host genomic data and clinical phenomic data;

[0010] The four-dimensional data is input into a preset bacterial-metabolite joint network analysis model to construct an interaction network of bacteria and metabolites, and extract key network topology features;

[0011] A host-bacteria interaction correlation matrix is generated based on the key network topology features and the host genomic data, and a functional gene synergy degree index is calculated;

[0012] The clinical phenomic data and the functional gene synergy degree index are combined to generate a symptom-microorganism dynamic response index through a dynamic response algorithm;

[0013] A preset transfer learning framework is used to jointly model the interaction network, the host-bacteria interaction correlation matrix and the symptom-microorganism dynamic response index, and output a micro-ecological transplantation efficacy prediction result.

[0014] Optionally, the step of establishing a multi-omics data fusion subsystem to collect four-dimensional data of a target patient comprises:

[0015] Obtain the metagenomic data of the target patient, and label the bacterial composition and functional gene expression;

[0016] Collect metabolomic data, extract the concentration of various short-chain fatty acids, the concentration of bile acids and the concentration of tryptophan, and the metabolic pathway activity index;

[0017] Analyze the host genomic data, mark single nucleotide polymorphism sites and epigenetic modification regions;

[0018] Integrate clinical phenotype group data, including symptom score scale results and historical treatment records, and align the metagenomic data, the metabolomic data, and the host genomic data in time series.

[0019] Optionally, the microbiota-metabolite joint network analysis model comprises:

[0020] Construct a microbial co-occurrence network based on the microbiota abundance in the metagenomic data, wherein the nodes represent genus taxonomic units, and the edge weights represent the strength of symbiotic or competitive relationships between species;

[0021] Map the metabolite concentrations in the metabolomic data to the microbial co-occurrence network to generate a metabolite-microbiota association subnetwork;

[0022] Identify the microbiota-metabolite core module with bidirectional regulation by traversing the association subnetwork using a random walk algorithm.

[0023] Optionally, the step of generating a host-microbiota interaction association matrix comprises:

[0024] Extract immune regulatory genes and metabolism-related genes in the host genomic data as a candidate interaction gene set;

[0025] Calculate the Pearson correlation coefficient between the expression level of each gene in the candidate interaction gene set and the microbiota abundance of the microbiota-metabolite core module;

[0026] Screen gene-microbiota pairs with absolute correlation coefficients exceeding a first threshold value to construct a multi-dimensional association matrix with genes as rows and microbiota as columns.

[0027] Optionally, the implementation of the dynamic response algorithm comprises:

[0028] Divide the dynamic response period according to the severity of symptoms in the clinical phenotype group data;

[0029] In each period, calculate the lag correlation between the change rate of the microbiota abundance of the microbiota-metabolite core module and the symptom score;

[0030] Determine the optimal lag window based on the maximum information coefficient and calculate the dynamic response weight of the microbiota to the symptoms in each period.

[0031] Optionally, the construction steps of the transfer learning framework comprise:

[0032] Pre-train a basic prediction model with the multi-omics data of healthy donors and the corresponding microecological transplantation efficacy labels as input;

[0033] Freeze the feature extraction layer of the basic prediction model and add an adaptive parameter adjustment layer;

[0034] The combined modeling data of the target patient is input into the adjusted model, and the prediction result is optimized by a domain adaptation loss function.

[0035] Optionally, the method further comprises:

[0036] The efficacy prediction result is evaluated in terms of confidence distribution of the prediction result in the feature space of the transfer learning framework.

[0037] If the confidence is lower than a second threshold value, an incremental learning mechanism is triggered, the current prediction data is added to the training set, and the model parameters are re-optimized.

[0038] Optionally, the method further comprises a data preprocessing step:

[0039] The metagenomic data is filtered for low-abundance flora, and genera with a relative abundance exceeding a third threshold value are retained.

[0040] The metabolomic data is batch effect corrected, and a linear regression model based on quality control samples is used to eliminate instrument detection bias.

[0041] The host genomic data is subjected to linkage disequilibrium analysis, and redundant single nucleotide polymorphism sites are removed.

[0042] Optionally, the method further comprises:

[0043] A secure storage protocol for multi-omics data is established, wherein different dimensions of data are divided according to privacy levels and encryption strengths.

[0044] During data transmission, dynamic key fragmentation technology is used to segment and transmit the encrypted data in real time.

[0045] The present application also provides an IBS micro-ecological transplantation intelligent prediction system based on multi-omics driving, which comprises:

[0046] A data acquisition module is used to acquire four-dimensional data of a target patient through a multi-omics data fusion subsystem, wherein the four-dimensional data includes metagenomic data, metabolomic data, host genomic data and clinical phenotypic data.

[0047] A network analysis module is used to input the four-dimensional data into a preset flora-metabolite joint network analysis model to construct an interaction network of flora and metabolites and extract key network topological features.

[0048] An association calculation module is used to generate a host-flora interaction association matrix based on the key network topological features and host genomic data, and calculate a functional gene synergy degree index.

[0049] a dynamic response module, configured to combine the clinical phenotype group data and the functional gene synergy index, and generate a symptom-microorganism dynamic response index through a dynamic response algorithm;

[0050] a joint modeling module, configured to jointly model the interaction network, the host-microbiota interaction correlation matrix and the symptom-microorganism dynamic response index by using a preset transfer learning framework, and output a microecological transplantation efficacy prediction result.

[0051] Compared with the prior art, the present application has the following beneficial effects:

[0052] From the perspective of data collection and integration, by establishing a multi-omics data fusion subsystem, the macrogenomic data, metabolomic data, host genomic data and clinical phenotype group data of the target patient are comprehensively collected and aligned in time sequence. This multi-dimensional data integration method can comprehensively and dynamically reflect the interaction relationship between the intestinal microecosystem and the host. Compared with the traditional analysis method relying on only a single data type, the information source is greatly enriched, the analysis bias caused by information loss is avoided, and a solid data foundation is laid for subsequent accurate prediction.

[0053] In terms of analysis model construction, the preset microbiota-metabolite joint network analysis model can deeply mine the complex interaction relationship between microbiota and metabolites by constructing a microbial co-occurrence network, generating a metabolite-microbiota correlation subnetwork and identifying a core module, and extract key network topological features. This helps to reveal the internal relationship between microorganisms and metabolic levels in the pathogenesis of IBS, provides a deeper perspective for understanding the occurrence and development process of the disease, and provides a more biologically meaningful index for microecological transplantation efficacy prediction.

[0054] The host-microbiota interaction correlation matrix is generated and the functional gene synergy index is calculated, further integrating the host genomic data and the microbiota information, and analyzing the interaction between the host and the microbiota at the gene level. This analysis method takes into account the influence of host genetic factors on the intestinal microecosystem, and can more accurately evaluate the potential influence of individual differences on the efficacy of microecological transplantation, making the prediction result more personalized and targeted.

[0055] By combining the clinical phenotype group data and the functional gene synergy index, a symptom-microorganism dynamic response index is generated through a dynamic response algorithm, fully considering the dynamic change relationship between disease symptoms and microorganisms. The index can reflect the influence degree of microorganisms on symptoms in different time periods in real time, capture the subtle changes between the microecosystem and clinical symptoms, and compared with the traditional static analysis method, better reflect the dynamic development process of the disease, thereby providing a more timely and accurate basis for efficacy prediction.

[0056] By using the preset transfer learning framework to jointly model the multi-omics data, not only can the multi-omics data of healthy donors and the corresponding micro-ecological transplantation efficacy labels be pre-trained, but also the model can be better adapted to the data characteristics of the target patients through the adaptive parameter adjustment layer and the domain adaptation loss function, so as to optimize the prediction result. This transfer learning method improves the generalization ability of the model and reduces the dependence on a large number of target patient samples, and can still achieve accurate prediction in the case of limited sample size, greatly improving the prediction efficiency and accuracy.

[0057] In addition, the credibility of the efficacy prediction result is evaluated, and the incremental learning mechanism is triggered when the confidence is low, which can continuously optimize the model. By adding new prediction data to the training set to re-optimize the model parameters, the model can learn more diverse data features and gradually improve the prediction ability for different patients, further improving the reliability of the prediction.

[0058] In terms of data processing and security protection, a series of effective measures are implemented. The data preprocessing steps, such as filtering low-abundance flora for metagenomic data, batch effect correction for metabolomic data, and linkage disequilibrium analysis for host genomic data, improve the data quality and reduce the interference of noise and redundant information on the prediction result. At the same time, the established multi-omics data security storage protocol and the adopted dynamic key fragmentation technology guarantee the security of the data in the storage and transmission process, protect the privacy of patients, and provide a safe and reliable environment for the clinical application of multi-omics data. BRIEF DESCRIPTION OF DRAWINGS

[0059] Figure 1 The working principle diagram of the IBS micro-ecological transplantation intelligent prediction method based on multi-omics driving according to the present application is shown in the figure;

[0060] Figure 2 The workflow diagram of the flora-metabolite joint network analysis model is shown in the figure;

[0061] Figure 3 The workflow diagram of the efficacy prediction result evaluation and model optimization is shown in the figure;

[0062] Figure 4 The workflow diagram of the data preprocessing is shown in the figure. DETAILED DESCRIPTION

[0063] The technical solutions in the embodiments of the present application will be described clearly and completely below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, not all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.

[0064] Please refer toFigures 1-4 The present application relates to a multi-omics driven IBS micro-ecological transplantation intelligent prediction method and system, which will be described in detail below.

[0065] A multi-omics data fusion subsystem is established to collect four-dimensional data of the target patient, including metagenomic data, metabolomic data, host genomic data, and clinical phenomic data. This subsystem integrates data obtained from different dimensions, providing a comprehensive information base for subsequent analysis.

[0066] The four-dimensional data is input into a preset bacterial-metabolite joint network analysis model to construct the interaction network of bacteria and metabolites, and extract key network topology features. This step mines the internal relationship between bacteria and metabolites through a specific model, and finds network features that are important for subsequent analysis.

[0067] Based on the key network topology features and host genomic data, a host-bacterial interaction correlation matrix is generated, and a functional gene synergy degree index is calculated. In this way, the interaction between host genes and bacteria is analyzed, and the micro-ecological system is further understood.

[0068] The clinical phenomic data and the functional gene synergy degree index are combined, and a symptom-microorganism dynamic response index is generated through a dynamic response algorithm. This index reflects the dynamic correlation between microorganisms and patient symptoms, providing a more accurate basis for prediction.

[0069] The interaction network, the host-bacterial interaction correlation matrix, and the symptom-microorganism dynamic response index are jointly modeled using a preset transfer learning framework, and the micro-ecological transplantation efficacy prediction result is output. With the help of the transfer learning framework, the existing data and models are fully utilized to improve the accuracy and reliability of the prediction.

[0070] The other technical features of the present application will be described in detail through specific embodiments.

[0071] Example 1:

[0072] The establishment of a multi-omics data fusion subsystem to collect four-dimensional data of the target patient specifically includes: obtaining metagenomic data of the target patient, and labeling the composition of bacterial groups and the expression level of functional genes. The specific operation is as follows: using professional gene sequencing equipment to perform metagenomic sequencing on the fecal sample of the patient to obtain the original sequencing data. Through bioinformatics analysis tools such as MetaPhlAn, HUMAnN, etc., the sequencing data is processed to determine the types and relative abundances of various bacterial groups in the sample, i.e., to label the composition of bacterial groups. At the same time, the expression level of functional genes is quantitatively analyzed using these tools to identify genes related to micro-ecological function and record their expression levels.

[0073] Metabolome data is collected, and concentrations of various short-chain fatty acids, bile acids, and tryptophan, as well as metabolic pathway activity indicators, are extracted. Metabolome mass spectrometry data is collected from the patient's blood, feces, or other suitable biological samples. Metabolic data processing software such as XCMS, MZmine, etc. is used to perform peak identification, alignment, and quantitative analysis on the mass spectrometry data. Various short-chain fatty acids, bile acids, or tryptophan are important products of intestinal microbial metabolism and are closely related to the pathogenesis of IBS. Through specific detection methods and data analysis means, the concentrations of various short-chain fatty acids, bile acids, or tryptophan in the sample are accurately determined. For metabolic pathway activity indicators, by mapping metabolites to known metabolic pathway databases such as KEGG, Reactome, etc., using pathway enrichment analysis algorithms, the activity level of each metabolic pathway is calculated to reflect changes in metabolic function.

[0074] Host genome data is analyzed, and single nucleotide polymorphism sites and epigenetic modification regions are marked. High-throughput sequencing technology is used to sequence the host genome, obtaining a large amount of genetic sequence information. SNP calling tools such as GATK, SAMtools, etc. are used to analyze the sequencing data and identify single nucleotide polymorphism (SNP) sites. Changes in these sites can affect the host's susceptibility to disease and interaction with microorganisms. At the same time, epigenetic research techniques such as chromatin immunoprecipitation sequencing (ChIP-seq), whole-genome bisulfite sequencing (WGBS), etc. are used to determine epigenetic modification regions, including DNA methylation, histone modification, etc. These modifications can regulate gene expression, thereby affecting host-microbe interactions and the development of disease.

[0075] Clinical phenotype group data, including symptom score scale results and historical treatment records, are integrated, and the metagenome data, metabolome data, and host genome data are aligned in time sequence. Collect clinical phenotype information of patients, such as using IBS symptom severity score scale (IBS-SSS) to quantitatively score symptoms such as abdominal pain, abdominal distension, and changes in bowel habits of patients. At the same time, historical treatment records of patients are sorted, including drugs used, treatment time, treatment effect, etc. In order to facilitate comprehensive analysis, the metagenome data, metabolome data, and host genome data collected at different time points are aligned in time sequence to ensure the time consistency of the data, so that subsequent analysis can accurately reflect the dynamic changes of the disease.

[0076] Example 2:

[0077] The construction process of the microbial-metabolite combined network analysis model is as follows: a microbial co-occurrence network is constructed based on the microbial community abundance in the metagenomic data. In this network, each microbial genus classification unit is regarded as a node, and the edge weight is determined by calculating the correlation between different microbial genera. If two microbial genera often appear together in a sample, it indicates that there is a symbiotic relationship between them, and the edge weight is positive; on the contrary, if two microbial genera rarely appear together, it may indicate a competitive relationship, and the edge weight is negative. The specific calculation method can use the Spearman correlation coefficient or the Pearson correlation coefficient, and the formula is:

[0078]

[0079] wherein r represents the correlation coefficient, x i and y i represent the abundance of the two microbial genera in the i-th sample, and are the average abundance of the two microbial genera in all samples, and n is the number of samples. When |r| exceeds a certain threshold (such as 0.6), it is considered that there is a significant co-occurrence relationship between the two microbial genera, and a connection is established in the network.

[0080] The metabolite concentration in the metabolomic data is mapped to the microbial co-occurrence network to generate a metabolite-microbial community correlation subnetwork. The metabolites are regarded as new nodes, and the connections between the metabolite nodes and the microbial genus nodes are added in the microbial co-occurrence network according to the known biochemical relationship between the metabolites and the microbial genera or the correlation obtained by data analysis. For example, if a certain microbial genus can produce a specific metabolite, or the concentration changes of the two show significant correlation, the corresponding edge is established in the network. In this way, a correlation subnetwork containing microbial community and metabolite information is constructed, which more comprehensively shows the interaction relationship in the micro-ecosystem.

[0081] The random walk algorithm is used to traverse the correlation subnetwork to identify the microbial-metabolite core module with bidirectional regulation. The random walk algorithm is an algorithm for random exploration on graph structure. In the correlation subnetwork, starting from a randomly selected node, move to the connected nodes according to a certain probability, and repeat this process continuously. In each moving process, the nodes and edges passed are recorded. After a large number of random walk steps, analyze which nodes and edges are frequently accessed. The microbial-metabolite core module with bidirectional regulation is usually the part of the subnetwork that is frequently accessed in the random walk process. In this way, the key core module that plays a key role in the interaction between microbial community and metabolite can be found, which provides an important basis for subsequent analysis.

[0082] Example 3:

[0083] In generating the host-microbiota interaction correlation matrix, immune regulation genes and metabolism-related genes in the host genome data are extracted as the candidate interaction gene set. From existing gene databases such as NCBI, Ensembl, etc., obtain gene information related to immune regulation and metabolism. Combined with the pathogenesis of IBS and existing research results, genes that may interact with intestinal flora are screened to form the candidate interaction gene set. These genes may be involved in regulating the host's immune response, affecting intestinal barrier function, or participating in the processing of microbial metabolites, etc.

[0084] The Pearson correlation coefficient between the expression level of each gene in the candidate interaction gene set and the abundance of the core module of the flora-metabolite is calculated. Using gene expression profile data and flora-metabolite core module abundance data, the Pearson correlation coefficient between the expression level of each gene and the abundance of each flora-metabolite core module is calculated. The formula for calculating the Pearson correlation coefficient is the same as the formula for calculating the correlation coefficient between genera in Example 2 (in the formula, x i represents the expression of the gene in the i-th sample, y i represents the abundance of the core module of the flora in the i-th sample). The coefficient reflects the degree of linear correlation between gene expression and flora abundance, and the value range is between -1 and 1. A positive value indicates a positive correlation, that is, when the expression level of the gene increases, the abundance of the flora also tends to increase; a negative value indicates a negative correlation, that is, when the expression level of the gene increases, the abundance of the flora tends to decrease.

[0085] The gene-flora pairs with an absolute value of the correlation coefficient exceeding a first threshold (e.g., 0.5) are screened, and a multi-dimensional correlation matrix with genes as rows and flora as columns is constructed. According to the calculated Pearson correlation coefficient, gene-flora pairs with strong correlation are screened. These gene-flora pairs are arranged into a multi-dimensional matrix, with rows representing different genes and columns representing different flora core modules. The elements in the matrix are the correlation coefficients of the corresponding gene-flora pairs, thus constructing the host-microbiota interaction correlation matrix, which directly shows the interaction relationship between host genes and flora.

[0086] Example 4:

[0087] The implementation of the dynamic response algorithm is as follows: the dynamic response period is divided according to the symptom severity in the clinical phenotype group data. First, the symptom severity is quantified, for example, the symptoms are divided into mild (75<IBS-SSS score≤175), moderate (175<IBS-SSS score≤300) and severe (IBS-SSS score>300) according to the IBS-SSS score. Then, in chronological order, the patient's course is divided into different periods, each period corresponding to a different symptom severity change phase. For example, when the symptoms change from mild to moderate, a new period is divided; after the symptoms remain stable at the moderate level for a period of time, if there is a significant change, a new period is divided. In this way, the relationship between microorganisms and symptoms can be more accurately analyzed according to the dynamic changes of symptoms.

[0088] In each period, the lag correlation between the core module of the change rate of the abundance of the flora-metabolite and the symptom score is calculated. For each divided period, the change rate of the abundance of the core module of the flora at different time points is calculated. Let t1 and t2 be two time points in the period, and the formula for calculating the change rate of the abundance of the flora is:

[0089]

[0090] At the same time, the symptom score at the corresponding time point is recorded. Then, the lag correlation between the change rate of the abundance of the flora and the symptom score is calculated. The lag correlation refers to the correlation between the change in the abundance of the flora and the change in the symptoms in time. For example, the correlation between the change rate of the abundance of the flora at t1 and the symptom score at t1+k (k is the lag time, k=1, 2,...) is calculated. In this way, it can be found whether there is a time delay in the response of microorganisms to the change in symptoms.

[0091] The optimal lag window is determined based on the maximum information coefficient, and the dynamic response weight of the flora to the symptoms in each period is calculated. The maximum information coefficient (MIC) is an index for measuring the complex correlation between two variables. After calculating the correlation between the change rate of the abundance of the flora and the symptom score at different lag times, the MIC is used to find the lag time with the strongest correlation, which is the optimal lag window. After determining the optimal lag window, the dynamic response weight of the flora to the symptoms in each period is calculated according to the correlation strength of the change rate of the abundance of the flora and the symptom score at the optimal lag window. The stronger the correlation, the greater the dynamic response weight, indicating that the influence of the flora on the symptoms is more significant. In this way, an index that can accurately reflect the dynamic response relationship between symptoms and microorganisms is generated.

[0092] Example 5:

[0093] The construction steps of the transfer learning framework are as follows: pre-training a basic prediction model with the multi-omics data of healthy donors and the corresponding micro-ecological transplantation efficacy labels. Collect a large amount of metagenomic data, metabolomic data, host genomic data of healthy donors, and the efficacy data of these donors after micro-ecological transplantation (efficacy labels can be represented by cure rate, symptom improvement degree, etc.). Select a suitable machine learning model, such as deep neural network (DNN), random forest (RF), etc., use the multi-omics data of these healthy donors as input, and the efficacy label as output to train the model. During training, adjust the parameters of the model to enable the model to learn the relationship between the multi-omics data of healthy donors and the efficacy of micro-ecological transplantation, and obtain a pre-trained basic prediction model.

[0094] Freeze the feature extraction layer of the basic prediction model and add an adaptive parameter adjustment layer. After pre-training, the basic prediction model has learned important features in the multi-omics data. In order to adapt to the data characteristics of the target patient, the feature extraction layer is frozen so that its parameters are no longer updated. Then, an adaptive parameter adjustment layer is added to the model. This adjustment layer can fine-tune the model according to the data of the target patient, for example, by adding a fully connected layer, a convolutional layer, etc., and training the parameters of the adjustment layer using the data of the target patient, so that the model can better adapt to the micro-ecosystem and clinical characteristics of the target patient.

[0095] Input the joint modeling data of the target patient into the adjusted model and optimize the prediction results through the domain adaptation loss function. Input the interaction network, correlation matrix, and dynamic response index of the target patient into the model with the adaptive parameter adjustment layer. In order to make the model's prediction on the target patient data more accurate, use the domain adaptation loss function to optimize the model. The role of the domain adaptation loss function is to minimize the difference between the source domain (healthy donor data) and the target domain (target patient data), so that the model can better predict on the target patient data. For example, the maximum mean discrepancy (MMD) can be used as the domain adaptation loss function, and its formula is:

[0096]

[0097] Where x and y represent the data of the source domain and the target domain, respectively, x i and y j are the samples in the source domain and the target domain, n and m are the number of samples in the source domain and the target domain, φ is a function that maps samples to a reproducing kernel Hilbert space (RKHS), denotes the norm in RKHS. By continuously adjusting the parameters of the adaptive parameter adjustment layer, the domain adaptation loss function is minimized, thereby optimizing the model's prediction of the efficacy of micro-ecological transplantation for the target patient.

[0098] Example 6:

[0099] The method of the present application also includes confidence evaluation of the therapeutic effect prediction result and related data processing and security measures. The confidence evaluation of the therapeutic effect prediction result is performed in the manner of calculating the confidence distribution of the prediction result in the feature space of the transfer learning framework. In the transfer learning framework, the distribution of the prediction result in the feature space can be obtained through the prediction of the model on the target patient data. The confidence of the prediction result is evaluated by using some statistical methods, such as calculating the probability density function or confidence interval of the prediction result. For example, Gaussian Mixture Model (GMM) can be used to model the prediction result, estimate the probability of the prediction result belonging to different categories, and thus obtain the confidence distribution of the prediction result.

[0100] If the confidence is lower than the second threshold value (such as 0.6), the incremental learning mechanism is triggered to add the current prediction data to the training set and re-optimize the model parameters. When the confidence of the prediction result is lower than the set threshold value, it indicates that the reliability of the model for the prediction result is low. In order to improve the accuracy of the model, the incremental learning mechanism is triggered. The multi-omics data of the current target patient and the corresponding prediction result (even if the prediction result may not be accurate) are added to the training set, and then the model is retrained and optimized. In the retraining process, the model can learn more sample information, especially the data features similar to the target patient, so as to adjust the parameters of the model and improve the prediction ability of the model for similar patients.

[0101] In terms of data preprocessing, low-abundance bacterial flora filtering is performed on metagenomic data to retain bacterial genera with relative abundance exceeding a third threshold value (such as 0.01%). There are a large number of bacterial flora with very low relative abundance in metagenomic data, which may be caused by experimental errors or environmental contaminants, have little effect on the analysis results, and increase the computational burden. By setting a relative abundance threshold, low-abundance bacterial flora is filtered out, only dominant bacterial flora with biological significance is retained, and the accuracy and efficiency of data analysis are improved.

[0102] Batch effect correction is performed on the metabolomic data, and a linear regression model based on quality control samples is used to eliminate instrument detection bias. During the collection of metabolomic data, due to differences in experimental conditions, instrument states and other factors of different batches, batch effects may occur, affecting the accuracy of the data. Using quality control samples (the same standard sample is added in each batch experiment), the metabolomic data is corrected by a linear regression model. Let y ij be the metabolite measurement value of the i-th sample in the j-th batch, x ij be the corresponding covariate (such as batch number), the linear regression model can be expressed as:

[0103] y ij = β0+ β1xij + ∈ ij

[0104] where β0 and β1 are regression coefficients, ∈ ij is the error term. By analyzing the quality control samples, the regression coefficients are estimated, and then the data of all samples are corrected to eliminate the instrument detection bias caused by batch effects.

[0105] Linkage disequilibrium analysis is performed on the host genome data to remove redundant single nucleotide polymorphism sites. Linkage disequilibrium refers to the non-random association between different sites in the genome. In the host genome data, there are a large number of single nucleotide polymorphism (SNP) sites, some of which have high linkage disequilibrium and carry similar information. Through linkage disequilibrium analysis, the linkage disequilibrium coefficient (such as r 2 ) between different SNP sites is calculated. When r 2 exceeds a certain threshold (such as 0.8), it indicates that there is a strong linkage disequilibrium relationship between the two sites. One of the sites is selected to remain, and the other redundant sites are removed. This can reduce the data dimension and reduce the computational complexity, while avoiding the overfitting problem caused by redundant information.

[0106] In terms of data security, a multi-omics data security storage protocol is established, in which different dimensions of data are divided into encryption strengths according to privacy levels. The metagenomic data, metabolomic data, host genome data, and clinical phenotyping data are divided into different levels according to their privacy sensitivity, for example, the host genome data and clinical phenotyping data involve patient privacy, and are set to high privacy level; the metagenomic data and metabolomic data have relatively low privacy sensitivity, and are set to medium privacy level. For high-privacy-level data, high-strength encryption algorithms such as AES-256 are used for encryption storage; for medium-privacy-level data, relatively weak but still secure encryption algorithms such as AES-128 are used for encryption storage. In this way, while ensuring data security, encryption resources can also be reasonably allocated according to the importance and sensitivity of the data.

[0107] In the data transmission process, the dynamic key fragmentation technology is used to segment the encrypted data for transmission and real-time verification. The dynamic key fragmentation technology is to divide the encryption key into multiple fragments, and dynamically generate and update these fragments during data transmission. Specifically, at the sending end, the encrypted data is divided into multiple data segments according to certain rules, and a corresponding key fragment is generated for each data segment. These key fragments are transmitted together with the data segments, but the transmission paths can be different, increasing the security of data transmission. At the receiving end, after receiving the data segments and key fragments, a real-time verification mechanism is used to ensure the integrity and accuracy of the data. The verification process can use hash check and other methods to calculate the hash of the received data segment and compare it with the hash value provided by the sending end. If they are consistent, it means that the data has not been tampered with during transmission, and the verification is passed. If the verification fails, a retransmission mechanism is triggered to require the sending end to resend the data segment and the corresponding key fragment, thereby ensuring the security and reliability of data transmission during transmission. Through this dynamic key fragmentation technology and real-time verification mechanism, the data is effectively prevented from being stolen or tampered with during transmission, and the secure transmission of multi-omics data is ensured.

[0108] Correspondingly, the embodiment of the application also provides an IBS micro-ecological transplantation intelligent prediction system based on multi-omics driving, comprising:

[0109] A data acquisition module is configured to acquire four-dimensional data of a target patient through a multi-omics data fusion subsystem, wherein the four-dimensional data comprises metagenomic data, metabolomic data, host genomic data, and clinical phenomic data.

[0110] A network analysis module is configured to input the four-dimensional data into a preset bacterial flora-metabolite joint network analysis model to construct an interaction network of bacterial flora and metabolites and extract key network topology features.

[0111] An association calculation module is configured to generate a host-bacterial flora interaction association matrix based on the key network topology features and the host genomic data and calculate a functional gene synergy degree index.

[0112] A dynamic response module is configured to combine the clinical phenomic data and the functional gene synergy degree index to generate a symptom-microorganism dynamic response index through a dynamic response algorithm.

[0113] A joint modeling module is configured to use a preset transfer learning framework to jointly model the interaction network, the host-bacterial flora interaction association matrix, and the symptom-microorganism dynamic response index and output a micro-ecological transplantation efficacy prediction result.

[0114] The system of the embodiment can be used to perform Figure 1The technical solutions, implementation principles and technical effects of the method embodiments shown are similar, and thus will not be described herein.

[0115] It should be noted that, in this document, the terms such as first and second are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between such entities or operations. Also, the terms "comprises", "comprising", or any other variations thereof are intended to cover a non-exclusive inclusion, so that a process, method, article, or apparatus that comprises a list of elements does not include only those elements but can also include other elements not expressly listed or inherent to such process, method, article, or apparatus.

[0116] Although the embodiments of the present application have been shown and described, it will be understood by those of ordinary skill in the art that various changes, modifications, alternatives, and variations can be made thereto without departing from the principles and spirit of the application, and the scope of the present application is defined by the appended claims and their equivalents.

Claims

1. An IBS micro-ecological transplantation intelligent prediction method based on multi-omics driving, characterized in that, The application relates to a method for predicting the efficacy of microecological transplantation. The method comprises the following steps:

1. Collecting four-dimensional data of a target patient by establishing a multi-omics data fusion subsystem, wherein the four-dimensional data comprises metagenomic data, metabolomic data, host genomic data and clinical phenomic data; 2. Inputting the four-dimensional data into a preset bacterial flora-metabolite joint network analysis model to construct an interaction network of bacterial flora and metabolites and extract key network topological features; 3. Generating a host-bacterial flora interaction correlation matrix based on the key network topological features and the host genomic data and calculating a functional gene synergy degree index; 4. Combining the clinical phenomic data and the functional gene synergy degree index to generate a symptom-microorganism dynamic response index by a dynamic response algorithm, wherein the dynamic response algorithm comprises the following steps: 1) Dividing dynamic response time periods according to the severity of symptoms in the clinical phenomic data: first, quantifying the severity of symptoms, and dividing the symptoms into mild (IBS-SSS score <75), moderate (75<=IBS-SSS score<175) and severe (IBS-SSS score >=175) according to the IBS-SSS score; then, dividing the course of the patient into different time periods in chronological order, and each time period corresponds to a different symptom severity change stage: a new time period is divided when the symptoms change from mild to moderate; when the symptoms remain stable at the moderate level for a period of time, a new time period is divided if there is a significant change; in this way, the relationship between microorganisms and symptoms can be more accurately analyzed according to the dynamic changes of the symptoms; 2) In each time period, the lag correlation between the core module bacterial flora abundance change rate and the symptom score is calculated: for each divided time period, the change rate of the core module bacterial flora abundance at different time points is calculated, and the formula for calculating the change rate of the bacterial flora abundance is: (t2-t1) / t1, wherein t1 and t2 are two time points in the time period; meanwhile, the symptom score at the corresponding time point is recorded; then, the lag correlation between the bacterial flora abundance change rate and the symptom score is calculated, and the lag correlation refers to the correlation between the bacterial flora abundance change and the symptom change when the bacterial flora abundance change lags behind the symptom change in time; in this way, it is found whether there is a time delay in the response of the microorganisms to the symptom change; 3) Determining the optimal lag window based on the maximum information coefficient MIC and calculating the dynamic response weight of the bacterial flora to the symptoms in each time period: the MIC is an index for measuring the complex correlation between two variables, and after calculating the correlation between the bacterial flora abundance change rate and the symptom score at different lag times, the MIC is used to find the lag time with the strongest correlation, which is the optimal lag window; after the optimal lag window is determined, the dynamic response weight of the bacterial flora to the symptoms in each time period is calculated according to the correlation strength of the bacterial flora abundance change rate and the symptom score at the optimal lag window; the stronger the correlation, the greater the dynamic response weight, and the more significant the influence of the bacterial flora on the symptoms; in this way, an index that can accurately reflect the dynamic response relationship between the symptoms and the microorganisms is generated; 5. Joint modeling of the interaction network, the correlation matrix and the dynamic response index by using a preset transfer learning framework to output a microecological transplantation efficacy prediction result. 2.The multi-omics driven IBS micro-ecological transplantation intelligent prediction method according to claim 1, characterized in that, The step of establishing the four-dimensional data fusion subsystem comprises: Obtaining metagenomic sequencing data of a target patient, and labeling the composition of flora and the expression amount of functional genes; Collecting metabolomic mass spectrometry data, extracting short-chain fatty acid concentration and metabolic pathway activity indicators; Analyzing host genome sequencing data, marking single nucleotide polymorphism sites and epigenetic modification regions; Integrating clinical phenotype group data, including symptom score scale results and historical treatment records, and aligning the metagenomic data, metabolomic data and host genomic data in time sequence. 3.The multi-omics driven IBS micro-ecological transplantation intelligent prediction method according to claim 1, characterized in that, The flora-metabolite joint network analysis model comprises: Constructing a microbial co-occurrence network based on the flora abundance in the metagenomic data, wherein the nodes represent genus classification units, and the edge weight represents the strength of the symbiotic or competitive relationship between species; Mapping the metabolite concentration in the metabolomic data to the microbial co-occurrence network to generate a metabolite-flora correlation subnetwork; Identifying a flora-metabolite core module with bidirectional regulation by traversing the correlation subnetwork through a random walk algorithm. 4.The multi-omics driven IBS micro-ecological transplantation intelligent prediction method according to claim 3, characterized in that, The step of generating a host-flora interaction correlation matrix comprises: Extracting immune regulatory genes and metabolism-related genes in the host genomic data as a candidate interaction gene set; Calculating the Pearson correlation coefficient between the expression level of each gene in the candidate interaction gene set and the abundance of the flora core module; Screening gene-flora pairs with absolute correlation coefficients exceeding a first threshold value to construct a multi-dimensional correlation matrix with genes as rows and flora as columns. 5.The multi-omics driven IBS micro-ecological transplantation intelligent prediction method based on claim 1, characterized in that, The construction step of the transfer learning framework comprises: Pre-training a basic prediction model, with the input being the multi-omics data of healthy donors and the corresponding microecological transplantation efficacy labels; Freezing the feature extraction layer of the basic prediction model and adding an adaptive parameter adjustment layer; Inputting the joint modeling data of the target patient into the adjusted model to optimize the prediction results through a domain adaptation loss function. 6.The multi-omics driven IBS micro-ecological transplantation intelligent prediction method according to claim 5, characterized in that, The method further comprises: Performing credibility evaluation on the efficacy prediction results by calculating the confidence distribution of the prediction results in the feature space of the transfer learning framework; If the confidence is lower than a second threshold value, triggering an incremental learning mechanism to add the current prediction data to the training set and re-optimize the model parameters. 7.The multi-omics driven IBS micro-ecological transplantation intelligent prediction method based on claim 1, characterized in that, The method further comprises a data preprocessing step: Filtering low-abundance flora from the metagenomic data, retaining genera with a relative abundance exceeding a third threshold value; Performing batch effect correction on the metabolomic data, using a linear regression model based on quality control samples to eliminate instrument detection bias; Performing linkage disequilibrium analysis on the host genomic data to remove redundant single nucleotide polymorphism sites. 8.The multi-omics driven IBS micro-ecological transplantation intelligent prediction method according to claim 7, characterized in that, The method further comprises: Establishing a secure storage protocol for multi-omics data, wherein different dimensions of data are divided into encryption strengths according to privacy levels; In the data transmission process, dynamic key fragmentation technology is used to segment and transmit the encrypted data and perform real-time verification.

9. An IBS micro-ecological transplantation intelligent prediction system based on multi-omics driving, characterized by, Comprise: A data acquisition module for acquiring four-dimensional data of a target patient through a multi-omics data fusion subsystem, wherein the four-dimensional data comprises metagenomic data, metabolomic data, host genomic data and clinical phenotype group data; a network analysis module configured to input the four-dimensional data into a preset microbiota-metabolite joint network analysis model, to construct a microbiota-metabolite interaction network, and to extract key network topological features; a correlation calculation module configured to generate a host-microbiota interaction correlation matrix based on the key network topological features and host genome data, and to calculate a functional gene synergy degree index; a dynamic response module configured to combine the clinical phenotype group data and the functional gene synergy degree index, and to generate a symptom-microorganism dynamic response index through a dynamic response algorithm; a joint modeling module configured to perform joint modeling on the interaction network, the correlation matrix, and the dynamic response index by using a preset transfer learning framework, and to output a microecological transplantation efficacy prediction result.

Citation Information

Patent Citations

  • Marker combination for diagnosing Basal type pancreatic ductal adenocarcinoma (PDAC) and application thereof

    CN115125312A

  • System and method for drug target and biomarker discovery and diagnosis using a multidimensional multiscale module map

    WO2016118771A1