Identification method and system for disease-related microorganism key global regulation species
By constructing a Bayesian network model of healthy samples and simulating the impact of intervention assessment, the problem of insufficient causal relationships and global regulation in the identification of key microbial regulatory species in existing technologies is solved, and high-precision identification and targeted intervention of disease-related microbial regulatory species are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- JIANGNAN UNIV
- Filing Date
- 2026-02-05
- Publication Date
- 2026-05-15
AI Technical Summary
Existing technologies, when identifying key regulatory species of disease-related microorganisms, suffer from limited accuracy, poor targeting, and insufficient interpretability because they cannot effectively infer causal relationships between species and lack global regulatory capabilities.
We construct a Bayesian network model based on healthy samples, assess the influence of each species on the global state of the network through simulated interventions, quantify its influence using the HITS algorithm and a greedy strategy, and identify potential global regulatory species.
It enables precise identification of key global regulatory species of disease-related microorganisms, improves the accuracy and interpretability of intervention target selection, and can provide clear causal clues from the perspective of causal association and global impact, supporting targeted intervention.
Smart Images

Figure CN122050520A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to a method and system for identifying key global regulatory species of disease-related microorganisms, belonging to the interdisciplinary field of microbiology and computer science. Background Technology
[0002] In recent years, the biological functions of the gut microbiota as the human "second genome" have received widespread attention. Clinical studies have shown that the gut microbiota forms a complex symbiotic network with the host in maintaining health through pathways such as metabolite secretion (e.g., short-chain fatty acids, secondary bile acids), immune regulation (Treg cell differentiation, IgA production), and neuroendocrine regulation (gut-brain axis signaling). Large-scale cohort studies (such as MetaHIT and the American Gut Project) have confirmed that gut microbiota dysbiosis is significantly associated with metabolic syndrome (obesity, type II diabetes), autoimmune diseases (inflammatory bowel disease, rheumatoid arthritis), cardiovascular diseases (atherosclerosis), and neurodegenerative diseases (Alzheimer's disease, Parkinson's disease). For example, the correlation between an elevated Firmicutes / Bacteroidetes ratio and enhanced energy absorption in obese patients has been supported by multi-omics evidence, while the absence of butyrate-producing bacteria (such as Faecalibacterium prausnitzii) has been found to be closely related to intestinal barrier dysfunction in IBD patients.
[0003] Although microbiome technologies (such as 16S rRNA gene sequencing and metagenomic assembly of genomes using MAGs) can systematically analyze the characteristics of gut microbiota composition, existing methods still have significant limitations in translational applications for disease prevention. First, insufficient species resolution: 16S rRNA V3-V4 region sequencing struggles to distinguish functional heterogeneity at the species / strain level, while metagenomics, although improving classification accuracy, has limited sensitivity for detecting low-abundance species (<0.1% relative abundance). Second, lack of causal inference: Traditional differential species analysis (such as LEfSe and DESeq2) can only identify co-occurring species associated with disease phenotypes, failing to distinguish whether microbiota changes are a cause or a secondary consequence of disease. Furthermore, indirectness in functional prediction: Microbiota functional annotation based on KEGG / COG databases fails to accurately reflect metabolic activity and host-microbe interaction mechanisms within the gut microenvironment. Finally, ambiguity in intervention targets: While machine learning models (such as random forests and XGBoost) can improve disease prediction accuracy, they lack interpretable identification of key regulatory species and their pathways of action. Therefore, there is an urgent need to develop a more precise technical framework to overcome the limitations of existing methods and promote the translational application of microbiome in disease prevention.
[0004] Currently, intervention strategies targeting the gut microbiota mainly focus on probiotic / prebiotic supplementation, fecal microbiota transplantation (FMT), and dietary adjustments. However, these methods generally suffer from two key scientific problems. First, the inefficiency of broad-spectrum interventions: non-specific regulation easily leads to the disruption of the microbiota ecosystem's stability; clinical studies show that approximately 35% of FMT patients experience failed microbiota remodeling. Second, the blindness of targeting mechanisms: existing methods are mostly based on empirical species selection (such as Lactobacillus and Bifidobacterium), lacking systematic identification and causal verification of core regulatory nodes. Microbial network analysis tools developed in recent years (such as SparCC and MENAP) attempt to identify keystone species by constructing species interaction networks, but these tools, based on Pearson correlation coefficients or sparse regression-based association inference methods, cannot effectively analyze complex relationships, including the regulatory role of host factors (such as genetic background and immune status) on microbiota structure, cross-species signal transduction pathways mediated by microbiota metabolites, and dynamic succession patterns and feedback mechanisms over time. Therefore, existing intervention strategies still face significant challenges in terms of targeting and effectiveness.
[0005] Currently, methods for predicting microbe-disease associations based on statistics, network analysis, and machine learning have been proposed. However, these methods typically employ simple feature concatenation or linear weighting, making it difficult to fully integrate multimodal information from different data sources (such as microbe functional similarity, disease symptom information, and association networks). Furthermore, traditional methods are insufficient in capturing global structural information and local high-order interactions, resulting in low predictive performance and generalization ability of the models in large-scale, high-noise data environments, and an inability to accurately identify key regulatory species. Summary of the Invention
[0006] [Technical Issues] The technical problem to be solved by this invention is: how to overcome the technical defects of existing technologies in identifying key regulatory species of disease-related microorganisms, which are limited in accuracy, poor targeting and insufficient interpretability due to the inability to effectively infer causal relationships between species and the lack of assessment of the global regulatory capacity of species. In order to provide a method and system that can accurately and interpretably identify key global regulatory species from the perspective of causal association and global influence.
[0007] [Technical Solution] To improve the accuracy of identifying key global regulatory species of disease-related microorganisms, this invention provides a method and system for identifying key global regulatory species of disease-related microorganisms, the technical solution of which is as follows: In a first aspect, the present invention provides a method for identifying key globally regulatory species of disease-related microorganisms, comprising: Step S1: From the public dataset, based on the differences in the prevalence and abundance of species in healthy samples and disease samples, identify the set of health-related species, and construct a species composition abundance table based on the healthy samples; Step S2: Based on the species composition abundance table of the healthy samples, construct a Bayesian network model representing the causal relationship between species under healthy conditions, and obtain the causal network structure and the conditional probability distribution of each species. Step S3: Based on the causal network structure and conditional probability distribution, the influence of each health-related species on the global state of the network is evaluated through simulated intervention. The simulated intervention includes sequentially simulating changes in the state of each target species, and quantifying the impact of the state change on the rest of the network based on the network node importance ranking algorithm and iterative optimization strategy to calculate the intervention score, and determining potential global regulatory species based on the intervention score.
[0008] Optionally, in step S1, the set of health-related species includes healthy prevalent species and healthy scarce species determined by comparing the prevalence ratio and average relative abundance difference of species in healthy samples and disease samples.
[0009] Optionally, step S2 includes: preprocessing the relative abundance data of species in the healthy samples, and constructing a directed acyclic graph network structure of the health-related species based on the preprocessed data using a causal structure learning algorithm.
[0010] Optionally, the preprocessing includes logarithmic transformation and standardization of the relative abundance data of species.
[0011] Optionally, after constructing the directed acyclic graph network structure, discrete latent variables are introduced to model the conditional probability distribution of nodes, wherein the discrete latent variables correspond to multiple Gaussian distribution components, and the expectation-maximization algorithm is used to learn the model parameters.
[0012] Optionally, in step S3, the network node importance ranking algorithm is the HITS algorithm, and the iterative optimization strategy is a greedy strategy.
[0013] Optionally, the intervention score is calculated using the following scoring function:
[0014] in, The intervention score is for a specific target species. The total number of network nodes. For node indexing, To simulate the node after intervention in the target species The change in state For nodes The mean abundance difference between the healthy group and the disease group, where sign(·) is the sign function. For nodes calculated based on the HITS algorithm Hub weights.
[0015] Secondly, the present invention provides a system for identifying key globally regulatory species of disease-related microorganisms, comprising: The data preprocessing module is configured to identify a set of health-related species from the microbiome dataset based on the differences in the prevalence and abundance of species between healthy and disease samples, and to construct a species composition abundance table based on the healthy samples. The causal network modeling module, connected to the data preprocessing module, is configured to construct a Bayesian network model under healthy conditions based on the species composition abundance table, and obtain the causal network structure and conditional probability distribution among health-related species. The global intervention analysis module, connected to the causal network modeling module, is configured to calculate the intervention score of each health-related species by simulating intervention based on the causal network structure and conditional probability distribution, and to determine potential globally regulated species based on the intervention score.
[0016] Optionally, the global intervention analysis module is configured to use the HITS algorithm to calculate the importance weights of network nodes and perform iterative optimization based on a greedy strategy to calculate the intervention score.
[0017] Thirdly, the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method for identifying key global regulatory species of disease-related microorganisms as described above.
[0018] [Beneficial Effects] First, addressing the limitation of traditional methods in distinguishing between causal relationships and co-occurrences among microbial species, this invention constructs a Bayesian network model based on healthy samples and employs discrete latent variables to model species abundance distribution, thereby characterizing potential directed causal driving relationships among species. This makes the identified key species not only associated with disease but also potentially the cause of changes in network state, providing clear causal clues for subsequent mechanism research and targeted intervention, overcoming the shortcomings of existing technologies such as vague targets and weak interpretability.
[0019] Second, addressing the limitation of traditional methods to localized difference analysis, which fail to assess the cascading impact of single-species changes on the entire microbial network, this invention employs simulated intervention analysis, combined with a network node importance ranking algorithm and iterative optimization strategy, to calculate the intervention score for each species. This allows for precise identification of key regulatory species at the core hub level, whose state changes can trigger global reprogramming, from both network topology and dynamic response perspectives. This significantly improves the accuracy and reliability of intervention target selection.
[0020] Third, addressing the issue that microbiome data typically exhibits high dimensionality, sparsity, and non-normal distribution, which can easily lead to unstable model learning, this invention preprocesses species abundance data and introduces a latent variable model based on a Gaussian mixture distribution for probabilistic characterization. This enables the Bayesian network to more flexibly and accurately reflect the true data distribution, thereby learning more robust inter-species conditional dependencies. This lays a solid foundation for subsequent reliable causal inference and intervention analysis.
[0021] Fourth, this invention organically combines the construction of a health benchmark network with global intervention analysis, forming a closed-loop analysis process. This framework can systematically output a high-confidence list of potential regulatory species, and its modular and systematic design facilitates integration into auxiliary diagnostic or microecological therapy target screening platforms, effectively promoting the advancement of microbiome research results into preclinical research and translational medicine applications. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a flowchart illustrating the method for identifying key global regulatory species of disease-related microorganisms provided by the present invention.
[0024] Figure 2 This is a graph showing the performance results of the model provided by this invention in terms of health prediction and discrimination.
[0025] Figure 3 This is an intervention score and cumulative intervention score curve of potential globally regulated species provided by the present invention.
[0026] Figure 4 This is a map of potential global regulatory species that have a high impact on the UC disease network, as discovered in this invention.
[0027] Figure 5This is a graph showing the changes in the health status of the intestinal flora under fecal microbiota transplantation intervention provided by the present invention.
[0028] Figure 6 This is a graph showing the changes in the health status of gut microbiota under different dietary conditions, provided by the present invention.
[0029] Figure 7 This is a diagram showing the changes in the health status of the gut microbiota under antibiotic intervention, provided by the present invention. Detailed Implementation
[0030] The technical solution of the present invention will now be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0031] Example 1 This embodiment provides a method for identifying key globally regulatory species of disease-related microorganisms. The overall process of this method can be found in [link to relevant documentation]. Figure 1 It mainly includes the following three core steps.
[0032] Step S1: This step aims to screen a set of species that are indicative of health status from a publicly available microbiome dataset containing labels for both healthy and disease samples, and to construct a base data table for subsequent modeling.
[0033] First, species relative abundance data and corresponding clinical labels were collected from the publicly available GMWI2 database. To quantitatively characterize the health of the samples, the Gut Microbiota Health Index (GMWI) was introduced. For a sample i, its GMWI value was calculated using the following formula:
[0034] GMWI i refers to samples i GMI value, Let represent the sum of the relative abundances of all healthy, prevalent species in sample i. GMWI represents the sum of the relative abundances of all healthy rare species in sample i. The higher the GMWI value, the closer the microbial composition of the sample is to a healthy state.
[0035] The set of health-related species, comprised of both prevalent and scarce species, is defined based on a comparative analysis of healthy and disease samples. Specifically, it involves calculating the prevalence ratio and average relative abundance difference of each species between healthy and disease samples. To achieve optimal classification, multiple candidate combinations of prevalence ratio and difference thresholds are pre-defined, and health discrimination models are constructed based on the species sets selected from these combinations. By comparing the equilibrium accuracy of each model on the validation set, the species set and its threshold combination corresponding to the optimal equilibrium accuracy are selected as the final classification criteria. Based on this optimized threshold combination, species with high prevalence and significantly higher abundance in healthy samples than in disease samples are classified as prevalent healthy species. Conversely, species with low prevalence and significantly lower abundance in healthy samples compared to disease samples are classified as healthy rare species. .
[0036] Step S2: This step utilizes the health sample data obtained in Step S1 to learn the causal interactions between health-related species under steady-state conditions, constructing a probabilistic graphical model (BN model). The core of this model is to obtain the causal network structure and the conditional probability distribution of each node. See also... Figure 2 The model constructed using this method demonstrates superior performance in distinguishing between healthy and diseased samples.
[0037] S2.1 Data Preprocessing The relative abundance table of health-related species was preprocessed to better align with the modeling assumptions. Preprocessing included logarithmic transformation and standardization, as detailed in the following formulas:
[0038]
[0039] In the formula, Indicates sample i In species j The value after standardization It is a sample i In species j The original relative abundance value, It is a sample i In species j The value after logarithmic transformation. It is a local minimum value, with a value of 1 × 10. -7 , It is a species j The average value of X1, It is a species j The standard deviation of X1.
[0040] S2.2 Learning Causal Network Structure Based on the preprocessed health sample data matrix Y, a Markov Chain Monte Carlo (MCMC) causal structure learning algorithm is used to infer the most probable directed acyclic graph (DAG) network structure among health-related species. This algorithm is a traditional DAG-MCMC implementation, its core being the generation of candidate graph structures in the DAG space by adding, deleting, or reversing single edges, and using an ancestor matrix to ensure the acyclicity of the generated graph. This method iterates through all possible species pairs, selecting the graph with the highest posterior probability from all possible DAGs for each species pair as the optimal connection for that pair, progressively constructing the entire network structure. Specifically, based on a preset node order, the posterior probabilities of three possible connection directions (i.e., from node A to node B, from node B to node A, or no connection) between each pair of nodes are compared sequentially, and the direction with the highest posterior probability is selected as the final connection, iteratively constructing the overall network in this way.
[0041] Key adjustable parameters of the algorithm include: a scoring function (used to evaluate the quality of the graph structure, with Bayesian scoring as the default), the number of samples (defaulted to 100 times the number of nodes), and the burn period (defaulted to 5 times the number of nodes). Through multiple iterations and evaluation using the scoring function, the algorithm optimizes the graph structure and finally outputs a set of sampled graph structures and their related statistics.
[0042] The calculation formula is as follows:
[0043]
[0044] In the formula, Refers to each possible directed acyclic graph d In the sampled image set S The posterior probability in S It is a collection of sampled images. Represents a set S The number of sampled images in the middle d It is the set of all possible directed acyclic graphs. D One of the elements, It is an indicator function, when the sampling graph s With directed acyclic graph d When they are the same, ;otherwise, , This represents the optimal structure selected from all possible directed acyclic graphs, maximizing the posterior probability. The network represents the potential causal relationships between species using sparse directed edges.
[0045] S2.3 Hybrid Model Parameter Learning To address the complex distribution of microbial abundance data and improve model fitting, a common discrete latent variable is introduced into each continuous variable node within the obtained directed acyclic graph structure. This latent variable divides the data into three subpopulations, each modeled using a Gaussian distribution, with the following initial parameter set:
[0046] Furthermore, based on the new causal network structure and the relative abundance data of healthy samples, the EM step is used to calculate the conditional probability (CPD) of each node and obtain the optimal model parameters. This yields a Bayesian inference engine incorporating species information. The E-step calculates the posterior probability distributions of the samples belonging to three Gaussian distributions (latent variables). The M-step updates the model parameters (mean of the Gaussian distribution) based on the latent variable probability weights obtained in the E-step. Covariance Mixing coefficient And the conditional probability distribution parameters in causal networks). The core formula for the EM step is as follows: E-Step:
[0047] In the formula, s Represents the number of iterations. k Representing the k A Gaussian distribution, Indicates the first s In the next iteration, given the observed data Category The posterior probability, In the first s In the nth iteration, the th k The weights or mixing coefficients of the nth Gaussian distribution. It represents the weights or mixing coefficients of the nth Gaussian distribution. k The contribution of each Gaussian distribution to the entire mixture model, and the sum of all weights is 1. Indicates the first s In the next iteration, the observed data In the k The probability density function under a Gaussian distribution, here and They represent the first k The mean vector and covariance matrix of a Gaussian distribution. Indicates all K The sum of weighted probability density functions of Gaussian distributions.
[0048] M-step:
[0049]
[0050]
[0051] In the formula, n This represents the total number of samples. In the M-step, the model parameters are updated iteratively to better fit the data distribution. In the EM-step, this process is repeated until the parameters converge or the preset number of iterations is reached.
[0052] Step S3: This step assesses the global influence of each health-related species in the causal network to identify key regulatory species. The core idea is that a species whose state change triggers the most dramatic cascading effects in the rest of the network is likely to have a global regulatory role. For example... Figure 3 , Figure 4 As shown: In specific implementation, based on the causal network structure and conditional probability distribution obtained in step S2, a simulation intervention analysis method combining the HITS (Hyperlink-Induced Topic Search) algorithm and a greedy optimization strategy is adopted. First, the directed acyclic graph structure of the causal network is transformed into an adjacency matrix, and the HITS algorithm is applied iteratively to calculate the Hub weight of each network node (species). This weight characterizes the importance of the node in the network structure. During the iteration process, the Hub value and Authority value of each node are initialized to 1. Then, the Authority value (the sum of the Hub values of all neighbors pointing to this node) and the Hub value (the sum of the Authority values of all neighbors pointed to by this node) are alternately updated. Normalization is performed after each update until the changes in the Hub value and Authority value between two adjacent iterations are both less than a preset tolerance, at which point convergence occurs. The converged Hub score is taken as the structural importance weight (hub_score) of the node. In a preferred embodiment, the preset tolerance is... .
[0053] Subsequently, simulated interventions were conducted to quantify the global impact of species. Interventions were defined as the vector of the difference in average abundance for each species between the healthy and diseased groups, with the direction representing the ideal intervention direction that would shift the species from a diseased state to a healthy state. During the simulated intervention, each target species was treated as an intervention node, and its state was set to the average abundance in the healthy state (i.e., the corresponding baseline intervention was applied), while the states of the remaining species were considered unknown. Inference was performed based on the conditional probability distribution of the Bayesian network to calculate the expected state change vector (dx) of all other species in the network (i.e., the "remaining part of the network") under this external intervention.
[0054] Based on the above calculations, the intervention scoring function for a single species is defined as: Total Network Score ,in Let k be the total number of network nodes, k be the node index, and sign(·) be the sign function used to determine the direction of the state change of node k. Is it aligned with the overall health goals? Consistent with the previous definition, hub_score(k) represents the HITS Hub weight of node k. This score comprehensively reflects the degree of influence of the target species intervention on the overall network state towards a healthier direction and the importance distribution of this influence in the network structure.
[0055] To determine the priority of key species, a greedy optimization strategy is used to generate an intervention priority ranking: All species are used as the initial candidate set. In each iteration, it is assumed that one species is removed from the current candidate set. The intervention amount corresponding to the "remaining species set after removal" is recalculated, the network state change dx after the intervention is simulated, and the overall network score is calculated. The scores of all candidate species after "removal" are compared, and the species that causes the most significant drop in score is identified as the most critical node, added to the final key species ranking list, and removed from the candidate set. This process is repeated until the candidate set is empty, resulting in a species sequence and its corresponding intervention scores arranged in descending order of global influence. The cumulative intervention influence is calculated based on this sequence. Species contributing the top 80% of the cumulative influence (i.e., species with a cumulative intervention score of 0.8) are typically identified as potential global control species. This threshold (0.8) is an empirical reference value, which can be flexibly adjusted by technicians based on the inflection point of the cumulative intervention score curve in the specific dataset or actual needs. Figure 3 The species intervention scores and cumulative curves calculated based on ulcerative colitis data are presented. Figure 4 This visualizes the core position of potential key species with high intervention scores in the network topology.
[0056] To verify the universality and effectiveness of the model constructed by this method, this embodiment applies it to multiple longitudinal intervention cohort data. The specific steps are as follows: 1. Extract the three longitudinal cohort datasets of FMT treatment, dietary fiber intake changes, and antibiotic interference separately, and extract feature columns according to health-related species. If a species is not in the dataset, fill it with 0.
[0057] 2. Species are sequentially assumed to be in a hidden state. Based on the Bayesian inference engine using healthy samples and the observed values of the remaining health-related species in the entire longitudinal cohort, the health status of each sample is predicted. The generality of the BN model is verified by observing whether it can correctly identify the health status of samples at different time points. The health network model constructed in this invention can track changes in the network's health index under different perturbations. Compared to the GMWI2 study, this model is not only sensitive to known health changes but also able to identify unknown health changes. See details... Figure 5 , Figure 6 and Figure 7 .like Figure 5 As shown, in the fecal microbiota transplantation (FMT) treatment cohort, the model outputs a significantly higher probability of healthy samples as treatment is successful, and it can capture subtle fluctuations in the recovery process. Figure 6 The model demonstrates that it can distinguish the differentiated effects of different dietary fiber interventions on gut microbiota health. Figure 7 The results demonstrate that the model can accurately track the sharp decline and slow recovery of health status caused by antibiotic use. These results confirm that the health network model constructed by this method can not only stably represent health benchmarks, but also sensitively and reliably track the dynamic changes of the microbial community under various perturbations.
[0058] In summary, this embodiment constructs a microbial causal network under healthy conditions and innovatively introduces simulated intervention analysis to quantify the global influence of species, thereby achieving a deeper identification of key regulatory species related to diseases from "correlation" to "causation," and from "local" to "global." This method overcomes the limitation of traditional differential analysis, which can only identify accompanying changes, and can output a list of potential regulatory species with high confidence, providing a new technical means for revealing disease mechanisms and developing targeted microbial ecological intervention strategies.
[0059] Example 2 This embodiment provides a system for identifying key globally regulatory species of disease-related microorganisms, used to implement the method described in Embodiment 1. The system includes the following functional modules: The data preprocessing module is configured to identify a set of health-related species from the microbiome dataset based on the differences in the prevalence and abundance of species between healthy and disease samples, and to construct a species composition abundance table based on the healthy samples. The causal network modeling module, connected to the data preprocessing module, is configured to construct a Bayesian network model under healthy conditions based on the species composition abundance table, and obtain the causal network structure and conditional probability distribution among health-related species. The global intervention analysis module, connected to the causal network modeling module, is configured to calculate intervention scores for each health-related species based on causal network structure and conditional probability distribution, and to identify potential globally regulated species based on these intervention scores. Specifically, the global intervention analysis module uses the HITS algorithm to calculate the importance weights of network nodes and employs a greedy strategy for iterative optimization to calculate the intervention scores.
[0060] Example 3 This embodiment provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method for identifying key global regulatory species of disease-related microorganisms as described in Embodiment 1.
[0061] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for identifying key global regulatory species of disease-associated microorganisms, characterized in that, include: Step S1: From the public dataset, based on the differences in the prevalence and abundance of species in healthy samples and disease samples, identify the set of health-related species, and construct a species composition abundance table based on the healthy samples; Step S2: Based on the species composition abundance table of the healthy samples, construct a Bayesian network model representing the causal relationship between species under healthy conditions, and obtain the causal network structure and the conditional probability distribution of each species. Step S3: Based on the causal network structure and conditional probability distribution, the influence of each health-related species on the global state of the network is evaluated through simulated intervention. The simulated intervention includes sequentially simulating changes in the state of each target species, and quantifying the impact of the state change on the rest of the network based on the network node importance ranking algorithm and iterative optimization strategy to calculate the intervention score, and determining potential global regulatory species based on the intervention score.
2. The method according to claim 1, characterized in that, In step S1, the set of health-related species includes healthy prevalent species and healthy scarce species determined by comparing the prevalence ratio and average relative abundance difference of species in healthy samples and disease samples.
3. The method according to claim 1, characterized in that, Step S2 includes: preprocessing the relative abundance data of species in the healthy samples, and constructing a directed acyclic graph network structure of the health-related species based on the preprocessed data using a causal structure learning algorithm.
4. The method according to claim 3, characterized in that, The preprocessing includes logarithmic transformation and standardization of the relative abundance data of species.
5. The method according to claim 3, characterized in that, After constructing the directed acyclic graph network structure, discrete latent variables are introduced to model the conditional probability distribution of nodes. These discrete latent variables correspond to multiple Gaussian distribution components, and the expectation-maximization algorithm is used to learn the model parameters.
6. The method according to claim 1, characterized in that, In step S3, the network node importance ranking algorithm is the HITS algorithm, and the iterative optimization strategy is a greedy strategy.
7. The method according to claim 6, characterized in that, The intervention score is calculated using the following scoring function: in, The intervention score is for a specific target species. The total number of network nodes. For node indexing, To simulate the post-intervention node of the target species The change in state For nodes The mean abundance difference between the healthy group and the disease group, where sign(·) is the sign function. For nodes calculated based on the HITS algorithm Hub weights.
8. A system for identifying key globally regulatory species of disease-associated microorganisms, characterized in that, include: The data preprocessing module is configured to identify a set of health-related species from the microbiome dataset based on the differences in the prevalence and abundance of species between healthy and disease samples, and to construct a species composition abundance table based on the healthy samples. The causal network modeling module, connected to the data preprocessing module, is configured to construct a Bayesian network model under healthy conditions based on the species composition abundance table, and obtain the causal network structure and conditional probability distribution among health-related species. The global intervention analysis module, connected to the causal network modeling module, is configured to calculate the intervention score of each health-related species by simulating intervention based on the causal network structure and conditional probability distribution, and to determine potential globally regulated species based on the intervention score.
9. The system according to claim 8, characterized in that, The global intervention analysis module is configured to use the HITS algorithm to calculate the importance weights of network nodes and perform iterative optimization based on a greedy strategy to calculate the intervention score.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the method for identifying key global regulatory species of disease-associated microorganisms as described in any one of claims 1 to 7.