Intestinal microorganism metabolome multi-disease marker screening system
By employing modules such as batch calibration, covariance analysis, constraint network construction, and stability assessment, the system addresses the issues of consistency, stability, and redundancy in the joint analysis of gut microbiota and metabolomics, thereby improving the efficiency and accuracy of disease biomarker screening and optimizing the diagnostic performance for multiple diseases.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 中国人民解放军总医院第八医学中心
- Filing Date
- 2026-04-27
- Publication Date
- 2026-06-26
AI Technical Summary
Existing combined analysis of gut microbiota and metabolomics in disease biomarker screening suffers from batch effects, insufficient biomarker stability, functional redundancy, and uneven disease discrimination, resulting in poor data consistency, low reproducibility, high cost, and low diagnostic specificity.
The system employs a batch correction module to eliminate batch effects, a covariance analysis module to assess stability, a constrained network construction module to utilize metabolic pathway topology, a biomarker screening module to integrate network topology and disease discriminative power, a functional redundancy dimensionality reduction module to optimize biomarker sets, a stability assessment module to ensure ecological stability, and a cross-disease optimization module to output multiple disease biomarker panels.
It improves the consistency and replication rate of multicenter study data, identifies key biomarkers, reduces the number of biomarkers, enhances diagnostic stability and specificity, optimizes the balance between testing costs and information acquisition, and achieves synergistic optimization of multi-disease discrimination.
Smart Images

Figure CN122290710A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multi-omics data integration technology, and more specifically, to a system for screening multi-omics disease biomarkers of gut microbiota metabolomics. Background Technology
[0002] Currently, the combined analysis of gut microbiota and metabolomics faces several challenges in the screening of disease biomarkers. The batch effect, prevalent in non-targeted metabolomics testing, makes direct comparison of data from different experimental batches difficult, affecting the reliability and consistency of multicenter study results. Traditional microbe-metabolite association analyses often rely on simple Pearson or Spearman correlation coefficients, lacking assessment of association robustness, resulting in low biomarker reproducibility in validation cohorts and hindering clinical application. Existing screening methods often overlook rich prior knowledge of metabolic pathways, leading to selected biomarkers that, while statistically significant, lack biological interpretability, making it difficult for clinicians to understand their disease mechanism significance. Furthermore, biomarker screening methods that ignore metabolic network topology fail to identify important biomarkers located at key nodes in the metabolic network, focusing only on single-point differences while ignoring systemic changes. In multi-disease studies, the problem of significant functional redundancy among biomarkers is particularly prominent, increasing detection costs while limiting information gain. In clinical applications, the variability of biomarkers in healthy individuals is often overlooked, making many biomarkers susceptible to influences such as dietary habits, environmental factors, and even circadian rhythms, reducing diagnostic specificity. When dealing with various digestive system diseases such as inflammatory bowel disease, irritable bowel syndrome, and colorectal cancer, existing methods struggle to balance the discriminative power of different disease groups. Improving the diagnostic rate of one disease often comes at the cost of decreased discriminative power for others. The determination of biomarker quantities lacks scientific basis; clinical researchers often subjectively set these numbers based on experience or experimental costs, failing to find the optimal balance between information content and testing costs.
[0003] In view of this, the present invention proposes a multi-group disease biomarker screening system for gut microbiota metabolomics to solve the above problems. Summary of the Invention
[0004] To overcome the aforementioned deficiencies of the prior art and to achieve the above objectives, the present invention provides the following technical solution: a multi-group disease biomarker screening system for gut microbiota metabolomics, comprising: The batch correction module is used to acquire gut microbial species abundance data and untargeted metabolomics mass spectrometry raw data from multiple groups of disease subjects and healthy control groups, and to perform batch-to-batch drift correction based on the internal reference metabolite comparison value on the raw metabolomics mass spectrometry data of different detection batches, generating a corrected metabolite abundance matrix. The covariance analysis module is used to calculate the cross-sample abundance covariance trajectory between each metabolite and each microbial species for each metabolite in the corrected metabolite abundance matrix, and extract the directional consistency ratio of the covariance trajectory, which is denoted as the covariance stability index. The constraint network construction module is used to calculate the pathway topological distance between metabolite pairs based on the number of enzymatic reaction steps between metabolites in the metabolic pathway database, and to construct a pathway-constrained metabolite association network by combining the covariant stability index. The biomarker screening module is used to calculate the pathway constraint betweenness centrality for each metabolite node in the pathway constraint metabolite association network, and fuse it with the standardized effect size of the metabolite among the disease groups to obtain the network enhancement discrimination score. The candidate biomarker set is then screened based on the network enhancement discrimination score. The functional redundancy dimensionality reduction module is used to compare the functional annotations of each pair of metabolites in the candidate biomarker set to the corresponding metabolic pathways, calculate the functional coverage overlap, and perform stepwise elimination according to the maximum coverage and minimum redundancy criteria to generate a simplified biomarker set. The stability assessment module is used to calculate the coefficient of variation of the abundance of members within the ecological functional module composed of each metabolite and its associated microbial species in the simplified biomarker set in a healthy control group sample, and is recorded as the ecological stability baseline value. The cross-disease optimization output module is used to perform joint optimization of the discriminative power of the simplified biomarker set across disease groups, with the ecological stability baseline value meeting the preset stability threshold as a constraint, and output the final multi-group disease biomarker panel.
[0005] The technical effects and advantages of the intestinal microbiome metabolomics multi-protocol disease biomarker screening system of this invention: This invention eliminates batch effect interference, ensuring the consistency and comparability of multi-center study data and improving the repeatability of biomarkers across different cohorts. A stability assessment mechanism effectively distinguishes between accidental statistical associations and genuine biological associations, addressing the issue of insufficient biomarker stability in traditional methods. This invention fully utilizes the topological characteristics of metabolic networks, enabling the identification of key biomarkers located at metabolic regulatory hubs, rather than relying solely on single-point statistical differences. The functional redundancy dimensionality reduction method increases the information density of the biomarker set, significantly reducing the number of required biomarkers while maintaining diagnostic performance, thus optimizing the balance between detection costs and information acquisition. The ecological stability assessment mechanism effectively eliminates candidate biomarkers highly susceptible to environmental fluctuations, enhancing the stability and reliability of diagnostic results. In multi-disease diagnostic scenarios, this invention achieves synergistic optimization of the discriminative power of various diseases, avoiding the problem of sacrificing the discriminative power of other diseases while improving the diagnostic rate of one disease. Dynamic learning capabilities enable the system to continuously optimize its analysis strategy from historical data, resulting in virtuous cycle improvement. Attached Figure Description
[0006] Figure 1 This is a schematic diagram of the intestinal microbiome multi-group disease biomarker screening system of the present invention. Detailed Implementation
[0007] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0008] This application provides a system for screening multiple disease biomarkers based on the gut microbiota metabolome. The system's execution entities include, but are not limited to, those mounted on the system: a metabolomics analysis platform, a microbiome data processing center, a clinical biomarker screening system, and a multi-omics integrated analysis platform, which can be considered as general computing nodes in this application.
[0009] Please see Figure 1 In this embodiment of the invention, the gut microbiota metabolomics multi-group disease biomarker screening system includes: The batch correction module acquires gut microbial species abundance data and untargeted metabolomics mass spectrometry raw data from multiple disease subjects and healthy controls. It performs inter-batch drift correction based on internal reference metabolite comparison values on raw metabolomics mass spectrometry data from different batches, generating a corrected metabolite abundance matrix. This module first simultaneously acquires gut microbial sequencing data and metabolomics mass spectrometry data from multiple disease and healthy control samples. Then, addressing the batch effect problem commonly encountered in untargeted metabolomics, it uses the internal reference metabolite comparison value method for correction. By screening for internal reference metabolite pairs stably expressed in all batches, inter-batch correction factors are calculated to eliminate systematic biases between different batches. This correction method based on internal reference ratios does not rely on external standards; it utilizes the inherent structure of the data for standardization, ensuring the comparability and consistency of metabolite abundance data and providing a high-quality basic data matrix for subsequent analysis.
[0010] The covariance analysis module calculates the cross-sample abundance covariance trajectory between each metabolite and each microbial species in the corrected metabolite abundance matrix, extracting the directional consistency ratio of the covariance trajectory, denoted as the covariance stability index. This module reveals the potential functional associations between metabolites and microbial species by analyzing the synergistic patterns of abundance changes. The module uses a random sampling strategy to generate multiple sample subsets, assesses the sign consistency between metabolite and microbial abundance changes in each subset, and quantifies the stability of this covariance relationship using statistical methods. The covariance stability index not only reflects the strength of the association but, more importantly, assesses its robustness, providing a reliable edge weight indicator for constructing a microbial-metabolite functional network.
[0011] The constraint network construction module calculates the pathway topological distance between metabolite pairs based on the number of enzymatic reaction steps between metabolites in the metabolic pathway database, and constructs a pathway-constrained metabolite association network by combining this with a covariant stability index. This module combines prior biological knowledge with data-driven association analysis to build a pathway-constrained metabolite association network. By tracking the number of enzymatic reaction steps between metabolites in the metabolic pathway database, the pathway topological distance between metabolite pairs is calculated as a biological constraint for network construction. When the pathway distance between metabolite pairs meets a threshold condition, an association edge is established for that pair of metabolites based on the microbial covariant stability index, forming a constraint network that conforms to both the biological pathway structure and reflects the association patterns of experimental data, providing a structured network foundation for biomarker screening.
[0012] The biomarker screening module calculates the pathway constraint betweenness centrality for each metabolite node in the pathway-constrained metabolite association network and fuses it with the normalized effect size of the metabolite across disease groups to obtain a network enhancement discrimination score. A set of candidate biomarkers is then selected based on this score. This module integrates network topology characteristics and disease discrimination ability to screen candidate biomarkers with dual significance. First, the betweenness centrality of each metabolite in the constraint network is calculated to measure its importance as an information flow hub. Then, the normalized effect size of the metabolite across different disease groups is combined to assess its ability to distinguish diseases. Finally, the network enhancement discrimination score is obtained by multiplying these two indicators, and the set of candidate biomarkers is selected based on the score, ensuring that the selected biomarkers occupy a key position in the metabolic network and possess significant disease discrimination ability.
[0013] The functional redundancy dimensionality reduction module compares the functional annotations of each metabolite pair within the candidate biomarker set to their respective metabolic pathways, calculates functional coverage overlap, and performs a stepwise elimination process based on the maximum coverage and minimum redundancy criterion to generate a streamlined biomarker set. This module eliminates redundancy in the candidate biomarker set through functional annotation comparison, optimizing the functional coverage of the biomarker combinations. For each pair of metabolites, the metabolic pathway functional annotations are compared, and the Jaccard similarity coefficient is calculated as the functional coverage overlap. Then, metabolites with low functional redundancy are progressively selected according to their discrimination scores from highest to lowest, ensuring that the final biomarker set minimizes information redundancy while maintaining maximum functional coverage. This functional annotation-based dimensionality reduction method preserves biological meaning, avoids the potential reduction in interpretability caused by purely statistical dimensionality reduction, and provides a functionally clear and streamlined biomarker set for subsequent stability assessments.
[0014] The stability assessment module is used to calculate the coefficient of variation (COP) of the abundance of each member within an ecofunctional module composed of each metabolite and its associated microbial species in a simplified biomarker set, using a healthy control group sample. This COP is recorded as the ecostability baseline value. This module assesses the ecostability of biomarkers, ensuring that the selected biomarkers remain relatively stable under healthy conditions. The module first identifies the associated microbial species set for each metabolite biomarker, constituting an ecofunctional module. Then, it calculates the COP of the abundance of each member within this module in a healthy control group sample and uses a weighted average of the COP to derive the ecostability baseline value. This index measures the natural fluctuation range of the biomarker-related ecofunctional module under healthy conditions, providing stability constraints for cross-disease optimization. This ensures that the finally selected biomarkers can distinguish diseases and remain relatively stable under healthy conditions, improving the biological reliability of the biomarkers.
[0015] The cross-disease optimization output module performs joint optimization of the discriminative power of a simplified biomarker set across disease groups, using the ecological stability baseline value meeting a preset stability threshold as a constraint, and outputs a final panel of multiple disease biomarkers. This module optimizes the multi-disease discriminative performance of biomarker combinations under ecological stability constraints. First, metabolites with ecological stability baseline values exceeding the threshold are removed to ensure that the selected biomarkers have sufficient stability in a healthy state. Then, a multi-class discriminative model is constructed, employing a forward stepwise selection strategy, using the comprehensive discriminative performance of multiple disease groups as the objective function, to progressively select the optimal biomarker. The marginal contribution of each newly added biomarker to the discriminative sensitivity of each disease group is evaluated, and the selection process terminates when the contribution falls below the threshold. The final output panel of multiple disease biomarkers satisfies the ecological stability requirements and achieves efficient differentiation of multiple diseases, providing practical value for clinical applications.
[0016] The modules are connected via wired and / or wireless means to enable data transmission between them.
[0017] In this embodiment of the invention, the detailed implementation steps for performing inter-batch drift correction based on the internal reference metabolite contrast value on raw metabolomics mass spectrometry data from different detection batches to generate a corrected metabolite abundance matrix include: Metabolites with detection rates greater than a preset detection threshold and coefficients of variation less than a preset variation threshold are selected from all test batches and marked as candidate internal control metabolites. This step is crucial for identifying stably expressed metabolites and selecting reliable batch calibration references. The screening process first calculates the detection rate of each metabolite in each batch, i.e., the proportion of samples that can be detected, typically setting a detection threshold of 90% or higher; simultaneously, it calculates the coefficient of variation within each batch, i.e., the ratio of standard deviation to mean, reflecting the dispersion of abundance, typically setting a variation threshold below 30%. Metabolites that meet both conditions are marked as candidate internal control metabolites. These metabolites have high detection stability and low abundance variability, serving as reliable reference points for batch calibration and providing a stable basis for subsequent ratio calculations.
[0018] Candidate internal control metabolites are paired, and the stability coefficient of the abundance ratio of each pair is calculated across batches. The abundance ratio method is an effective strategy to eliminate batch effects by calculating the relative abundance relationship of metabolite pairs, thus reducing the impact of fluctuations in absolute abundance measurements. The pairing process combines all candidate internal control metabolites in pairs to form a set of candidate internal control metabolite pairs; then, the abundance ratio of each pair of metabolites is calculated in each sample of each batch, forming a batch-specific ratio distribution; finally, the stability coefficient of these ratio distributions across batches is evaluated. The stability coefficient calculation considers the ratio of intra-batch variance to inter-batch variance. Stable internal control pairs should maintain a similar ratio distribution across different batches, i.e., the intra-batch variance is greater than the inter-batch variance. A higher stability coefficient indicates that the metabolite pair is more suitable as a calibration benchmark.
[0019] The top K pairs of candidate internal control metabolites with the highest stability coefficients are selected and denoted as anchored internal control pairs. The selection of anchored internal control pairs is a core step in batch calibration, determining the most reliable combination of calibration benchmarks. The selection process first ranks all candidate internal control metabolite pairs in descending order of stability coefficient, and then selects the top K pairs as anchored internal control pairs. The value of K needs to balance calibration stability and computational complexity, and is typically set to 5-10. Using multiple pairs of internal controls instead of a single pair improves calibration robustness and reduces the impact of random fluctuations. The selection of anchored internal control pairs fully considers the biological characteristics of the metabolites and batch-to-batch technical variations, ensuring the reliability and consistency of subsequent calibration processes.
[0020] Using the median abundance ratio of the anchored internal control pairs as a benchmark, the scaling factor of each batch relative to the benchmark batch is calculated. Scaling factor calculation is a crucial step in achieving batch-to-batch standardization, quantifying systematic bias between different batches. The calculation process first selects a batch as the benchmark batch, typically choosing the batch with the largest sample size or highest quality; then, the median abundance ratio of each anchored internal control pair in each batch is calculated; next, the ratio deviation of each batch relative to the benchmark batch is calculated; finally, the batch scaling factor is determined by the combined bias of the anchored internal control pairs. The scaling factor calculation formula is: ; Among them, SF b Where K is the scaling factor for batch b, K is the number of anchored intrinsic parameter pairs, and M is the scaling factor for batch b. i,b M is the median abundance ratio of the i-th anchorage pair in batch b. i,ref This represents the median abundance ratio of the i-th anchorage pair in the baseline batch. The scaling factor directly reflects the systematic bias between batches, providing a quantitative basis for global correction.
[0021] Linear correction was performed on the abundance of all metabolites in each batch using a scaling factor, generating a corrected metabolite abundance matrix. Linear correction is the final step in eliminating batch effects; by systematically adjusting metabolite abundance values, it ensures comparability between data from different batches. The correction process applies the corresponding scaling factor to all metabolites in each batch, adjusting their abundance values. The correction formula is: ; in, To correct the abundance value of metabolite i in sample j of batch b, X i,j,b The original abundance values before correction, SF b is the scaling factor for batch b. The corrected metabolite abundance matrix eliminates systematic differences between batches while preserving biological variations among samples, providing a high-quality data foundation for subsequent covariance analysis and network construction, and ensuring the reliability of multi-batch data integration analysis.
[0022] In this embodiment of the invention, the detailed implementation steps for calculating the stability coefficient of the abundance ratio of each pair of candidate internal control metabolites across batches include: For each pair of candidate internal control metabolites, the abundance ratio of the two metabolites is calculated for each sample within each testing batch, forming a ratio vector for that batch. Constructing the ratio vector is a fundamental step in assessing the stability of internal control metabolite pairs, eliminating the influence of absolute abundance fluctuations by calculating the relative abundance relationship. The construction process first calculates the abundance ratio of the candidate internal control metabolite pair for each sample within each batch, i.e., the abundance of the first metabolite divided by the abundance of the second metabolite; then, all sample ratios within the same batch are combined into a single vector, forming a batch-specific ratio distribution. This ratio-based method effectively captures the relative relationships between metabolites, which are generally more stable than absolute abundance and better reflect systematic biases between batches, providing a reliable data foundation for subsequent stability assessments.
[0023] Calculate the between-batch and within-batch variances of the ratio vector across batches. Analysis of variance (ANOVA) is a key statistical method for assessing ratio stability, quantifying stability by comparing the degree of variation between and within batches. The analysis process first calculates the variance of the ratio vector within each batch, reflecting the natural biological variation within the batch; then, it calculates the variance of the median of the ratios between different batches, reflecting the technical deviation between batches; finally, it compares the relative magnitudes of these two variances to assess the batch stability of the ratio. A stable internal reference should be characterized by within-batch variance greater than between-batch variance, i.e., biological variation exceeds technical deviation, indicating that the ratio is not significantly affected by batch effects. This evaluation method based on variance decomposition conforms to statistical principles and provides a theoretical basis for calculating the stability coefficient.
[0024] The stability coefficient is calculated by dividing the within-batch variance by the between-batch variance. The stability coefficient is a core indicator for quantifying the reliability of internal parameters, and it visually represents batch stability through the variance ratio. The calculation formula is: ; Among them, SC i,j V is the stability coefficient of the internal reference pair composed of metabolites i and j. within (R i,j V represents the average variance of the internal reference comparison value across all batches. between (R i,j R represents the variance of the internal reference comparison value across batches. i,j This represents the abundance ratio of metabolites i to j. A stability coefficient greater than 1 indicates that intra-batch variability is greater than inter-batch variability, and the ratio remains relatively stable across batches; the higher the coefficient, the more suitable the internal control pair is for batch calibration. This indicator directly reflects the ability of the internal control pair to resist batch effects, providing a quantitative basis for selecting the optimal combination of internal controls.
[0025] When the stability coefficient exceeds a preset stability screening threshold, the candidate internal control metabolite pair is retained. Final screening is a crucial measure to ensure the quality of the internal control; threshold screening ensures the reliability of the calibration benchmark. The screening process involves setting an appropriate stability screening threshold (typically 1.5-2.0) to retain only candidate internal control pairs with stability coefficients exceeding the threshold—that is, metabolite pairs with intra-batch variance significantly greater than inter-batch variance. These highly stable internal control pairs effectively resist batch effects, maintain consistency across batches, and provide a reliable benchmark for subsequent batch calibration. The threshold setting must balance rigor and practicality, ensuring that a sufficient number of candidate internal control pairs are selected while maintaining sufficiently high stability to guarantee successful batch calibration.
[0026] In this embodiment of the invention, the detailed implementation steps for calculating the cross-sample abundance covariance trajectory between the covariance trajectory and each microbial species, and extracting the directional consistency ratio of the covariance trajectory as the covariance stability index include: After randomly shuffling all samples within each disease group, a predetermined number of subset sequences are drawn. Subset selection is a fundamental step in assessing covariant stability, evaluating the robustness of associations through multiple random samplings. The process begins by randomly shuffling samples within the same disease group while maintaining the proportion of samples from each group. Then, multiple subsets are drawn, each time representing representative samples from all disease groups, typically with a selection ratio of 60%-80% of the total sample. Finally, multiple subset sequences are formed for subsequent covariant analysis. The predetermined number of samplings is usually set between 100 and 500; a larger number of samplings results in greater stability but also increases computational complexity. This random sampling-based method can simulate natural data fluctuations, effectively assess the robustness of microbe-metabolite associations, reduce the impact of random associations, and provide diverse test datasets for calculating covariant stability indices.
[0027] For each extracted subset of sample sequences, the sign consistency of the changes in the abundance of the target metabolite and the target microbial species is calculated. Sign consistency is the proportion of sample pairs in which the abundance of both increases or decreases simultaneously out of all sample pairs. Sign consistency analysis is a simplified method for capturing covariance patterns, assessing the strength of association through the consistency of the direction of change. The analysis process first organizes sample pairs in temporal or spatial order within the sample subset; then, it determines the direction of change in the abundance of metabolites and microorganisms in each sample pair, recording whether they change in the same direction (simultaneous increase or simultaneous decrease); finally, it calculates the proportion of sample pairs changing in the same direction in the total sample pairs, serving as the sign consistency index. The sign consistency value ranges from [0,1], where a value of 0.5 indicates no association, greater than 0.5 indicates a positive correlation, less than 0.5 indicates a negative correlation, and extreme values of 1 or 0 indicate complete consistency or complete opposites. This method, based on signs rather than precise numerical values, is insensitive to outliers and can robustly capture major covariance trends, providing a reliable basic measure for stability assessment.
[0028] The mean and standard deviation of sign consistency across a predetermined number of sampling iterations are calculated. Statistical analysis is a crucial step in quantifying covariant stability, assessing the robustness of an association by integrating the results of multiple samplings. The analysis process summarizes all sign consistency values obtained from the sampling and calculates their mathematical mean and standard deviation. The mean reflects the overall strength and direction of the association, while the standard deviation reflects the degree of fluctuation of the association across different sample subsets. Stable associations should exhibit a high mean (far from 0.5) and a low standard deviation, i.e., a consistent association pattern across most samplings. This statistical method based on multiple sampling effectively distinguishes between stable and random associations, providing a statistical basis for calculating the covariant stability index.
[0029] Dividing the mean of the sign consistency by the standard deviation yields the covariance stability index. The covariance stability index is a core indicator for quantifying the robustness of microbial-metabolite associations, comprehensively considering both the strength and consistency of the association. The calculation formula is: ; Among them, CSI m,t μ(SC) is the covariant stability index of microorganism m and metabolite t. m,t ) represents the mean of sign consistency across all samples, σ(SC) m,t ) represents the standard deviation of sign consistency, SC m,t This index represents the set of symbol consistency values. Formally similar to the signal-to-noise ratio, the mean represents the "signal" strength, and the standard deviation represents the "noise" level. A higher index indicates a more stable and reliable association. Microbe-metabolite pairs with high covariance stability typically represent genuine biological associations, rather than accidental correlations caused by data fluctuations. This provides a reliable weighting indicator for constructing robust microbe-metabolite functional networks and forms the core foundation for subsequent network analysis.
[0030] In this embodiment of the invention, the detailed implementation steps for calculating the pathway topological distance between metabolite pairs based on the number of enzymatic reaction steps between each metabolite in the metabolic pathway database, and constructing a pathway-constrained metabolite association network in combination with the covariant stability index, include: In metabolic pathway databases, a global directed graph of metabolic reactions is constructed, using metabolites as nodes and enzymatic reactions as edges. Constructing this global directed graph is a fundamental step in network analysis, establishing functional connections between metabolites. The construction process utilizes public metabolic pathway databases (such as KEGG and MetaCyc) to extract information on metabolites and their enzymatic reactions. Each metabolite is designated as a node in the graph, and each enzymatic reaction as a directed edge, with the edge direction following the substrate-to-product direction. For reversible reactions, bidirectional edges are used. The constructed global graph typically contains thousands of nodes and tens of thousands of edges, comprehensively covering known metabolic networks. This pathway-based network construction method introduces prior biological knowledge into data analysis, providing the biological background and constraints for metabolite associations and laying a graph theory foundation for subsequent topological distance calculations.
[0031] For each pair of metabolites in the corrected metabolite abundance matrix, the shortest path length between them is calculated in the global metabolic response directed graph, denoted as the pathway topological distance. Calculating the pathway topological distance is a crucial step in quantifying the functional relationships between metabolites, using graph theory algorithms to determine the biochemical transformation distances between them. The calculation process utilizes shortest path algorithms such as Dijkstra's or Floyd-Warshall to calculate the shortest connection path between each pair of metabolites in the global metabolic response directed graph; the path length represents the number of edges traversed, signifying the minimum number of enzymatic reaction steps required to transform one metabolite into another. Metabolite pairs with small topological distances typically reside in the same or adjacent metabolic pathways, exhibiting more direct biochemical transformation relationships; metabolite pairs with large distances have more distant metabolic relationships and may belong to different functional modules. This graph theory-based distance calculation quantifies the relationships between metabolites into discrete values, providing biologically meaningful constraints for network construction.
[0032] When the pathway topological distance is less than a preset topological distance threshold, edges are established between metabolite pairs in the pathway-constrained metabolite association network. Pathway-constrained edge establishment is a crucial step in constructing a biologically sound network, using topological distance screening to establish functionally relevant connections. The establishment process sets an appropriate topological distance threshold (typically 2-4 steps of response distance), establishing connections only between metabolite pairs whose pathway distance is less than this threshold. This constraint ensures that the connected metabolites in the network genuinely have close biochemical metabolic relationships, rather than being arbitrary connections based solely on statistical correlation. The introduction of pathway constraints significantly reduces network complexity and false positive connections, making the network structure more biologically consistent, improving the interpretability and reliability of subsequent analyses, and providing a high-quality network foundation for biomarker screening.
[0033] The weight of each edge is set to the mean of the product of the covariance stability indices of the metabolites with respect to their respective sets of associated microbial species. Edge weight assignment is a crucial step in integrating multi-source information, incorporating microbial association data into the metabolite network. The assignment process first determines the set of microbial species associated with each metabolite node, typically selecting the species with the highest covariance stability indices; then, it calculates the product of the covariance stability indices between the respective microbial sets of two metabolites; finally, it takes the mean of these products as the edge weight. The weight calculation formula is: ; Among them, W i,j M represents the weight of the edge between metabolites i and j. i and M j CSI represents the sets of microbial species associated with metabolites i and j, respectively. m,i Let be the covariant stability index between microorganism m and metabolite i. This weight design considers both the strength of the association between metabolites and microorganisms and reflects the synergistic effects among microbial communities, providing edge weights with both biological significance and data support for network analysis, thereby enhancing the information content and analytical value of the network.
[0034] In this embodiment of the invention, the detailed implementation steps for calculating the pathway constraint betweenness centrality for each metabolite node and fusing it with the standardized effect size of the metabolite across disease groups to obtain the network enhancement discrimination score include: In a pathway-constrained metabolite association network, the proportion of all shortest paths passing through each metabolite node to the total number of shortest paths in the network is denoted as the pathway-constrained betweenness centrality. Betweenness centrality calculation is a crucial step in evaluating the importance of a node network, quantifying the value of a node as an information flow hub. The calculation process first identifies the shortest paths between all node pairs in the pathway-constrained network; then, it counts the frequency with which each node appears as an intermediate node on these paths; finally, it divides this frequency by the total number of shortest paths in the network to obtain the standardized betweenness centrality value. The calculation formula is: ; in, For nodes betweenness centrality, σ st This represents the total number of shortest paths from node s to node t. For the journey from s to t, passing through nodes The number of shortest paths. Metabolite nodes with high betweenness centrality are typically located at the intersection of multiple metabolic pathways, controlling metabolic flow and information transmission within the network, and may have broader regulatory roles in biological systems. This indicator quantifies the functional importance of metabolites as a network topological property, providing a systems biology perspective for biomarker screening.
[0035] For each metabolite, its standardized effect size was calculated between each disease group and the healthy control group. Standardized effect size calculation is a crucial step in assessing disease discriminative power, quantifying the degree of variability of metabolites across disease states. The calculation process involved comparing each disease group with the healthy control group, using effect sizes such as Cohen's d or Hedge's g to standardize inter-group differences in metabolite abundance. Standardized effect sizes consider both inter-group and intra-group variability, providing a more comprehensive assessment of discriminative power than simple p-values or fold changes. Calculating effect sizes across different disease groups allows for cross-disease comparisons, identifying metabolites sensitive to specific or multiple diseases, and providing a quantitative basis for multi-disease biomarker screening. The standardization of effect sizes eliminates dimensional differences between metabolites, making the discriminative power of each metabolite comparable and laying a statistical foundation for subsequent fusion analyses.
[0036] The maximum absolute value of the standardized effect size is recorded as the maximum cross-group effect size for that metabolite. Extracting the maximum cross-group effect size is a crucial step in determining discriminative potential, identifying the optimal discriminative scenario for a metabolite. For each metabolite, the extraction process compares the absolute values of its standardized effect sizes across all disease groups and healthy control groups, selecting the maximum value as the metabolite's characteristic indicator. Choosing the maximum absolute value, rather than the average, focuses on discovering metabolites with strong discriminative power for specific diseases, rather than disease-wide biomarkers with moderate discriminative power. This strategy aligns with the principles of precision medicine, identifying disease-specific candidate biomarkers while allowing the discovery of shared biomarkers with high discriminative power across multiple diseases. The maximum cross-group effect size, as a representative value of a metabolite's discriminative potential, provides the core input for calculating the disease discriminative dimension in the network-enhanced discriminative score.
[0037] The product of pathway constraint betweenness centrality and maximum cross-group effect size is used as the network enhancement discrimination score. The network enhancement discrimination score is an innovative indicator that integrates network topology and disease discrimination, comprehensively evaluating the dual value of metabolites as biomarkers. The calculation formula is as follows: ; in, For nodes The network enhancement discrimination score, For path constraint betweenness centrality, This represents the maximum cross-group effect size. This product-based fusion strategy emphasizes the importance of two dimensions: a metabolite can only achieve a high score if it is critical to its network position and has significant disease discriminative power. The network enhancement design overcomes the limitations of traditional biomarker selection based solely on statistical differences, introducing a network perspective from systems biology. This ensures that the selected biomarkers are not only statistically significant but also biologically interpretable, providing a comprehensive evaluation criterion for subsequent screening and improving the bioreliability and clinical translation potential of the biomarkers.
[0038] In this embodiment of the invention, the detailed implementation steps for calculating functional coverage overlap and performing stepwise elimination according to the maximum coverage and minimum redundancy criterion to generate a simplified set of markers include: For each metabolite in the candidate biomarker set, all pathway functional categories it participates in within metabolic pathway databases are extracted, constructing a functional annotation set for that metabolite. Constructing the functional annotation set is a fundamental step in functional analysis, establishing the association between metabolites and biological functions. The construction process utilizes metabolic pathway databases (such as KEGG and MetaCyc) to extract information on all metabolic pathways involved by each candidate biomarker metabolite; these pathways are then categorized according to their biological functions, such as energy metabolism, amino acid metabolism, and lipid metabolism; finally, a functional annotation set corresponding to each metabolite is formed, reflecting the scope of biological processes involved. This pathway-based functional annotation method places metabolites within a broader biological context, revealing their potential physiological roles and disease-related mechanisms, providing a functional-level explanatory framework for dimensionality reduction analysis, and ensuring that biomarker selection is based not only on statistical characteristics but also on biological significance.
[0039] For each pair of candidate biomarker metabolites, the Jaccard similarity coefficient of their functional annotation sets is calculated, denoted as the functional coverage overlap. Calculating the functional coverage overlap is a crucial step in assessing functional redundancy and quantifying the functional similarity between metabolites. The Jaccard similarity coefficient, a classic measure of set similarity, is used to calculate the ratio of the intersection size to the union size of the two metabolite functional annotation sets. The formula is: ; Where J(A,B) represents the functional coverage overlap between metabolites A and B, where A and B are the functional annotation sets for the two metabolites, respectively. Let the size of the intersection of the two sets be denoted by . The value represents the union size. The overlap value ranges from [0,1], where 0 indicates complete functional complementarity and 1 indicates complete functional overlap. This indicator intuitively reflects the degree of functional redundancy between metabolites, providing a quantitative basis for subsequent redundancy elimination and serving as the core computational foundation for achieving the principle of maximum coverage and minimum redundancy.
[0040] Metabolites are selected into the simplified biomarker set in descending order of their network enhancement discrimination scores. Prioritizing high-scoring metabolites is the first step in constructing a high-quality biomarker set, ensuring that core biomarkers possess optimal network and discrimination characteristics. The selection process first sorts candidate biomarkers in descending order of their network enhancement discrimination scores, forming a priority queue; then, starting from the top of the queue, each metabolite is examined to determine its suitability for inclusion in the simplified set. This score-based priority strategy ensures that the most important biomarkers are considered first, guaranteeing the core quality of the simplified set. Simultaneously, the rigorous sorting mechanism ensures good determinism and repeatability of the algorithm results, avoiding the instability that may arise from random selection and providing a clear processing order for subsequent functional redundancy screening.
[0041] Each time a new metabolite is selected, its functional coverage overlap with all previously selected metabolites is checked. If any functional coverage overlap exceeds a preset overlap threshold, the metabolite is skipped. Redundancy checking is a crucial step in achieving minimum redundancy, preventing functionally redundant biomarkers from entering the reduced set. The checking process calculates the functional coverage overlap of each candidate metabolite with all metabolites already selected in the reduced set for each candidate metabolite. An appropriate overlap threshold is set (typically 0.5-0.7). If any overlap exceeds the threshold, the metabolite is considered to provide significantly redundant functional information and is skipped. This threshold-based redundancy elimination strategy effectively reduces information redundancy in the biomarker set, ensuring that the final selected biomarkers are complementary rather than repetitive, greatly improving information density and analytical efficiency while maintaining the diversity of biological interpretations. This provides functionally complementary biomarker combinations for the comprehensive differentiation of multiple diseases.
[0042] The selection process terminates when the number of pathway functional categories covered by the union of all functional annotations in the reduced biomarker set reaches a preset coverage ratio. Coverage checking is the termination condition for achieving maximum coverage, ensuring the functional integrity of the reduced set. The checking process calculates the union of functional annotations for all metabolites in the current reduced biomarker set and assesses the proportion of pathway functional categories it covers relative to the total categories. An appropriate coverage ratio threshold is set (typically 80%-90%). When the coverage ratio reaches or exceeds this threshold, the reduced set is considered to have sufficient functional coverage breadth, and the selection process terminates. This coverage-based termination strategy avoids blindly pursuing a large number of biomarkers, instead focusing on the quality and breadth of functional coverage. It ensures that the reduced set can reflect changes in multiple biological processes, providing a comprehensive functional perspective for the differentiation of multiple diseases, while controlling the size of the biomarker set, enhancing the feasibility and cost-effectiveness of practical applications.
[0043] In this embodiment of the invention, the detailed implementation steps for calculating the coefficient of variation of the abundance of members within an ecological functional module composed of each metabolite and its associated microbial species in a simplified biomarker set, and recording it as the ecological stability baseline value, include: For each metabolite in the simplified biomarker set, microbial species with a covariant stability index greater than a preset association threshold are selected, forming ecological functional modules with the metabolite. Ecological functional module construction is a fundamental step in stability assessment, identifying the microbial association network of the metabolite. The construction process first involves screening all microbial species for each simplified biomarker metabolite, selecting those with a covariant stability index exceeding a preset association threshold (typically 2.0-3.0). Then, these highly associated microbial species are combined with the metabolite itself to form metabolite-centric ecological functional modules. These modules reflect the functional association network between the metabolite and the microbial community, representing potential ecological functional units, such as microbial participants and metabolites in specific metabolic pathways. Module construction, based on data-driven covariant relationships, captures stable associations observed experimentally, rather than relying solely on known relationships reported in the literature, enabling the discovery of new functional associations and providing a more comprehensive ecological network context for stability assessment.
[0044] In the entire sample of the healthy control group, the coefficient of variation (CV) of abundance for each member within the ecological functional module was calculated. CV calculation is a fundamental step in quantifying stability, assessing the natural range of fluctuations under healthy conditions. The calculation process used all samples from the healthy control group as the baseline population, calculating the CV of abundance for each member (including metabolites and related microorganisms) within the ecological functional module, i.e., the ratio of standard deviation to mean. As a dimensionless indicator, the CV eliminates dimensional differences between different types of data, making the stability of metabolites and microorganisms comparable. The selection of the healthy control group as a baseline reflects the normal range of fluctuations under physiological conditions, providing a reference for abnormal states. This variation analysis method based on a healthy population conforms to the basic principles of clinical biomarker research, providing reliable basic data for subsequent stability assessments.
[0045] The weighted average of the abundance variation coefficients of all members within an ecological functional module is taken, with the weights being the covariance stability indices corresponding to each member. This average is denoted as the ecological stability baseline value. The ecological stability baseline value is the core indicator for comprehensively assessing the module's stability, reflecting overall stability through weighted integration. The calculation formula is as follows: ; Wherein, ESB(m) is the baseline value for the ecological stability of metabolite m, and M m CSI is a set of members of an ecological functional module consisting of metabolite m. i,m CV is the covariance stability index between member i and metabolite m (1 for the metabolite itself). i(H) represents the coefficient of variation of the abundance of member i in the healthy control group. This weighted averaging method considers the differences in the strength of association between different members and core metabolites, with members having a more stable association having a greater influence on the module stability assessment. The baseline value of ecological stability directly reflects the degree of intrinsic fluctuation of the biomarker-related ecological network under healthy conditions, providing stability constraints for subsequent screening, ensuring that the selected biomarkers remain relatively stable under physiological conditions, and enhancing their reliability and specificity as disease indicators.
[0046] In this embodiment of the invention, the detailed implementation steps for performing cross-disease group discriminative joint optimization on a simplified biomarker set, using the ecological stability baseline value meeting a preset stability threshold as a constraint, and outputting the final multi-group disease biomarker panel, include: A stable biomarker set is obtained by removing metabolites with an ecostability baseline value greater than a preset stability threshold from a simplified biomarker set. Stability screening is a crucial step in ensuring biomarker reliability, eliminating candidates with excessive natural fluctuations. The screening process sets an appropriate stability threshold (typically 0.3-0.5), removing metabolites with ecostability baseline values exceeding this threshold from the simplified set, and retaining highly stable metabolites to form the stable biomarker set. A high stability baseline value indicates that the metabolite and its associated microbial network fluctuate significantly under healthy conditions, potentially influenced by factors such as diet and environment, making it unsuitable as a disease biomarker. This stability-based screening strategy reduces interference from physiological fluctuations, enhances the specificity and reproducibility of biomarkers, and provides a stable characteristic basis for subsequent multi-disease differentiation, representing an important quality control step to improve clinical application value.
[0047] A feature input matrix for a multi-class discriminant model is constructed using the abundance of metabolites in a stable biomarker set. Feature matrix construction is a fundamental step in training the discriminant model, preparing the data structure for multi-disease classification. The construction process uses metabolites from the stable biomarker set as feature variables, organizing the abundance values of these metabolites for all samples into a feature matrix, where rows represent samples and columns represent metabolites. Simultaneously, disease classification labels for the samples are prepared, forming a supervised learning dataset. The multi-class discriminant model can be trained and evaluated using algorithms such as random forests, support vector machines, or deep learning for multi-class disease classification problems. This biomarker-based feature matrix construction method significantly reduces the dimensionality of the original data, reduces the risk of overfitting, improves the model's generalization ability, provides a high-quality starting point for subsequent stepwise feature selection, and enhances the discriminative power and interpretability of the final biomarker panel.
[0048] A forward stepwise selection strategy is employed, using the area under the receiver operating characteristic (ROC) curve (AUC) of the macro-average values for multiple disease categories as the objective function. Metabolites are progressively added to the feature input matrix. Forward stepwise selection is a classic method for optimizing biomarker combinations, constructing an optimal feature set through iteration. The selection process begins with an empty set. Each iteration examines all unselected metabolites, evaluating their impact on classification performance upon addition to the current feature set. The metabolite with the greatest improvement is selected and added to the feature set, the model is updated, and the next iteration begins. The objective function is the macro-average AUC (Area Under the ROC Curve), a weighted average of the AUCs for each disease category, ensuring the model has good discriminative ability across all disease categories and avoiding over-reliance on large categories. While this greedy search strategy is not globally optimal, it is computationally efficient, effectively constructing near-optimal biomarker combinations, balancing discriminative performance and computational complexity, and providing a practical solution for clinical applications.
[0049] After each metabolite is added, the marginal contribution of that metabolite to the discrimination sensitivity of each disease group is calculated. Marginal contribution calculation is a crucial step in evaluating feature value, quantifying the incremental utility of the newly added biomarker. The calculation process first records the model's classification sensitivity in each disease group before adding the new metabolite, as a baseline; then, after adding the metabolite, the model is retrained, and the updated sensitivity is recorded; finally, the difference between the two sensitivities is calculated to obtain the marginal contribution value. The calculation formula is: ; Where MC(m) is the marginal contribution value of metabolite m, D is the set of all disease groups, and Sens i (F) represents the classification sensitivity of feature set F for disease group i, and FU{m} represents adding metabolite m to the current feature set F. Choosing the minimum value rather than the average as the marginal contribution ensures that the added metabolite has a positive effect on all disease groups, avoiding sacrificing the discriminative power of other diseases while improving the discriminative power of some diseases, thus meeting the balance requirement of multi-disease biomarker screening.
[0050] The addition process terminates when the marginal contribution value falls below a preset contribution threshold, and the selected metabolite combinations are output as the final multi-group disease biomarker panel. Setting the termination condition is a crucial step in controlling the number of biomarkers, avoiding overfitting and redundant additions. An appropriate contribution threshold (typically 0.02-0.05) is set during the termination process. When the marginal contribution value of a newly added metabolite falls below this threshold, it is considered that the marginal utility of continuing to add features is insufficient, the selection process terminates, and the currently selected metabolite combinations are output as the final biomarker panel. This marginal contribution-based termination strategy avoids the limitations of subjectively setting the number of biomarkers. Instead, it dynamically determines the optimal size based on the information content of the data itself, ensuring both discriminative performance and model complexity, enhancing the practicality and economy of the biomarker panel, and providing a streamlined and efficient disease diagnostic tool for clinical applications.
[0051] In this embodiment of the invention, the detailed implementation steps for calculating the marginal contribution value of the metabolite to the discrimination sensitivity of each disease group include: Before adding the metabolite, the classification sensitivity of the multi-class discriminant model for each disease group was recorded and denoted as the baseline sensitivity vector. Recording the baseline sensitivity is the first step in calculating the marginal contribution and establishes a reference point for performance comparison. The recording process first trains the multi-class discriminant model using the current feature set and evaluates the model's performance using cross-validation. Then, the classification sensitivity of the model for each disease group is calculated separately, i.e., the proportion of correctly identified diseases. Finally, the sensitivity values of all disease groups are organized into a vector form as the baseline performance record before adding the new feature. Choosing sensitivity over accuracy better meets the needs of clinical applications, especially for imbalanced disease distribution scenarios. The baseline sensitivity vector comprehensively reflects the disease discrimination ability of the current biomarker combination, providing a clear comparative benchmark for evaluating the contribution of the new biomarker and is a necessary prerequisite for calculating marginal utility.
[0052] After adding the metabolite, the multi-class discriminant model was retrained, and the classification sensitivity for each disease group was recorded, denoted as the updated sensitivity vector. Recording the updated sensitivity is the second step in calculating the marginal contribution, obtaining performance data after the feature addition. The recording process involves adding the metabolite under investigation to the current feature set, retraining the multi-class discriminant model using the expanded feature set, employing the same cross-validation settings as the baseline model to ensure comparability, and calculating the classification sensitivity of the updated model for each disease group to form the updated sensitivity vector. This step maintains the consistency of the evaluation method, ensuring that performance changes truly stem from feature addition rather than differences in the evaluation method. It provides reliable updated performance data for calculating the marginal contribution, reflecting the change in discriminative power brought about by the added metabolite, and serves as direct evidence for evaluating the value of the feature.
[0053] The element-wise difference between the updated sensitivity vector and the baseline sensitivity vector is calculated. This difference calculation is a crucial step in quantifying performance improvement, directly reflecting the degree of discriminative improvement for each disease group. The calculation process performs element-wise subtraction between the updated and baseline sensitivity vectors to obtain the change in sensitivity for each disease group. Positive values indicate an improvement in discriminative power for a specific disease after the addition of the metabolite, negative values indicate a decrease in discriminative power, and zero values indicate no change. This vector difference-based calculation method preserves the independent evaluation results for each disease group, rather than simply averaging performance. It reveals the differentiated impact of the new feature on different diseases, providing complete change data for subsequent minimum value extraction, and is a fundamental step in balancing the discriminative power of multiple diseases.
[0054] The minimum value among the element-wise differences is taken as the marginal contribution value to ensure that the added metabolite positively improves the discriminative power of all disease groups. Minimum value extraction is the final step in defining the marginal contribution, emphasizing comprehensive improvement across all diseases. The extraction process selects the minimum element from the sensitivity difference vector as the final marginal contribution value, representing the degree of improvement of the metabolite for the disease with the smallest improvement across all disease groups. This minimum-value-based evaluation strategy embodies the "barrel principle," focusing on improving the weakest link to ensure that the added biomarker comprehensively improves system performance, rather than significantly improving only some diseases at the expense of discriminative power in others. The minimum marginal contribution serves as a guideline for feature selection, effectively guiding biomarker screening, balancing the discriminative needs of different diseases, and providing a scientific evaluation standard for constructing a balanced multi-disease biomarker panel.
[0055] This invention achieves a systematic screening of multiple disease biomarkers from the gut microbiota metabolome through batch correction, covariance analysis, constrained network construction, biomarker screening, functional redundancy dimensionality reduction, stability assessment, and cross-disease optimized output. The network-enhanced screening method of this invention can accurately identify biologically significant and stable disease biomarkers and effectively integrate microbiome and metabolome data.
[0056] The above are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art can still modify the technical solutions described in the foregoing embodiments or make equivalent substitutions for some of the technical features. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
[0057] It should be noted that all formulas in this manual are calculated by removing dimensions and taking their numerical values. The formulas are derived from software simulations based on a large amount of collected data to obtain the most recent real-world results. The preset parameters and thresholds in the formulas are set by those skilled in the art according to the actual situation.
[0058] Although embodiments of the invention have been shown and described, those skilled in the art will understand that various changes, modifications, substitutions and alterations can be made to these embodiments without departing from the principles and spirit of the invention, the scope of which is defined by the claims and their equivalents.
Claims
1. A multi-protocol disease biomarker screening system for gut microbiota metabolomics, characterized in that, include: The batch correction module is used to acquire gut microbial species abundance data and untargeted metabolomics mass spectrometry raw data from multiple groups of disease subjects and healthy control groups, and to perform batch-to-batch drift correction based on the internal reference metabolite comparison value on the raw metabolomics mass spectrometry data of different detection batches, generating a corrected metabolite abundance matrix. The covariance analysis module is used to calculate the cross-sample abundance covariance trajectory between each metabolite and each microbial species for each metabolite in the corrected metabolite abundance matrix, and extract the directional consistency ratio of the covariance trajectory, which is denoted as the covariance stability index. The constraint network construction module is used to calculate the pathway topological distance between metabolite pairs based on the number of enzymatic reaction steps between each metabolite in the metabolic pathway database, and to construct a pathway constraint metabolite association network in combination with the covariant stability index. The biomarker screening module is used to calculate the pathway constraint betweenness centrality for each metabolite node in the pathway constraint metabolite association network, and fuse it with the standardized effect size of the metabolite among the disease groups to obtain the network enhancement discrimination score, and screen the candidate biomarker set based on the network enhancement discrimination score. The functional redundancy dimensionality reduction module is used to compare the functional annotations of each pair of metabolites in the candidate biomarker set to the corresponding metabolic pathways, calculate the functional coverage overlap, and perform stepwise elimination according to the maximum coverage and minimum redundancy criterion to generate a simplified biomarker set. The stability assessment module is used to calculate the coefficient of variation of the abundance of members within the ecological functional module composed of each metabolite and its associated microbial species in the simplified biomarker set in a healthy control group sample, and record it as the ecological stability baseline value. The cross-disease optimization output module is used to perform cross-disease group discriminative joint optimization on the simplified biomarker set with the ecological stability baseline value meeting the preset stability threshold as a constraint, and output the final multi-group disease biomarker panel.
2. The system according to claim 1, characterized in that, The process of performing inter-batch drift correction based on the internal reference metabolite contrast value on the raw metabolomics mass spectrometry data of different detection batches to generate a corrected metabolite abundance matrix includes: Metabolites with a detection rate greater than the preset detection threshold and a coefficient of variation less than the preset variation threshold in all test batches are screened and marked as candidate internal control metabolites. The candidate internal control metabolites were paired up in pairs, and the stability coefficient of the abundance ratio of each pair of candidate internal control metabolites was calculated between batches. The top K pairs of candidate internal control metabolites with the highest stability coefficients are selected and denoted as anchored internal control pairs. Using the median of the inter-batch abundance ratio of the anchored intrinsic reference pair as a benchmark, calculate the scaling factor of each batch relative to the benchmark batch; The abundance of all metabolites in each batch is linearly corrected using the scaling factor to generate the corrected metabolite abundance matrix.
3. The system according to claim 1, characterized in that, The calculation of its cross-sample abundance covariance trajectory with each microbial species, and the extraction of the directional consistency ratio of the covariance trajectory, denoted as the covariance stability index, includes: After randomly shuffling all samples within disease groups, a sample subset sequence is drawn a predetermined number of times. For each sample subset sequence extracted, the sign consistency of the changes in the abundance of the target metabolite and the target microbial species is calculated. The sign consistency is the proportion of sample pairs in which the abundance of both increases or decreases simultaneously to the total number of sample pairs. The mean and standard deviation of the sign consistency in the preset number of extractions are calculated. The covariant stability index is obtained by dividing the mean of the sign consistency by the standard deviation.
4. The system according to claim 1, characterized in that, The process of calculating the pathway topological distance between metabolite pairs based on the number of enzymatic reaction steps between metabolites in the metabolic pathway database, and constructing a pathway-constrained metabolite association network by combining the covariant stability index, includes: In the metabolic pathway database, a global directed graph of metabolic reactions is constructed with metabolites as nodes and enzyme-catalyzed reactions as edges; For each pair of metabolites in the corrected metabolite abundance matrix, the shortest path length between them is calculated in the global metabolic response directed graph and denoted as the pathway topology distance. When the pathway topological distance is less than a preset topological distance threshold, an edge is established between the metabolite pairs in the pathway-constrained metabolite association network. The weight of this side is set to the mean of the product of the covariant stability indices of the metabolite with respect to the respective set of microbial species.
5. The system according to claim 1, characterized in that, The process of calculating the pathway-constrained betweenness centrality for each metabolite node and fusing it with the standardized effect size of that metabolite across disease groups yields a network enhancement discrimination score, including: In the pathway-constrained metabolite association network, the proportion of the total number of shortest paths passing through each metabolite node to the total number of shortest paths in the network is calculated and denoted as the pathway-constrained betweenness centrality. For each metabolite, its standardized effect size was calculated between each disease group and the healthy control group. The maximum absolute value of the standardized effect size is taken as the maximum cross-group effect size of the metabolite. The product of the pathway constraint betweenness centrality and the maximum cross-group effect size is used as the network enhancement discrimination score.
6. The system according to claim 1, characterized in that, The calculation function covers overlap, and a step-by-step elimination process is performed according to the maximum coverage and minimum redundancy criterion to generate a simplified set of markers, including: For each metabolite in the candidate biomarker set, extract all pathway function categories it participates in in the metabolic pathway database, and construct a functional annotation set for that metabolite; For each pair of candidate biomarker metabolites, calculate the Jaccard similarity coefficient of their functional annotation sets, which is denoted as the functional coverage overlap. Select the simplified markers in descending order of the network enhancement discrimination scores; Each time a new metabolite is selected, the functional coverage overlap between the new metabolite and all previously selected metabolites is checked. If any functional coverage overlap exceeds a preset overlap threshold, the metabolite is skipped. The selection process terminates when the number of pathway function categories covered by the union of all functional annotation sets of the simplified marker set reaches a preset coverage ratio.
7. The system according to claim 1, characterized in that, The ecological functional module composed of each metabolite and its associated microbial species in the simplified biomarker set is used to calculate the coefficient of variation of the abundance of members within the module in a healthy control group sample, which is recorded as the ecological stability baseline value, including: For each metabolite in the simplified biomarker set, a microbial species with a covariance stability index greater than a preset association threshold is selected, and together with the metabolite, they constitute the ecological functional module of the metabolite. In all samples of the healthy control group, the abundance variation coefficient of each member within the ecological functional module was calculated. The weighted average of the abundance variation coefficients of all members within the ecological functional module is taken, with the weight being the covariance stability index corresponding to each member, and is denoted as the ecological stability baseline value.
8. The system according to claim 1, characterized in that, The process involves performing cross-disease group discriminative joint optimization on the simplified biomarker set, using the ecological stability baseline value meeting a preset stability threshold as a constraint, to output a final panel of multiple disease biomarkers, including: From the simplified biomarker set, metabolites whose ecological stability baseline value is greater than the preset stability threshold are removed to obtain a stable biomarker set; The feature input matrix of the multi-class discrimination model is constructed using the abundance of metabolites in the set of stable biomarkers; A forward stepwise selection strategy is adopted, with the area under the macro-average receiver operating characteristic curve of multiple disease classifications as the objective function, and metabolites are added to the feature input matrix stepwise. After each metabolite is added, the marginal contribution of that metabolite to the discrimination sensitivity of each disease group is calculated. When the marginal contribution value is lower than the preset contribution threshold, the addition is terminated, and the selected metabolite combinations are output as the final multi-group disease biomarker panel.
9. The system according to claim 2, characterized in that, The calculation of the stability coefficient of the abundance ratio of each pair of candidate internal control metabolites across batches includes: For each pair of candidate internal control metabolites, the ratio of their abundance is calculated for each sample in each detection batch to form the ratio vector for that batch. Calculate the inter-batch variance and within-batch variance of the ratio vector across all batches; The stability coefficient is calculated by dividing the intra-batch variance by the inter-batch variance. When the stability coefficient is greater than the preset stability screening threshold, the candidate internal reference metabolite pair is retained.
10. The system according to claim 8, characterized in that, The calculation of the marginal contribution of the metabolite to the discrimination sensitivity of each disease group includes: Before adding the metabolite, the classification sensitivity of the multi-class discrimination model for each disease group was recorded and denoted as the baseline sensitivity vector. After adding the metabolite, the multi-class discrimination model was retrained and the classification sensitivity of each disease group was recorded, denoted as the updated sensitivity vector. Calculate the element-by-element difference between the updated sensitivity vector and the reference sensitivity vector; The minimum value among the element-by-element differences is taken as the marginal contribution value.