Intestinal flora associated disease risk prediction system based on multi-data-set difference mutual identification

By constructing a multi-level microbial interaction network and dynamic evolution simulation, the network resilience and generating a threshold curve are quantified, which solves the problem of difficulty in distinguishing health from disease state in the existing technology, and achieves early and accurate early warning of intestinal flora-related diseases.

CN120015329AInactive Publication Date: 2025-05-16THE 3RD AFFILIATED HOSPITAL OF CHANGCHUN UNIVERSITY OF CHINESE MEDICINE
View PDF 0 Cites 6 Cited by

Patent Information

Application Number
CN202510490969.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-05-16
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing intestinal microbiota analysis methods ignore the complex interaction between microorganisms, making it difficult to distinguish between healthy networks and pre-disease states, limiting the accuracy of disease risk prediction, especially in multi-dataset environments.

Method used

Using the method of differential mutual verification of multiple data sets, by constructing a three-level microbial interaction network of species-genus-gate, dynamically evolve to simulate the time changes of network structure, quantify network resilience, identify fragile nodes and perturbation propagation paths, generate network resilience threshold curves, and achieve disease risk warning.

Benefits of technology

It has achieved early and accurate warnings for intestinal flora-related diseases, improved prediction accuracy and specificity, reduced false positive rates, and provided specific target suggestions for precise intervention.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120015329A_ABST
    Figure CN120015329A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of medical informatics, and discloses an intestinal flora associated disease risk prediction system based on multi-data set difference mutual identification, and the system comprises the steps: capturing the dynamic evolution characteristics of a network through constructing a species-genus-phylum three-level microbial interaction network, and employing a time sequence diagram convolution network; and designing a network disturbance experiment, introducing a network recovery force index to quantify the network recovery capability, constructing a network vulnerability map to identify key nodes and analyze a disturbance propagation path, and integrating stable and consistent network toughness characteristics in multiple data sets to establish an early warning mechanism. By means of the method, early and accurate early warning of disease risks is achieved, and a key time window is provided for clinical intervention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of medical informatics, and more specifically, to a risk prediction system for intestinal flora-associated diseases with mutual verification of differences in multiple data sets. Background Art

[0002] Existing methods for analyzing intestinal flora often ignore the complex interactions between microorganisms, or only focus on static network topology, lacking the assessment of network resilience. Especially in a multi-dataset environment, it is difficult to distinguish between a healthy network with strong adaptability and a fragile pre-disease state, limiting the accuracy of the early warning system. Summary of the invention

[0003] The present invention provides a risk prediction system for intestinal flora-associated diseases with mutual verification of differences in multiple data sets, which solves the technical problems of limitations of single data set analysis and insufficient network resilience assessment in related technologies.

[0004] The first aspect of the present invention provides a risk prediction system for intestinal flora-related diseases with mutual verification of differences in multiple data sets, comprising:

[0005] The network construction module is used to construct a species-genus-phylum three-level microbial interaction network to capture the microbial interaction relationships at different taxonomic levels;

[0006] Dynamic evolution module, which is used to model the dynamic evolution of multi-level networks through temporal graph convolutional networks and predict the temporal change pattern of network structure;

[0007] The disturbance simulation module is used to simulate network disturbances and track the recovery process, quantify the network's ability to recover from disturbances, and calculate the network resilience index;

[0008] Vulnerability analysis module, used to identify vulnerable nodes and disturbance propagation paths in the network and build a network vulnerability map;

[0009] A multi-dataset integration module is used to integrate the network resilience features of multiple independent datasets and identify stable and consistent features in multiple datasets;

[0010] The threshold curve generation module generates a network resilience threshold curve based on the disease outcome information in the historical data;

[0011] The risk warning module is used to monitor the network resilience index of real-time collected samples. When the index at multiple consecutive time points is lower than the corresponding threshold curve, a disease risk warning is triggered.

[0012] Furthermore, the construction of the species-genus-phylum three-level microbial interaction network includes: quality control filtering, splicing and clustering of the original sequencing data, obtaining the operational classification unit table, and converting it into a relative abundance matrix of each classification level; using the SparCC algorithm to calculate the association strength between microorganisms at each level and constructing a network adjacency matrix; establishing vertical connections between networks at different levels through taxonomic relationships to form a multi-level network with hierarchical interconnection.

[0013] Furthermore, the temporal graph convolutional network includes: a graph convolutional layer, a temporal gating unit, an attention mechanism layer and a fusion layer, wherein the update formula of the graph convolutional layer is:

[0014]

[0015] in, Indicates The node feature matrix of the layer, Indicates The node feature matrix of the layer, For time point The adjacency matrix of plus self-loops, is the degree matrix, is the weight matrix, is the activation function.

[0016] Furthermore, the quantified network's ability to recover from disturbances includes: constructing multiple disturbance modes to act on the network, including node deletion, edge weight reduction, and topological structure reorganization; using a stochastic differential equation model to simulate the recovery trajectory of the network after the disturbance; and calculating the network resilience index NRI based on the network recovery trajectory, and the calculation formula is:

[0017]

[0018] in, represents the network recovery trajectory to calculate the network resilience index, Representation Network The structural stability measure, Indicates After the perturbation, the network is the number of disturbances, Represents the relative stability loss caused by a single disturbance.

[0019] Furthermore, the form of the stochastic differential equation model is:

[0020]

[0021] in, Indicates a small change in network status. Indicates time The network state vector, To recover the trend function, is the state-dependent noise intensity, is the Wiener process, Indicates a small time interval.

[0022] Furthermore, the method for constructing the network vulnerability map is:

[0023]

[0024] in, Representation Network Vulnerability maps, Represents the original complete network The network resilience index, Representation Network The node set of Indicates the removal of a node After the network, Representation Node Contribution to network resilience, Indicates the removal of a node The resilience index of the post-network.

[0025] Furthermore, the analysis of the disturbance propagation path adopts the transfer entropy method, and its calculation formula is:

[0026]

[0027] in, and is the time series state of two microbial nodes, , Respectively represent nodes , In time The status value of Representation Node In time The status value of represents the joint probability distribution of three variables, Shown in the known and Under the conditions The conditional probability of Indicates that only known Under the conditions The conditional probability of Represents a slave node To Node information flow.

[0028] Furthermore, the identification of stable and consistent features in multiple data sets adopts a multi-view learning method, and the optimization goal is:

[0029]

[0030] in, is the consensus feature matrix, is the dataset weight, is the feature matrix, is the Frobenius norm, for norm, is the regularization parameter, is the total number of data sets, Represents finding the matrix that minimizes the objective function .

[0031] Furthermore, the generation formula of the network resilience threshold curve is:

[0032]

[0033] in, Indicates time point The network resilience threshold curve value, and Represents healthy individuals at time The average network resilience index and its standard deviation, is the sensitivity adjustment parameter.

[0034] The beneficial effects of the present invention are:

[0035] The present invention realizes early and accurate early warning of intestinal flora-related diseases through multi-level network disturbance recovery force quantification and multi-dataset mutual verification, and has the following technical effects:

[0036] Early warning capability: By capturing the dynamic changing trend of the network resilience index, the system can identify disease risks 4-6 weeks before the onset of clinical symptoms, providing a critical time window for early intervention.

[0037] Biomarker innovation: The network resilience index and its response pattern to perturbations, as a new system-level biomarker, have higher predictive value than traditional single bacterial community abundance indicators.

[0038] Advantages of mutual verification of multiple data sets: By integrating the network resilience features of multiple independent data sets, the system's prediction accuracy is improved by 28% and specificity is improved by 23%, while maintaining a sensitivity of more than 90%.

[0039] Reduced false positive rate: The system can effectively distinguish temporary network fluctuations from persistent network resilience decline, reducing the false positive rate by 42%, reducing unnecessary interventions and patient anxiety.

[0040] Precision intervention guidance: By identifying vulnerable nodes and disturbance propagation paths in the network, the system can provide specific target recommendations for precision intervention and improve intervention effectiveness. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 It is a module diagram of a multi-dataset difference-verified intestinal flora-associated disease risk prediction system of the present invention. DETAILED DESCRIPTION

[0042] The subject matter described herein will now be discussed with reference to example implementations. It should be understood that the discussion of these implementations is only to enable those skilled in the art to better understand and implement the subject matter described herein, and the functions and arrangements of the elements discussed may be changed without departing from the scope of protection of the present specification. Various examples may omit, replace, or add various processes or components as needed. In addition, the features described in some examples may also be combined in other examples.

[0043] Implementation method 1: a system for predicting the risk of intestinal flora-related diseases based on the mutual verification of differences among multiple data sets, such as Figure 1 As shown, including:

[0044] The network construction module is used to construct a species-genus-phylum three-level microbial interaction network to capture the microbial interaction relationships at different taxonomic levels;

[0045] The specific steps include:

[0046] Step 1.1: Data preprocessing and relative abundance calculation;

[0047] Perform quality control filtering, splicing and clustering on the raw sequencing data to obtain the operational taxonomic unit (OTU) table, which is converted into a relative abundance matrix of each taxonomic level (species, genus, phylum) ,in Representation sample Microbial groups relative abundance.

[0048] Step 1.2: Hierarchical association network construction;

[0049] The SparCC algorithm was used to calculate the association strength between microorganisms at each level and construct a network adjacency matrix. ,in Represents the classification level. For any two microbial nodes and , and its correlation strength is calculated as:

[0050] in, Indicated at the classification level Next, the microbial node and The strength of the association between and Represent the abundance vectors of microorganisms i and j, respectively. Represents microorganisms calculated using the SparCC algorithm and The correlation coefficient between is the association strength threshold (set to 0.3 by default), is the statistical significance level.

[0051] Step 1.3: Building connections between layers;

[0052] Through taxonomic relationships, vertical connections between networks at different levels are established to form a multi-level network with hierarchical interconnection. At the level ,microorganism At the level ,and Subordinate to , then establish the vertical connection weight .

[0053] Dynamic evolution module, which is used to model the dynamic evolution of multi-level networks through temporal graph convolutional networks and predict the temporal change pattern of network structure;

[0054] The specific steps include:

[0055] Step 2.1: Time series feature extraction;

[0056] For each time point Network structure , extract the topological feature vector , including node degree centrality, clustering coefficient and betweenness centrality.

[0057] Step 2.2: Temporal graph convolutional network training;

[0058] The network feature sequence is input into the T-GCN model and trained by iterative update formula:

[0059]

[0060] in, Indicates The node feature matrix of the layer, Indicates The node feature matrix of the layer, For time point The adjacency matrix of plus self-loops, is the degree matrix, is the weight matrix, is the activation function.

[0061] The T-GCN model includes the following components:

[0062] Graph convolution layer: used to capture the spatial interaction relationship of microbial nodes. The input is the node feature matrix and adjacency matrix, and the output is the node representation based on the graph structure.

[0063] Sequential gating unit: processes the network state sequence at different time points, captures time dependencies, and uses a bidirectional LSTM structure to enhance long-term memory capabilities;

[0064] Attention mechanism layer: assign weights to features at different time points to highlight the contributions of important time points;

[0065] Fusion layer: Integrate spatial and temporal features to form a comprehensive representation.

[0066] In the early warning application of inflammatory bowel disease, the intestinal flora data collected continuously at 10 time points are input into the T-GCN model, which can accurately predict the changes in network status in the next 4 weeks and identify disease risks in advance. The model uses the Adam optimizer, with a learning rate of 0.001, a batch size of 32, and a training cycle of 200.

[0067] Step 2.3: Network evolution prediction;

[0068] After training, the T-GCN model can predict the network status at future time points. , providing a benchmark for evaluating network disturbance resilience. The prediction formula is:

[0069]

[0070] in, Indicates the past The network status sequence at each time point, represents the trained T-GCN model, Indicates the past The network status sequence at each time point, represents the time span of the forecast, Indicates the size of the historical time window used for prediction.

[0071] The disturbance simulation module is used to simulate network disturbances and track the recovery process, quantify the network's ability to recover from disturbances, and calculate the network resilience index;

[0072] The specific steps include:

[0073] Step 3.1: Construction of perturbation experiment model;

[0074] Construct a variety of perturbation modes to act on the network, including: node deletion (simulating the disappearance of microorganisms), edge weight reduction (simulating the weakening of interactions) and topological structure reorganization (simulating the reconstruction of microbial communities). , generate the perturbed network :

[0075]

[0076] in, represents the original microbial interaction network, represents the network after perturbation, represents the disturbance operation, Represents the network structure change operator. For each perturbation mode, set multiple intensity levels ,implement Random perturbation (default ).

[0077] Step 3.2: Recovery process model construction;

[0078] The stochastic differential equation model is used to simulate the recovery trajectory of the network after the disturbance:

[0079]

[0080] in, Indicates a small change in network status. Indicates time The network state vector, To recover the trend function, is the state-dependent noise intensity, is the Wiener process, Represents a small time interval. The recovery trend function is constructed based on the early network dynamic characteristics:

[0081]

[0082] in, To recover the trend function, represents the recovery rate parameter, represents the network state before the disturbance, Indicates the difference between the current state and the original state. Indicates the drive parameters, represents the energy gradient term.

[0083] The stochastic differential equation model is implemented by the following steps:

[0084] Parameter initialization: Determined based on the training data set , and Functional form, for the inflammatory bowel disease data, Usually 0.15-0.25, Take 0.05-0.1;

[0085] Numerical solution: The improved Euler-Maruyama method is used to discretize the differential equations, and the time step is set to 0.01;

[0086] Multi-trajectory simulation: Generate 100 recovery trajectories for each disturbance condition and obtain the statistical distribution;

[0087] Recovery feature extraction: Extract key features from simulation trajectories, including recovery rate, stabilization time, and equilibrium state deviation.

[0088] In the application of irritable bowel syndrome, for disturbances with decreased diversity, the model successfully simulated three typical modes of recovery of the microbial network from disturbances: rapid complete recovery, slow partial recovery, and continuous deterioration, providing key indicators for disease risk assessment.

[0089] Step 3.3: Network resilience index calculation;

[0090] Define the Network Resilience Index (NRI) based on the network recovery trajectory:

[0091]

[0092] in, represents the network recovery trajectory to calculate the network resilience index, Indicates After the perturbation, the network is the number of disturbances, represents the relative stability loss caused by a single disturbance, Representation Network The structural stability measure of is calculated by integrating parameters such as network connectivity, modularity and heterogeneity:

[0093]

[0094] in, For network connectivity, For modularity, is the network heterogeneity, For modularity, represents network heterogeneity, , , They respectively represent the importance of controlling network heterogeneity, modularity, and network heterogeneity in the total score. The lower the NRI value, the stronger the network resilience.

[0095] Vulnerability analysis module, used to identify vulnerable nodes and disturbance propagation paths in the network and build a network vulnerability map;

[0096] The specific steps include:

[0097] Step 4.1: Vulnerability map generation;

[0098] Generate a network vulnerability map by analyzing the impact of nodes on network resilience :

[0099]

[0100] in, Representation Network Vulnerability maps, Representation Node The contribution value of Represents the original complete network The network resilience index, Indicates that the node The resilience index of the post-network, Indicates the removal of a node After the network, Representation Network The node set of For each node in the network Calculate their contribution values.

[0101] Step 4.2: Identification of key nodes;

[0102] Based on the vulnerability map, identify the key node set :

[0103]

[0104] in, represents a set of key nodes, For a single node in the network, Representation Node The contribution value of represents the contribution threshold, Representation Network The node set of .

[0105] Step 4.3: Disturbance propagation path analysis;

[0106] The transfer entropy (TE) method is used to analyze the propagation path of disturbances in the network:

[0107]

[0108] in, Represents a slave node To Node The transfer entropy value of Indicates a point in time, , Respectively represent nodes and nodes At the point in time The status value of Representation node At the point in time The status value of Representation Node In time and The status and nodes In time The joint probability distribution of the states is Indicates that at a known node In time Status and nodes In time Under the condition of the state, the node In time The probability of the state Indicates that at a known node In time Under the condition of the state, the node In time The probability of the state Denotes a logarithmic function. By calculating the transfer entropy between all node pairs, a disturbance propagation network is constructed to identify the main propagation paths and key relay nodes.

[0109] A multi-dataset integration module is used to integrate the network resilience features of multiple independent datasets and identify stable and consistent features in multiple datasets;

[0110] The specific steps include:

[0111] Step 5.1: Feature extraction from multiple datasets;

[0112] From multiple independent datasets Extracting network resilience feature matrix ,in, , , Instrument room No. 1, 2, Datasets, is the total number of data sets, and each feature matrix contains: resilience index time series, key node stability indicators, disturbance propagation pattern characteristics, etc.

[0113] Step 5.2: consensus feature identification;

[0114] Through a multi-view learning approach, we identify stable and consistent network resilience features across multiple datasets:

[0115]

[0116] in, is the consensus feature matrix, is the dataset weight, is the feature matrix, is the Frobenius norm, for norm, is the regularization parameter, is the total number of data sets, Represents finding the matrix that minimizes the objective function .

[0117] The implementation details of the multi-view learning method include:

[0118] Dataset weighting: Assign weights based on dataset size, quality, and relevance ,Correlation was assessed by Jensen-Shannon divergence between datasets;

[0119] Iterative optimization: The alternating direction method of multipliers (ADMM) was used to solve the optimization problem, with the maximum number of iterations set to 500 and the convergence threshold set to 1e-6;

[0120] Feature importance evaluation: based on The sparse solution of the norm is used to calculate the feature importance score;

[0121] Stable feature selection: Perform 100 resamplings using the bootstrap method and select features that are selected in more than 95% of the resamplings.

[0122] The threshold curve generation module generates a network resilience threshold curve based on the disease outcome information in the historical data;

[0123] Generate network resilience threshold curve based on disease outcome information in historical data :

[0124]

[0125] in, Indicates time point The network resilience threshold curve value, and Represents healthy individuals at time The average network resilience index and its standard deviation, is the sensitivity adjustment parameter.

[0126] The risk warning module is used to monitor the network resilience index of real-time collected samples. When the index at multiple consecutive time points is lower than the corresponding threshold curve, a disease risk warning is triggered;

[0127] Network resilience index by monitoring samples collected in real time , when continuous time points (default )of Both are lower than the corresponding threshold curve When the disease risk warning is triggered:

[0128]

[0129] in, Indicates time point The network resilience index, Indicates the number of time points for continuous monitoring, the default setting is 3. , Respectively represent the time points arrive Continuous The network resilience index sequence and threshold curve value sequence at each time point, Indicates time point The network resilience threshold curve value, Indicates that when the conditions are met, the value of the warning signal is True and triggers the warning, otherwise it is False, indicating a normal state.

[0130] At the same time, the system provides potential intervention target recommendations based on the results of the disturbance propagation path analysis, including key microbial nodes and their regulation strategies.

[0131] The technical effects of this implementation are as follows:

[0132] This implementation method achieves early and accurate warning of intestinal flora-related diseases through quantification of multi-level network disturbance recovery and mutual verification of multiple data sets, and has the following technical effects:

[0133] Early warning capability: By capturing the dynamic changing trend of the network resilience index, the system can identify disease risks 4-6 weeks before the onset of clinical symptoms, providing a critical time window for early intervention.

[0134] Biomarker innovation: The network resilience index and its response pattern to perturbations, as a new system-level biomarker, have higher predictive value than traditional single bacterial community abundance indicators.

[0135] Advantages of mutual verification of multiple data sets: By integrating the network resilience features of multiple independent data sets, the system's prediction accuracy is improved by 28% and specificity is improved by 23%, while maintaining a sensitivity of more than 90%.

[0136] Reduced false positive rate: The system can effectively distinguish temporary network fluctuations from persistent network resilience decline, reducing the false positive rate by 42%, reducing unnecessary interventions and patient anxiety.

[0137] Precision intervention guidance: By identifying vulnerable nodes and disturbance propagation paths in the network, the system can provide specific target recommendations for precision intervention and improve intervention effectiveness.

[0138] An application example of implementation mode 1 is as follows:

[0139] This system is used in early risk warning scenarios for inflammatory bowel disease (IBD). Taking the IBD risk population monitoring project of the Department of Gastroenterology of a tertiary hospital as an example, the project tracked and observed 200 IBD high-risk populations (including first-degree relatives of IBD patients, those with a history of autoimmune diseases, etc.), and used this system to conduct disease risk warnings by regularly collecting intestinal flora samples.

[0140] Implementation process example:

[0141] Multi-level network construction example:

[0142] 40 subjects were randomly selected from 200 subjects for sample analysis, and intestinal flora samples were collected from each person at 10 time points at 2-week intervals. The raw data were obtained by 16SrRNA sequencing. After quality control and OTU clustering, an average of 324 species, 121 genera, and 42 phylum-level taxa were identified in each sample.

[0143] The SparCC algorithm was used to calculate the microbial association network at each level, with the association strength threshold set to 0.3 and the p-value threshold set to 0.05. Table 1 shows some of the association networks constructed at the species level:

[0144] Table 1: Species-level association network (partial)

[0145]

[0146] The vertical connections between levels are constructed through taxonomic relationships to form a multi-level network structure. In the three-layer network, there are 2,387 edges at the species level, 843 edges at the genus level, 105 edges at the phylum level, and 463 vertical connections between levels.

[0147] Example of capturing dynamic network evolution features:

[0148] The network structure of 40 subjects at 10 time points was dynamically modeled, and the topological characteristics of the network at each time point were extracted, including degree distribution, clustering coefficient, betweenness centrality, etc. Table 2 shows the changes in network characteristics of a subject at 5 consecutive time points:

[0149] Table 2: Changes in network characteristics at 5 consecutive time points

[0150]

[0151] The above feature sequence was input into the T-GCN model, and the model converged after 200 rounds of training. The average prediction accuracy of 10-fold cross validation was 87.3%, and the F1 score was 0.85.

[0152] Network disturbance resilience quantification example:

[0153] Three types of perturbation experiments are performed on the constructed network:

[0154] (1) Randomly delete 10% of the nodes;

[0155] (2) Randomly reduce the weight of 30% of the edges;

[0156] (3) Reorganize 15% of the network topology. Table 3 shows the comparison of the network resilience index (NRI) between the healthy group and the IBD risk group after facing different disturbances:

[0157] Table 3: Comparison of NRI values ​​between the healthy group and the IBD risk group under different disturbances

[0158]

[0159] The stochastic differential equation model was used to simulate the recovery trajectory of the network after the perturbation, and 100 simulated recovery trajectories were generated for each subject. Through the analysis of the recovery trajectories, three typical recovery modes were identified:

[0160] Type A: Rapid and complete recovery, accounting for about 78% of the healthy group and 31% of the IBD risk group;

[0161] Type B: Slow partial recovery, accounting for about 19% of the healthy group and 42% of the IBD risk group;

[0162] Type C: continuous deterioration, accounting for about 3% of the healthy group and 27% of the IBD risk group;

[0163] Examples of network vulnerability identification and propagation path analysis:

[0164] By calculating the vulnerability contribution of each node, a network vulnerability map was constructed. Table 4 lists the five microbial nodes with the highest contribution in the IBD risk group:

[0165] Table 4: Microbial nodes with the highest contribution to vulnerability in the IBD risk group

[0166]

[0167] Examples of mutual verification of multiple data sets and risk warning:

[0168] Three independent IBD cohort datasets (US, Chinese, and European) were integrated, and a multi-perspective learning method was applied to identify stable and consistent network resilience features. Through the multi-perspective learning method, 10 stable and consistent network resilience features were identified from the three datasets, and resilience threshold curves were constructed based on these features. 40 sample subjects were followed up for 12 months, and 12 subjects were continuously triggered by the system during the follow-up period. Among them, 9 were eventually clinically diagnosed with IBD, and the other 3 showed a significant increase in inflammatory markers but had not yet reached the clinical diagnostic criteria.

[0169] Technical effect verification:

[0170] Early warning capability verification:

[0171] Table 5 shows the comparison of the early warning capability of this system with traditional methods:

[0172] Table 5: Comparison of early warning capabilities of different methods

[0173]

[0174] As can be seen from Table 5, this system can provide disease risk warning 5.8 weeks before the onset of clinical symptoms on average, which is 3.7 weeks and 4.3 weeks earlier than single bacterial abundance indicators and intestinal inflammation markers, respectively. It also has significant advantages in sensitivity, specificity and positive predictive value.

[0175] Verification of the advantages of multiple data sets:

[0176] Table 6 compares the prediction performance of the single dataset and multi-dataset mutual verification methods:

[0177] Table 6: Performance comparison of single dataset and multi-dataset mutual verification methods

[0178]

[0179] Through mutual verification of multiple data sets, the prediction accuracy of this system increased from an average of 70.9% to 88.4%, the specificity increased from an average of 64.3% to 85.2%, and the false positive rate decreased from an average of 35.7% to 14.8%, fully verifying the superior performance of this system in a multi-dataset environment.

[0180] The above describes an embodiment of the present invention, but this embodiment is not limited to the above-mentioned specific implementation mode. The above-mentioned specific implementation mode is merely illustrative and not restrictive. Under the guidance of this embodiment, ordinary technicians in this field can also make more forms of equivalent embodiments, all of which are protected by this embodiment.

Claims

1. A risk prediction system for intestinal flora-related diseases based on the mutual verification of differences in multiple data sets, characterized in that: include: The network construction module is used to construct a species-genus-phylum three-level microbial interaction network to capture the microbial interaction relationships at different taxonomic levels; Dynamic evolution module, which is used to model the dynamic evolution of multi-level networks through temporal graph convolutional networks and predict the temporal change pattern of network structure; The disturbance simulation module is used to simulate network disturbances and track the recovery process, quantify the network's ability to recover from disturbances, and calculate the network resilience index; Vulnerability analysis module, used to identify vulnerable nodes and disturbance propagation paths in the network and build a network vulnerability map; A multi-dataset integration module is used to integrate the network resilience features of multiple independent datasets and identify stable and consistent features in multiple datasets; The threshold curve generation module generates a network resilience threshold curve based on the disease outcome information in the historical data; The risk warning module is used to monitor the network resilience index of real-time collected samples. When the index at multiple consecutive time points is lower than the corresponding threshold curve, a disease risk warning is triggered.

2. According to claim 1, a multi-dataset difference-based intestinal flora-associated disease risk prediction system is characterized by: The construction of the species-genus-phylum three-level microbial interaction network includes: quality control filtering, splicing and clustering of the original sequencing data, obtaining the operational classification unit table, and converting it into a relative abundance matrix of each classification level; using the SparCC algorithm to calculate the association strength between microorganisms at each level and constructing a network adjacency matrix; establishing vertical connections between networks at different levels through taxonomic relationships to form a multi-level network with hierarchical interconnection.

3. The intestinal flora-associated disease risk prediction system based on multi-dataset difference verification according to claim 2 is characterized in that: The temporal graph convolution network includes: a graph convolution layer, a temporal gating unit, an attention mechanism layer and a fusion layer, wherein the update formula of the graph convolution layer is: ; in, Indicates The node feature matrix of the layer, Indicates The node feature matrix of the layer, For time point The adjacency matrix of plus self-loops, is the degree matrix, is the weight matrix, is the activation function.

4. The intestinal flora-associated disease risk prediction system based on multi-dataset difference verification according to claim 3 is characterized in that: The method of quantifying the network's ability to recover from disturbances includes: constructing multiple disturbance modes to act on the network, including node deletion, edge weight reduction, and topological structure reorganization; using a stochastic differential equation model to simulate the recovery trajectory of the network after the disturbance; and calculating the network resilience index NRI based on the network recovery trajectory, and the calculation formula is: ; in, represents the network recovery trajectory to calculate the network resilience index, Representation Network The structural stability measure, Indicates After the perturbation, the network is the number of disturbances, Represents the relative stability loss caused by a single disturbance.

5. The intestinal flora-associated disease risk prediction system based on multi-dataset difference mutual verification according to claim 4 is characterized in that: The stochastic differential equation model is in the form of: ; in, Indicates a small change in network status. Indicates time The network state vector, To recover the trend function, is the state-dependent noise intensity, is the Wiener process, Indicates a small time interval.

6. The intestinal flora-associated disease risk prediction system based on multi-dataset difference mutual verification according to claim 5, characterized in that: The method for constructing the network vulnerability map is: ; in, Representation Network Vulnerability maps, Represents the original complete network The network resilience index, Representation Network The node set of Indicates the removal of a node After the network, Representation Node Contribution to network resilience, Indicates the removal of a node The resilience index of the post-network.

7. The intestinal flora-associated disease risk prediction system based on multi-dataset difference mutual verification according to claim 6, characterized in that: The analysis of the disturbance propagation path adopts the transfer entropy method, and its calculation formula is: ; in, and is the time series state of the two microbial nodes, , Respectively represent nodes , In time The status value of Representation Node In time The status value of represents the joint probability distribution of three variables, Shown in the known and Under the conditions The conditional probability of Indicates that only Under the conditions The conditional probability of Represents a slave node To Node information flow.

8. The intestinal flora-associated disease risk prediction system based on multi-dataset difference mutual verification according to claim 7 is characterized in that: The identification of stable and consistent features in multiple data sets adopts a multi-view learning method, and the optimization goal is: ; in, is the consensus feature matrix, is the dataset weight, is the feature matrix, is the Frobenius norm, for norm, is the regularization parameter, is the total number of data sets, Represents finding the matrix that minimizes the objective function .

9. The intestinal flora-associated disease risk prediction system based on multi-dataset difference mutual verification according to claim 8, characterized in that: The generation formula of the network resilience threshold curve is: ; in, Indicates time point The network resilience threshold curve value, and Represents healthy individuals at time The average network resilience index and its standard deviation, is the sensitivity adjustment parameter.

Citation Information

Cited By

  • Dynamic monitoring and analyzing system for multi-trophic-level interaction relation of ecological system

    CN121210891A

  • Colorectal cancer risk assessment system based on intestinal microbial diversity

    CN121545757A

  • Digestive system disease-based risk prediction method and device

    CN122050796A

  • Functional gastrointestinal disease risk prediction system based on multi-modal data

    CN122337677A

  • Functional gastrointestinal disorder risk prediction system based on multi-modal data

    CN122337677B