Prediction model construction system, storage medium and kit for Clostridium difficile infection and recurrent infection
By constructing a machine learning-based intestinal flora prediction model and using 16S rRNA sequencing and feature screening algorithms of fecal samples, the problem of predicting CDI and rCDI in Crohn's disease patients was solved, and high-sensitivity and specificity risk prediction was achieved to support individualized treatment strategies.
Patent Information
- Application Number
- CN202510947505.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-10
AI Technical Summary
Existing technologies are unable to effectively predict the occurrence of Clostridium difficile infection (CDI) and its recurrent infection (rCDI) in Crohn's disease patients, ignoring the complex interactive relationship between the overall imbalance of intestinal flora and infection, resulting in passive treatment paths and high recurrence rates.
A machine learning-based prediction model was constructed. By collecting stool samples and performing 16S rRNA sequencing, the LEfSe and Lasso feature screening algorithms were used in combination with the ElasticNet method to train CDI infection risk prediction models and rCDI recurrence risk prediction models, identify characteristic bacterial genus sets, and make predictions by calculating risk scores.
The model achieved accurate identification of CDI and rCDI in Crohn's disease patients. It performed well in the training set and independent test set, provided a forward-looking early warning of infection risk, assisted in individualized intervention treatment, and reduced the recurrence rate.
Smart Images

Figure CN120432173B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of medical prediction, and in particular to a prediction model construction system, storage medium, and molecular detection kit based on intestinal flora characteristics for Clostridioides difficile infection and recurrent CDI (rCDI), which are suitable for infection risk assessment, individualized intervention guidance, and precise clinical management. Background Art
[0002] Clostridium difficile infection (CDI) is an intestinal infectious disease with a steadily increasing incidence worldwide, particularly in patients receiving antibiotics or experiencing intestinal microbial imbalance. Patients with inflammatory bowel diseases, particularly Crohn's disease, are at the highest risk for CDI and its recurrence (rCDI) due to underlying chronic intestinal inflammation and frequent anti-infective treatment. The incidence of rCDI can be as high as 30%, significantly increasing hospitalization time, medical burden, and mortality risk.
[0003] Existing CDI diagnostics primarily focus on detecting Clostridium difficile toxins (such as tcdA / B), but overlook the complex interplay between overall gut microbiome imbalance and CDI pathogenesis. Furthermore, current technologies are unable to effectively predict the development of rCDI, resulting in a passive treatment path for patients and high recurrence rates. These existing diagnostic and predictive technologies generally overlook the role of host microbiome characteristics in the development of CDI and rCDI, and have failed to develop effective personalized risk prediction tools based on changes in individual gut microbiome profiles. This is particularly true in high-risk populations such as Crohn's disease, where traditional methods struggle to accurately reflect potential infection risk. Summary of the Invention
[0004] The purpose of the present invention is to provide a prediction model construction system, storage medium and kit for the occurrence of Clostridium difficile infection and recurrent infection and their applications.
[0005] To this end, the present invention provides a system for constructing a prediction model for the occurrence and recurrent infection of Clostridium difficile infection, comprising:
[0006] A collection module is used to collect stool samples from subjects with Crohn's disease and sequence the stool samples to obtain genus-level relative abundance information of the bacterial flora;
[0007] A feature selection module is used to apply LEfSe and Lasso feature screening algorithms to the stool sample and screen out a characteristic genus set based on the genus-level relative abundance information of the bacterial community;
[0008] A training module is used to train a CDI infection risk prediction model and an rCDI recurrence risk prediction model using the ElasticNet method based on the genus-level relative abundance information of the characteristic bacterial genus set;
[0009] The identification module is used to obtain the genus-level relative abundance information of the characteristic bacterial genus set of the individual to be tested, and input it into the trained CDI infection risk prediction model or rCDI recurrence risk prediction model to obtain the corresponding risk score.
[0010] Furthermore, in the above system, the feature selection module is used to perform LEfSe analysis on fecal samples, calculate the linear discriminant analysis LDA score of each bacterial genus between groups, and obtain a preliminary screening set of bacterial genera based on the LDA score; based on the preliminary screening set of bacterial genera, the Lasso feature selection algorithm is used to obtain a characteristic bacterial genus set.
[0011] Furthermore, in the above system, the stool sample includes: stool samples of subjects with Clostridium difficile infection and those without Crohn's disease during diagnosis or follow-up, as first samples; the acquisition module is used to obtain genus-level relative abundance information of the bacterial flora of the first sample; the first sample is split into a first training data set and a first test set;
[0012] The feature selection module is used to perform LEfSe analysis on the fecal samples in the first training data set, calculate the linear discriminant analysis LDA score of each bacterial genus between groups, and select bacterial genera with LDA>2 or LDA<–2 and p<0.05 to enter the first preliminary screening set; based on the first genus-level relative abundance matrix of the first preliminary screening set, the Lasso feature selection algorithm is used to obtain the corresponding first feature bacterial genus set with non-zero coefficients.
[0013] Furthermore, in the above system, the first characteristic bacterial genus set at least includes:
[0014] Veillonella; Bifidobacterium; Anaerostipes; Holdemanella; Romboutsia; Erysipelotrichaceae UCG-003; Clostridium innocuum group; Pyramidobacter; Parabacteroides; Weissella; Peptostreptococcus; Prevotella; Collinsella; Coprococcus; Bacteroides; Eubacterium eligens group; Eubacterium hallii group); Akkermansia and Ruminococcaceae UCG-013.
[0015] Furthermore, in the above system, the training module is used to perform a CDI infection risk prediction model based on the first training data set and output an infection risk score P1 representing the infection risk level;
[0016] ;
[0017] Among them, P 1,i represents the CDI infection risk score of the i-th subject;
[0018] x ij is the relative abundance information of the jth characteristic genus in the first characteristic genus for the i-th subject in the first sample; wherein, the relative abundance information of the jth characteristic genus in the first characteristic genus for the i-th subject in the first sample is the relative abundance information after log10 transformation and Z-score normalization;
[0019] β j is the weight coefficient of the jth characteristic bacterial genus in the first characteristic bacterial genus learned in the trained CDI infection risk prediction model;
[0020] β0 is the bias term;
[0021] The infection risk score P1 is a continuous variable between 0 and 1;
[0022] The training module is also used to obtain the recognition performance of the infection risk score of the trained CDI infection risk prediction model on the first training data set and the first test set respectively; if the recognition performance of the trained CDI infection risk prediction model on the first training data set and the first test set is less than a preset difference threshold, the trained CDI infection risk prediction model will be used as the final CDI infection risk prediction model.
[0023] Furthermore, in the above system, the stool sample includes: stool samples from subjects with recurrent Clostridium difficile infection and Crohn's disease who have not relapsed after a previous Clostridium difficile infection, as a second sample; the collection module is further used to obtain genus-level relative abundance information of the bacterial flora of the second sample; split the second sample into a second training data set and a second test set;
[0024] The feature selection module is used to perform LEfSe analysis on the fecal samples in the second training data set, calculate the linear discriminant analysis LDA score of each bacterial genus between groups, and select bacterial genera with LDA>2 or LDA<–2 and p<0.05 to enter the second preliminary screening set; based on the second genus-level relative abundance matrix of the second preliminary screening set, a second characteristic bacterial genus set with corresponding non-zero coefficients is obtained.
[0025] Furthermore, in the above system, the second characteristic bacterial genus set at least includes:
[0026] Veillonella; Megamonas; Marvinbryantia; Dorea; Lachnoclostridium; Butyricicoccus; Romboutsia; Anaerotruncus; Clostridium sensu stricto 1; and Staphylococcus.
[0027] Furthermore, in the above system, the training module is used to perform an rCDI recurrence risk prediction model based on the second training data set, and output a recurrence risk score P2, which represents the probability of recurrence. The scoring formula is as follows:
[0028] ;
[0029] P 2,i represents the recurrence risk score of the i-th patient with previous CDI;
[0030] x ikis the relative abundance information of the kth characteristic genus in the second characteristic genus for the i-th subject of the second sample; wherein, the relative abundance information of the kth characteristic genus in the second characteristic genus for the i-th subject of the second sample is the relative abundance information after log10 transformation and Z-score normalization;
[0031] γ k is the weight coefficient of the kth characteristic bacterial genus in the second characteristic bacterial genus learned in the trained rCDI recurrence risk prediction model;
[0032] σ() is the Sigmoid function, which maps the output to between 0 and 1;
[0033] recurrence risk score P2, a continuous variable between 0 and 1;
[0034] The training module is also used to obtain the recognition performance of the infection risk score of the trained rCDI recurrence risk prediction model on the second training data set and the second test set respectively; if the recognition performance of the trained rCDI infection risk prediction model on the second training data set and the second test set is less than a preset difference threshold, the trained rCDI recurrence risk prediction model is used as the final rCDI recurrence risk prediction model.
[0035] Furthermore, in the above system, the ElasticNet method integrates dual constraints of L1 regularization and L2 regularization, and the hyperparameters are set as follows: the α value is set to 0.5, that is, the L1:L2 ratio is 1:1; the λ2 value used to control the intensity of the penalty for the regression coefficient is based on 5-fold cross-validation to determine the optimal λ2.min value.
[0036] Furthermore, in the above system, the threshold of the trained CDI infection risk prediction model is τ1=0.718. If P1≥τ1, the patient is judged to be at high risk of CDI;
[0037] The threshold value of the trained rCDI recurrence risk prediction model is τ2 = 0.771. If P2 ≥ τ2, it indicates that the patient has a high possibility of recurrence.
[0038] According to another aspect of the present invention, a computer-readable storage medium is provided, on which computer-executable instructions are stored. When the instructions are executed by a processor, the processor executes the various modules of the system as described in any one of the above items.
[0039] According to another aspect of the present invention, there is also provided a microecological detection kit, comprising a stool collection component, including a stool sample collection tube of an individual to be tested and a DNA stabilizing solution;
[0040] DNA extraction kit for obtaining intestinal microbial DNA based on stool sample collection tubes and DNA stabilization solution;
[0041] The amplification and sequencing component is used to amplify and sequence the intestinal microbial DNA to obtain the genus-level relative abundance information of the first characteristic bacterial genus set or the second characteristic bacterial genus set; wherein,
[0042] The first characteristic bacterial genus set at least includes:
[0043] Veillonella; Bifidobacterium; Anaerostipes; Holdemanella; Romboutsia; Erysipelotrichaceae UCG-003; Clostridium innocuum group; Pyramidobacter; Parabacteroides; Weissella; Peptostreptococcus; Prevotella; Collinsella; Coprococcus; Bacteroides; Eubacterium eligens group; Eubacterium hallii group); Akkermansia and Ruminococcaceae UCG-013;
[0044] The second characteristic bacterial genus set includes at least:
[0045] Veillonella; Megamonas; Marvinbryantia; Dorea; Lachnoclostridium; Butyricicoccus; Romboutsia; Anaerotruncus; Clostridium sensu stricto 1; and Staphylococcus.
[0046] Specifically, the collection module is used to collect fecal samples from subjects diagnosed with Crohn's disease, including: samples from subjects who developed Clostridium difficile infection (CDI) and those who did not develop CDI during diagnosis or treatment follow-up, or samples from subjects who developed recurrent infection (rCDI) after previous CDI and those who did not develop recurrent infection (NR); extract intestinal microbiome 16S rRNA sequencing data from each fecal sample to obtain genus-level relative abundance values at the genus level of the bacterial community.
[0047] The feature selection module is used to select a first characteristic bacterial genus set based on the genus-level relative abundance information of the first sample using Lasso cross-validation and linear discriminant analysis (LDA); and to select a second characteristic bacterial genus set based on the genus-level relative abundance information of the second sample using Lasso cross-validation and linear discriminant analysis (LDA). Here, machine learning training is performed based on the relative abundance information of intestinal bacterial genus between the CDI and non-CDI groups, and between the rCDI and NR groups, to select characteristic bacterial genera for distinguishing infection status, which serve as the bacterial biomarker set after feature selection.
[0048] The training module is used to train the first ElasticNet model based on the genus-level relative abundance information of the first characteristic bacterial genus set to obtain a trained CDI infection risk prediction model, which is used to output the infection risk score P1; and to train the second ElasticNet model based on the genus-level relative abundance information of the second characteristic bacterial genus set to obtain a trained rCDI recurrence risk prediction model, which is used to output the recurrence risk score P 2。
[0049] Preferably, the training module is configured to train a first ElasticNet model based on the genus-level relative abundance information of a first bacterial genus set in a first sample to obtain a trained CDI infection risk prediction model, which outputs an infection risk score P1; and to train a second ElasticNet model based on the genus-level relative abundance information of a second bacterial genus set in a second sample to obtain a trained rCDI recurrence risk prediction model, which outputs a recurrence risk score P2. The CDI infection risk prediction model and its rCDI recurrence risk prediction model can be trained based on feature-selected bacterial genus abundance data. These models utilize the ElasticNet algorithm and are implemented using R (4.3.1) software and its glmnet package.
[0050] The identification module is used to obtain genus-level relative abundance information for the first characteristic bacterial genus set of the individual being tested and input it into the trained CDI infection risk prediction model to output an infection risk score P1. It also obtains genus-level relative abundance information for the second characteristic bacterial genus set of the individual being tested and inputs it into the trained rCDI recurrence risk prediction model to output a recurrence risk score P2. Here, the relative abundance values of the fecal genus flora of the individual with Crohn's disease being tested can be obtained and input into the aforementioned trained infection identification model to calculate and output an infection risk score to assist in determining whether the subject is at risk for CDI or rCDI.
[0051] Furthermore, in the above system, after the first sample is sequenced by 16S rRNA, the acquisition module obtains a first genus-level relative abundance matrix; after the second sample is sequenced by 16S rRNA, a second genus-level relative abundance matrix is obtained; all first samples are split into a non-overlapping first training data set and a first test set; all second samples are split into a non-overlapping second training data set and a second test set.
[0052] Here, a standardized 16S rRNA high-throughput sequencing platform can be used to extract genus-level relative abundance and construct the first set and the test set.
[0053] Furthermore, the feature selection module is used to perform LDAEffect Size analysis (LEfSe) on the stool samples in the first training data set, calculate the linear discriminant analysis (LDA) score of each bacterial genus between groups, and select bacterial genera with LDA>2 or LDA<–2 and p<0.05 to enter the first preliminary screening set;
[0054] The feature selection module is configured to divide the first genus-level relative abundance matrix of the first preliminary screening set into five non-overlapping data sets, each time using one of the five non-overlapping data sets that has not been selected as a first internal validation set, and each time using the remaining four data sets as a first training set; wherein both the first internal validation set and the first training set include non-overlapping Crohn's disease subjects with and without CDI;
[0055] The feature selection module is used to train the feature selection model for 5 rounds based on the first training set each time, using the Lasso feature selection algorithm to obtain the corresponding 5-round feature selection model. In each round of feature selection model, for each λn value in the hyperparameter sequence, a corresponding feature selection model is generated. Mk ( λn ), n=1-100; k=1-5, k represents the number of rounds; n represents the sequence number of λ value, l is one of the hyperparameters;
[0056] Feature selection module for each value in the hyperparameter sequence λn , input the first internal validation set corresponding to each round into the feature selection model of the corresponding round in the 5 feature selection models Mk ( λn ), and obtain each λn The error rate corresponding to the value on the first internal validation set of the corresponding round; the average error rate corresponding to each value is obtained from the 5 error rates corresponding to each value; based on the minimum average error rate among all values, the value with the lowest error rate is selected l .min, where the error rate is the lowest. l .min, get the corresponding first bacterial community characteristic genus set with non-zero coefficients.
[0057] Furthermore, in the above system, the feature selection module is used to perform LDA Effect Size analysis (LEfSe) on the stool samples in the second training data set, calculate the linear discriminant analysis (LDA) score of each bacterial genus between groups, and select bacterial genera with LDA>2 or LDA<–2 and p<0.05 to enter the second preliminary screening set;
[0058] The feature selection module is configured to divide the second genus-level relative abundance matrix of the second primary screening set into five non-overlapping data sets, each time using one of the five non-overlapping data sets that has not been selected as a second internal validation set, and each time using the remaining four data sets as a second training set; wherein the first and second internal validation sets and the second training set both include non-overlapping Crohn's disease subjects who have had a recurrent Clostridium difficile infection after a previous Clostridium difficile infection and subjects who have not had a relapse;
[0059] The feature selection module is used to train the feature selection model for 5 rounds based on the second training set each time, using the Lasso feature selection algorithm, to obtain the corresponding 5-round feature selection model. In each round of feature selection model, for each λn value in the hyperparameter sequence, a corresponding feature selection model is generated. Mk ( λn ), n=1-100; k=1-5, k represents the number of rounds; n represents the sequence number of λ value, l is one of the hyperparameters;
[0060] Feature selection module for each value in the hyperparameter sequence λn , input the second internal validation set corresponding to each round into the feature selection model of the corresponding round in the 5 feature selection models Mk ( λn ), and obtain each λnThe error rate corresponding to the value on the second internal validation set of the corresponding round; the average error rate corresponding to each value is obtained from the 5 error rates corresponding to each value; based on the minimum average error rate among all values, the value with the lowest error rate is selected l .min, where the error rate is the lowest. l .min, and obtain the second characteristic genus set corresponding to the non-zero coefficient.
[0061] Here. The feature selection module is used to divide the first characteristic bacterial genus set or the second characteristic bacterial genus set into five non-overlapping data sets, perform five-fold cross-validation, use the Lasso regression algorithm to model each round of training set, and generate feature selection models Mk(λn) for 100 hyperparameter λ values, where n=1–100 and k=1–5.
[0062] The feature selection module is used to calculate the average false positive rate on the five-round internal validation set for each λn value, select the λ.min value corresponding to the lowest average false positive rate, and output the corresponding bacterial genus as the final feature bacterial genus set.
[0063] Furthermore, in the above system, the intestinal flora genus after the feature selection includes but is not limited to:
[0064] Bacterial genera used for CDI prediction include: Ruminococcaceae UCG-013, Akkermansia, Faecalibacterium, etc. (decreased abundance) and Veillonella, Bifidobacterium, Anaerostipes, Holdemanella, etc. (increased abundance);
[0065] Furthermore, in the above system, the intestinal flora genus after the feature selection is the core characteristic flora genus relied on for constructing the infection recognition model, wherein:
[0066] The characteristic bacterial group genera used in the Clostridium difficile infection (CDI) prediction model, namely the first characteristic bacterial group, include the following 19 bacterial genera: Veillonella; Bifidobacterium; Anaerostipes; Holdemanella; Romboutsia; Erysipelotrichaceae UCG-003; [Clostridium] innocuumgroup; Pyramidobacter; Parabacteroides; Weissella; Peptostreptococcus; Prevotella; Collinsella; Coprococcus; Bacteroides; Eubacterium eligens group); Eubacterium hallii group; Akkermansia; Ruminococcaceae UCG-013.
[0067] The characteristic bacterial genera used in the recurrent infection (rCDI) prediction model, namely the second characteristic bacterial genus set, include the following 10 bacterial genera: Veillonella; Megamonas; Marvinbryantia; Dorea; Lachnoclostridium; Butyricicoccus; Romboutsia; Anaerotruncus; Clostridium sensu stricto 1; and Staphylococcus.
[0068] Furthermore, in the above system, the training module runs the glmnet package on the R software (version 4.3.1) platform and uses the ElasticNet method for training. ElasticNet integrates dual constraints of L1 regularization (Lasso) and L2 regularization (Ridge), and the hyperparameters are set as follows: the α (alpha) value is set to 0.5, that is, the L1:L2 ratio is 1:1; λ2 (lambda) is selected through 5-fold cross-validation to determine the optimal penalty parameter of the model.
[0069] A training module is configured to train, based on a first training data set, a Clostridium difficile infection (CDI) infection risk prediction model for identifying the risk of Clostridium difficile infection (CDI) in Crohn's disease patients. The CDI infection risk prediction model uses an ElasticNet algorithm on an R (4.3.1) software platform and outputs a bacterial community feature-based infection risk score P1 based on the selected bacterial community genera and their relative abundances of each subject with and without CDI events in the first training data set. P1 ranges from 0 to 1, representing a low to high probability of infection, and is calculated as follows:
[0070]
[0071] Among them, P 1,i represents the CDI infection risk score of the i-th subject; x ij is the relative abundance of the jth characteristic genus in the first characteristic genus for the i-th subject; β j represents the regression coefficient of the jth characteristic genus in the ElasticNet model corresponding to the first characteristic genus, β0 is the bias term; σ() is the sigmoid function, which maps the output to between 0 and 1;
[0072] A training module is configured to train, based on a second training data set, a rCDI recurrence risk prediction model for identifying the risk of recurrent infection (rCDI) in patients with prior CDI, wherein the rCDI recurrence risk prediction model utilizes an ElasticNet algorithm on the R (4.3.1) software platform and outputs a recurrence risk score P2 based on selected bacterial genera and their relative abundances from each subject with and without rCDI in the second training data set. P2 ranges from 0 to 1, representing a low to high probability of infection, and is calculated as follows:
[0073]
[0074] Among them, P 2,i represents the recurrence risk score of the i-th previous CDI patient; x ik is the relative abundance of the kth characteristic bacterial genus for the i-th subject; γ k represents the regression coefficient of the corresponding bacterial genus in the ElasticNet model, γ0 is the bias term; σ() is the sigmoid function, which maps the output to between 0 and 1;
[0075] The training module runs the glmnet package and uses the ElasticNet method for training, wherein ElasticNet integrates dual constraints of L1 regularization and L2 regularization, and the hyperparameters are set as follows: the α value is set to 0.5, that is, the L1:L2 ratio is 1:1; λ2 selects λ2.min through 5-fold cross-validation to determine the optimal penalty parameter of the model; here, the implementation of the ElasticNet algorithm relies on the glmnet software package, wherein the hyperparameter α used to control the ratio of L1 regularization and L2 regularization is set to 0.5, and the λ2 value used to control the penalty intensity of the regression coefficient is based on the 5-fold cross-validation to determine the optimal λ2.min value.
[0076] The training module is further configured to obtain the recognition performance of the infection risk score of the trained CDI infection risk prediction model on the first training data set and the first test set, respectively; if the difference in recognition performance of the trained CDI infection risk prediction model on the first training data set and the first test set is less than a preset difference threshold, the trained CDI infection risk prediction model is used as the final CDI infection risk prediction model;
[0077] The training module is also used to obtain the recognition performance of the infection risk score of the trained rCDI recurrence risk prediction model on the second training data set and the second test set respectively; if the recognition performance of the trained rCDI infection risk prediction model on the second training data set and the second test set is less than a preset difference threshold, the trained rCDI recurrence risk prediction model is used as the final rCDI recurrence risk prediction model.
[0078] The identification module is used to obtain the relative abundance data of the characteristic bacterial genera of the subject to be tested and input it into the trained CDI infection prediction model and rCDI recurrence prediction model to calculate the infection risk score P1 and recurrence risk score P2, respectively, to assist in determining whether the subject is at risk of initial CDI infection or recurrent infection.
[0079] The identification module includes the following functions:
[0080] Risk score calculation: For the input microbiome characteristic data of the test subject, the CDI infection prediction model and the rCDI recurrence prediction model are used to make predictions and calculate the corresponding:
[0081]
[0082] Among them, x j ,x k is the relative abundance of the corresponding characteristic bacterial genus; β j ,γ kis the weight coefficient learned in the training model; σ() is the Sigmoid function. The score is a continuous value between 0 and 1, indicating the relative probability of infection or recurrence.
[0083] Threshold determination mechanism: Based on model training and ROC curve analysis in an independent test set, the optimal Youden Index thresholds were set: τ1 = 0.718 for CDI risk and τ2 = 0.771 for rCDI risk. If P1 ≥ τ1, the patient is considered at high risk for CDI and should be managed with enhanced fecal microecological management or early intervention. If P2 ≥ τ2, the patient is at high risk for recurrence and is recommended to implement strict monitoring and recurrence prevention interventions. These thresholds can be dynamically optimized based on data from subsequent clinical trials.
[0084] Result output: The recognition module can present the scoring results in the form of electronic text, graphic dashboard or scale, so that doctors can intuitively understand the infection / recurrence risk and assist in deciding whether to conduct intervention treatment (such as fecal microbiota transplantation, probiotic supplementation, antibiotic tapering therapy, etc.).
[0085] According to another aspect of the present invention, a computer-readable storage medium is provided, on which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor is caused to perform the following steps: executing the steps of the modules of any one of the above-mentioned systems, including:
[0086] Step S1: The collection module collects fecal samples from subjects in the CDI group, non-CDI group, rCDI group, and NR group, extracts 16S rRNA sequencing data from the samples, and obtains the relative abundance matrix at the genus level of the bacterial community;
[0087] Step S2: Based on the bacterial genus abundance information, the feature selection module uses Lasso cross-validation and linear discriminant analysis (LDA) to jointly screen out the intestinal bacterial genera that are significantly different from the CDI or rCDI status as characteristic bacterial clusters;
[0088] Step S3: The training module constructs a CDI infection risk model or rCDI recurrence risk model based on the screened characteristic bacterial genera and their relative abundance information. The model is trained using the ElasticNet algorithm in R software (version 4.3.1) (alpha = 0.5, λ2 is selected through cross-validation.min), and outputs an infection risk score P, P ∈ [0,1];
[0089] Step S4: The recognition module receives the relative abundance of bacterial genus of the subject to be tested, substitutes it into the model, and outputs an infection probability score to assist in identifying the CDI infection status or rCDI recurrence risk.
[0090] According to another aspect of the present invention, there is also provided an application of an intestinal flora biomarker in an infection identification product, wherein the intestinal flora biomarker includes at least the following bacterial genera:
[0091] Bacterial genera used in the CDI infection identification model: Veillonella; Bifidobacterium; Anaerostipes; Holdemanella; Romboutsia; Erysipelotrichaceae UCG-003; [Clostridium] innocuum group; Pyramidobacter; Parabacteroides; Weissella; Peptostreptococcus; Prevotella; Collinsella; Coprococcus; Bacteroides; Eubacterium eligens group; Eubacterium hallii group); Akkermansia; Ruminococcaceae UCG-013.
[0092] The genera used in the rCDI recurrence prediction model are: Veillonella; Megamonas; Marvinbryantia; Dorea; Lachnoclostridium; Butyricicoccus; Romboutsia; Anaerotruncus; Clostridium sensu stricto 1; and Staphylococcus.
[0093] The above bacterial genera can be used as predictive biomarkers for CDI infection or rCDI recurrence.
[0094] According to another aspect of the present invention, a stool microbiome analysis and detection kit for identifying the risk of CDI or rCDI infection is provided, comprising:
[0095] A stool collection assembly, comprising a stool sample collection tube and DNA stabilization solution for the individual to be tested;
[0096] DNA extraction kit for obtaining intestinal microbial DNA based on stool sample collection tubes and DNA stabilization solution;
[0097] The amplification and sequencing component is used to amplify and sequence the intestinal microbial DNA to obtain the genus-level relative abundance information of the first characteristic bacterial genus set or the second characteristic bacterial genus set; wherein,
[0098] The first characteristic bacterial genus set at least includes:
[0099] Veillonella; Bifidobacterium; Anaerostipes; Holdemanella; Romboutsia; Erysipelotrichaceae UCG-003; Clostridium innocuum group; Pyramidobacter; Parabacteroides; Weissella; Peptostreptococcus; Prevotella; Collinsella; Coprococcus; Bacteroides; Eubacterium eligens group; Eubacterium hallii group); Akkermansia and Ruminococcaceae UCG-013;
[0100] The second characteristic bacterial genus set includes at least:
[0101] Veillonella; Megamonas; Marvinbryantia; Dorea; Lachnoclostridium; Butyricicoccus; Romboutsia; Anaerotruncus; Clostridium sensu stricto 1; and Staphylococcus.
[0102] Compared to existing technologies, this invention, based on large-scale prospective clinical samples and combined with 16S sequencing data from the intestinal microbiome, systematically constructed two infection risk prediction models specifically for Crohn's disease patients, one for predicting primary Clostridium difficile infection (CDI) and the other for predicting recurrent Clostridium difficile infection (rCDI). Using the ElasticNet algorithm in machine learning, this invention extracts the most discriminative microbial signatures from genus-level microbiome data, enabling accurate identification of CDI and rCDI status. The model demonstrated excellent performance metrics (AUC, sensitivity, specificity, F1 value, etc.) in both training and independent test sets. The fecal microbiome recognition model proposed in this invention not only has high sensitivity and specificity but also demonstrates good generalizability and clinical practical value.
[0103] Compared to CDI diagnostic methods that rely on patient symptoms or traditional pathogen testing, this method proactively identifies risks based on imbalances in the intestinal flora, providing early warning before a definitive diagnosis of infection is made. Furthermore, this method relies on only a small stool sample and, combined with high-throughput 16S rRNA sequencing, can obtain the required data. The method is non-invasive, simple to use, and cost-effective, demonstrating its potential for widespread adoption in primary care settings.
[0104] Furthermore, the model outputs a continuous quantitative score (infection probability value P), which can assist clinicians in determining the severity of individual microbiome disturbances and the timing of intervention, providing a basis for treatment strategies such as individualized fecal microbiota transplantation and probiotic administration. As a supplement to traditional pathogen diagnostic methods, this invention has broad translational potential in the management of CDI and rCDI throughout the entire process, and also provides a new technical approach for the predictive application of intestinal microbiome markers in the treatment of concurrent infections in inflammatory bowel disease. BRIEF DESCRIPTION OF THE DRAWINGS
[0105] Figure 1 This is a histogram of LDA analysis of intestinal flora genus level corresponding to whether Clostridium difficile infection (CDI) occurs in Crohn's disease patients in one embodiment of the present invention;
[0106] Figure 2 This is a histogram of LDA analysis of intestinal flora genus level corresponding to whether recurrent infection (rCDI) occurs in Crohn's disease patients with previous CDI in one embodiment of the present invention;
[0107] Figure 3 The CDI infection risk prediction model in one embodiment of the present invention identifies the occurrence of CDI on the first training data set and the first test set;
[0108] Figure 4 FIG1 is the performance of the rCDI recurrence risk prediction model in identifying rCDI on the second training data set and the second test data set in one embodiment of the present invention. DETAILED DESCRIPTION
[0109] The present invention is further described in detail below with reference to the accompanying drawings.
[0110] In a typical configuration of the present application, the terminal, the device of the service network and the trusted party all include one or more processors (CPUs), input / output interfaces, network interfaces and memories.
[0111] Memory may include non-permanent storage in a computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.
[0112] Computer-readable media includes both permanent and non-permanent, removable and non-removable media that can be implemented using any method or technology for information storage. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change RAM (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include non-transitory computer-readable media, such as modulated data signals and carrier waves.
[0113] The inventors discovered for the first time in their research that the risk of Clostridium difficile infection and recurrence in Crohn's disease patients is closely associated with structural changes in a specific group of bacterial genera in the intestine. Specifically, through microbial sequencing and systems bioinformatics analysis, the present invention identified a set of microbial signatures that are stably and differentially expressed in Crohn's disease (CDI) and rCDI states, encompassing a variety of metabolic and structural genus-level microorganisms. These signatures were statistically validated as significant across multicenter samples. These microbial signatures could serve as biomarker combinations for early identification of infection risk.
[0114] Based on this, the present invention further constructed a CDI and rCDI risk prediction model based on this bacterial flora combination, and established a corresponding combined molecular detection strategy, filling the current technical gap in the prediction and evaluation methods of Clostridium difficile infection and recurrent infection in high-risk populations for inflammatory bowel disease.
[0115] To further clarify the objectives, technical solutions, and advantages of the embodiments of the present invention, various embodiments of the present invention are described in detail below with reference to the examples. Experimental methods in the examples where specific conditions are not specified are generally performed under conventional conditions, such as those described in textbooks and experimental manuals, or under conditions recommended by the manufacturer. The methods are also performed under the recommended conditions of the supporting software.
[0116] In a typical embodiment of the present invention, a system for constructing an intestinal flora identification model, a storage medium, and a molecular detection auxiliary tool for predicting the risk of Clostridioides difficile infection (CDI) and recurrent CDI (rCDI) in Crohn's disease patients are provided, aiming to achieve early warning, precise intervention, and individualized management based on intestinal microecological characteristics. The method includes steps S1 to S4.
[0117] Step S1: A collection module is used to collect stool samples from Crohn's disease subjects, wherein the stool samples include: stool samples from Crohn's disease subjects with Clostridium difficile infection (CDI) and those without CDI during diagnosis or follow-up, as first samples, and stool samples from Crohn's disease subjects with recurrent Clostridium difficile infection (rCDI) and those without recurrence (NR) after a previous Clostridium difficile infection, as second samples; genus-level relative abundance information of the bacterial flora of the first sample and the second sample is obtained respectively;
[0118] Here, the collection module was used to collect fresh stool samples from patients with Crohn's disease and obtain bacterial flora abundance data through 16S rRNA amplification sequencing.
[0119] Preferably, the acquisition module is further used to obtain a first genus-level relative abundance matrix after 16S rRNA sequencing of the first sample; obtain a second genus-level relative abundance matrix after 16S rRNA sequencing of the second sample; split all first samples into a first non-overlapping training data set and a first test set; split all second samples into a second non-overlapping training data set and a second test set.
[0120] Optionally, all first samples may be split into a first training data set and a first test set that do not overlap with each other in a ratio of 8:2;
[0121] All second samples may be split into a second training data set and a second test set that do not overlap with each other in a ratio of 8:2.
[0122] Specific processes may include:
[0123] Step S11, grouping the subjects according to whether they were diagnosed with CDI or rCDI recurrence, wherein:
[0124] The CD group refers to patients with Crohn's disease only (Crohn's disease subjects without CDI), and the CD+CDI group refers to patients with Crohn's disease combined with Clostridium difficile infection (Crohn's disease subjects with Clostridium difficile infection);
[0125] The NR group consisted of patients with a history of CDI but no recurrence (non-recurrent Crohn's disease subjects), and the rCDI group consisted of patients with recurrent infection (previous C. difficile infection-followed recurrent C. difficile infection-followed Crohn's disease subjects).
[0126] In step S12, the stool sample was processed by a standard DNA extraction process and amplified using the 16S rRNA V3-V4 region with primers 341F / 806R;
[0127] Step S13: The sequencing platform is Illumina MiSeq or NovaSeq. The raw sequences are denoised by bioinformatics pipelines such as DADA2 and clustered into ASVs. They are then mapped to the Greengenes or SILVA database to obtain genus-level features.
[0128] In step S14, the standardized relative abundance information (such as % value) of the bacterial genus level is obtained for each sample, which is used as the core microbiome variable for subsequent modeling input.
[0129] Preferably, the data structure of the genus-level relative abundance information of the bacterial genus level output in step S14 is an abundance matrix of several rows (samples) × several columns (bacterial genera), which is input as a training data set.
[0130] Step S2: A feature selection module is used to select a first characteristic genus set based on the genus-level relative abundance information of the first sample by using Lasso cross-validation and linear discriminant analysis (LDA); and to select a second characteristic genus set based on the genus-level relative abundance information of the second sample by using Lasso cross-validation and linear discriminant analysis (LDA);
[0131] Here, the feature selection module is used to identify intestinal bacterial genera significantly associated with CDI or rCDI status in the training set. The specific process is as follows:
[0132] Step S21: Perform LDAEffect Size analysis (LEfSe) on the stool samples in the first training data set and the second training data set, calculate the linear discriminant analysis (LDA) score of each bacterial genus between groups, and select bacterial genera with LDA>2 or LDA<–2 and p<0.05 to enter the corresponding primary screening set;
[0133] In step S22 , Lasso cross-validation was used to run the glmnet package in R (version 4.3.1) on the bacterial genera in the initial screening set. The optimal λ.min was determined through 5-fold cross-validation, and bacterial genera with non-zero coefficients were screened.
[0134] Preferably, the feature selection module is used to perform LDAEffect Size analysis (LEfSe) on the stool samples in the first training data set, calculate the linear discriminant analysis (LDA) score of each bacterial genus between groups, and select bacterial genera with LDA>2 or LDA<–2 and p<0.05 to enter the first preliminary screening set;
[0135] The feature selection module is configured to divide the first genus-level relative abundance matrix of the first preliminary screening set into five non-overlapping data sets, each time using one of the five non-overlapping data sets that has not been selected as a first internal validation set, and each time using the remaining four data sets as a first training set; wherein both the internal validation set and the training set include non-overlapping Crohn's disease subjects with and without CDI;
[0136] The feature selection module is used to train the feature selection model for 5 rounds based on the first training set each time, using the Lasso feature selection algorithm to obtain the corresponding 5-round feature selection model. In each round of feature selection model, for each λn value in the hyperparameter sequence, a corresponding feature selection model is generated. Mk ( λn ), n=1-100; k=1-5, k represents the number of rounds; n represents the sequence number of λ value, l is one of the hyperparameters;
[0137] Feature selection module for each value in the hyperparameter sequence λn , input the first internal validation set corresponding to each round into the feature selection model of the corresponding round in the 5 feature selection models Mk ( λn ), and obtain each λn The error rate corresponding to the value on the internal validation set of the corresponding round; the average error rate corresponding to each value is obtained from the 5 error rates corresponding to each value; based on the minimum average error rate among all values, the value with the lowest error rate is selected l .min, where the error rate is the lowest. l .min, get the corresponding first bacterial community characteristic genus set with non-zero coefficients.
[0138] Preferably, the feature selection module is used to perform LDAEffect Size analysis (LEfSe) on the stool samples in the second training data set, calculate the linear discriminant analysis (LDA) score of each bacterial genus between groups, and select bacterial genera with LDA>2 or LDA<–2 and p<0.05 to enter the second preliminary screening set;
[0139] The feature selection module is configured to divide the second genus-level relative abundance matrix of the second primary screening set into five non-overlapping data sets, each time using one of the five non-overlapping data sets that has not been selected as a second internal validation set, and each time using the remaining four data sets as a second training set; wherein the first and second internal validation sets and the second training set both include non-overlapping Crohn's disease subjects who have had a recurrent Clostridium difficile infection after a previous Clostridium difficile infection and subjects who have not had a relapse;
[0140] The feature selection module is used to train the feature selection model for 5 rounds based on the second training set each time, using the Lasso feature selection algorithm, to obtain the corresponding 5-round feature selection model. In each round of feature selection model, for each λn value in the hyperparameter sequence, a corresponding feature selection model is generated. Mk ( λn ), n=1-100; k=1-5, k represents the number of rounds; n represents the sequence number of λ value, l is one of the hyperparameters;
[0141] Feature selection module for each value in the hyperparameter sequence λn , input the internal validation set corresponding to each round into the feature selection model of the corresponding round in the 5 feature selection models Mk ( λn ), and obtain each λn The error rate corresponding to the value on the second internal validation set of the corresponding round; the average error rate corresponding to each value is obtained from the 5 error rates corresponding to each value; based on the minimum average error rate among all values, the value with the lowest error rate is selected l .min, where the error rate is the lowest. l .min, and obtain the second characteristic genus set corresponding to the non-zero coefficient.
[0142] Figure 1 This is a bar chart of LDA analysis of intestinal flora genus level corresponding to whether Clostridium difficile infection (CDI) occurs in Crohn's disease patients in one embodiment of the present invention, wherein the right side of the vertical axis represents the bacterial genera significantly enriched in patients with combined CDI, the left side of the vertical axis represents the bacterial genera enriched only in Crohn's disease patients, and the horizontal axis represents the LDA score (log10 transformation), which indicates the genus discrimination ability.
[0143] like Figure 1As shown, preferably, the first characteristic genus set, i.e., the CDI prediction characteristic genus, includes:
[0144] Veillonella; Bifidobacterium; Anaerostipes; Holdemanella; Romboutsia; Erysipelotrichaceae UCG-003; [Clostridium] innocuum group; Pyramidobacter; Parabacteroides; Weissella; Peptostreptococcus; Prevotella; Collinsella; Coprococcus; Bacteroides; Eubacterium eligens group; Eubacterium hallii group); Akkermansia; Ruminococcaceae UCG-013.
[0145] Figure 2 This is a bar chart of LDA analysis of intestinal flora genus levels corresponding to whether recurrent infection (rCDI) occurs in Crohn's disease patients with previous CDI in one embodiment of the present invention. The right side of the vertical axis represents the bacterial genera significantly enriched in the rCDI group, and the left side of the vertical axis represents the bacterial genera enriched in the non-recurrent (NR) group. The horizontal axis represents the LDA score (log10 transformation), which indicates the ability to distinguish bacterial genera.
[0146] Better, Figure 2 As shown, the second characteristic genus set, i.e., the rCDI predicted characteristic genus, includes:
[0147] Veillonella; Megamonas; Marvinbryantia; Dorea; Lachnoclostridium; Butyricicoccus; Romboutsia; Anaerotruncus; Clostridium sensu stricto 1; Staphylococcus.
[0148] The bacterial genus finally determined was used as the input feature for modeling, and the genus-level microbiological indicators were log10 transformed and then entered the model training stage.
[0149] Step S3: Training module
[0150] The training module is used to construct two independent prediction models based on the collected genus-level microbial abundance characteristics, corresponding to the CDI infection risk scoring model and the rCDI recurrence risk scoring model, respectively. The ElasticNet algorithm is used to achieve microbial feature selection and modeling optimization.
[0151] Step S3.1: Model algorithm and platform
[0152] The training module runs the glmnet package on the R software platform (version 4.3.1) and uses the ElasticNet method for training. ElasticNet integrates dual constraints of L1 regularization (Lasso) and L2 regularization (Ridge), which is suitable for modeling high-dimensional and sparse features of bacterial communities, and combines variable selection and robustness. The hyperparameter settings are as follows:
[0153] The alpha value is set to 0.5, which means the L1:L2 ratio is 1:1;
[0154] λ2 (lambda) is selected through 5-fold cross-validation to determine the optimal penalty parameter of the model;
[0155] Here, in order to distinguish it from the λ used in the feature selection module, the λ used in the training module is expressed as λ2.
[0156] The input variables were the relative abundance data of bacterial communities at the genus level (%), which had been log10 transformed and Z-score standardized.
[0157] The training module is configured to develop a CDI infection risk prediction model based on the first training data set, outputting an infection risk score P1 representing the infection risk level; and to develop an rCDI recurrence risk prediction model based on the second training data set, outputting a recurrence risk score P2 representing the probability of recurrence. The scoring formula is as follows:
[0158] ;
[0159] ;
[0160] Among them, P 1,i represents the CDI infection risk score of the i-th subject;
[0161] x ijis the relative abundance information of the jth characteristic genus in the first characteristic genus for the i-th subject in the first sample; wherein, the relative abundance information of the jth characteristic genus in the first characteristic genus for the i-th subject in the first sample is the relative abundance information after log10 transformation and Z-score normalization;
[0162] β j is the weight coefficient of the jth characteristic bacterial genus in the first characteristic bacterial genus learned in the trained CDI infection risk prediction model;
[0163] β0 is the bias term;
[0164] P 2,i represents the recurrence risk score of the i-th patient with previous CDI;
[0165] x ik is the relative abundance information of the kth characteristic genus in the second characteristic genus for the i-th subject of the second sample; wherein, the relative abundance information of the kth characteristic genus in the second characteristic genus for the i-th subject of the second sample is the relative abundance information after log10 transformation and Z-score normalization;
[0166] γ k is the weight coefficient of the kth characteristic bacterial genus in the second characteristic bacterial genus learned in the trained rCDI recurrence risk prediction model;
[0167] σ() is the Sigmoid function, which maps the output to between 0 and 1;
[0168] The infection risk score P1 and recurrence risk score P2 are continuous variables ranging from 0 to 1;
[0169] Specifically, step S3.2: CDI infection risk prediction model training
[0170] The CD group (n=64) and the CD+CDI group (n=50) can be combined into the first training data set (a total of 114 samples), and a binary classification model can be constructed with whether CDI occurs as the dependent variable (0 / 1) and the genus abundance as the independent variable.
[0171] During training, the ElasticNet model outputs a CDI infection probability score P1 for each sample, ranging from 0 to 1, indicating a low to high risk of infection. P1 is calculated as follows:
[0172] Formula (1):
[0173]
[0174] Where: σ() is the Sigmoid function: ;X i is the standardized relative abundance of the i-th genus; β i is the regression coefficient corresponding to the genus; b(β 0 ) is the model bias term; n is the number of bacterial genera included in the model.
[0175] Figure 3 The CDI infection risk prediction model in one embodiment of the present invention identifies the occurrence of CDI on a first training data set and a first test set.
[0176] After the training is completed, Figure 3 As shown, performance on the first training dataset was as follows: AUC (area under the curve): 0.892 (95% CI: 0.880–0.904); sensitivity: 0.810, specificity: 0.908, and accuracy: 0.898. In the independent first test dataset (n = 58), the model prediction AUC was 0.890 (95% CI: 0.818–0.963), with a DeLong test P = 0.886, indicating no risk of overfitting.
[0177] Step S3.3: rCDI recurrence risk prediction model training
[0178] The rCDI model used the NR group (n=65) and the rCDI group (n=58) as the second training data set (a total of 123 samples), with recurrence as the dependent variable (0 / 1). The training logic was the same as that of the CDI model, and the output score P2 represented the recurrence risk level.
[0179] Formula (2):
[0180]
[0181] Where: X j : rCDI predicts the normalized abundance of related genus-level bacterial communities; γ j : regression coefficient of each genus; b′(γ0): model bias term; m is the number of bacterial genera included in the rCDI model.
[0182] Figure 4 FIG1 is the performance of the rCDI recurrence risk prediction model in identifying rCDI on the second training data set and the second test data set in one embodiment of the present invention.
[0183] After the training is completed, Figure 4As shown, performance on the second training dataset was as follows: AUC (area under the curve): 0.892 (95% CI: 0.816–0.968); sensitivity: 0.861, specificity: 0.910; accuracy: 0.893, F1-value: 0.82. In the independent second test set (n = 41), the model prediction AUC was 0.862 (95% CI: 0.849–0.874), with DeLong's test P = 0.406, indicating that the model was not significantly overfitting and had good generalization performance.
[0184] The scores P1 and P2 are both continuous variables (range 0-1), and the optimal discrimination threshold of the Youden index can be set (τ1=0.718 for CDI risk and τ2=0.771 for rCDI risk): If P1≥τ1, the patient is judged to be at high risk of CDI, and fecal microecological management or early intervention should be strengthened; if P2≥τ2, it indicates that there is a high possibility of recurrence, and strict monitoring and anti-recurrence intervention plan are recommended.
[0185] Step S3.4: Model performance evaluation and decision support mechanism construction
[0186] The training module is further configured to obtain the recognition performance of the infection risk score of the trained CDI infection risk prediction model on the first training data set and the first test set, respectively; if the difference in recognition performance of the trained CDI infection risk prediction model on the first training data set and the first test set is less than a preset difference threshold, the trained CDI infection risk prediction model is used as the final CDI infection risk prediction model;
[0187] The training module is also used to obtain the recognition performance of the infection risk score of the trained rCDI recurrence risk prediction model on the second training data set and the second test set respectively; if the recognition performance of the trained rCDI infection risk prediction model on the second training data set and the second test set is less than a preset difference threshold, the trained rCDI recurrence risk prediction model is used as the final rCDI recurrence risk prediction model.
[0188] like Figure 3 and 4 As shown in the figure, the performance of the trained CDI infection risk model and rCDI recurrence risk model was verified by evaluating the training data set and an independent test set. Model performance evaluation metrics include: AUC (area under the curve), sensitivity, specificity, accuracy, F1 value, average error rate (ERR), and Youden Index. Their specific meanings are as follows:
[0189] AUC indicates the ability of the model to distinguish positive and negative samples at different thresholds. The closer the AUC value is to 1, the better the model performance.
[0190] Sensitivity (Recall) indicates the ability of the model to identify positive patients (CDI or rCDI patients);
[0191] Specificity indicates the ability of the model to identify negative patients (non-infected or non-relapsed patients);
[0192] The accuracy is the proportion of samples correctly classified by the model to the total samples;
[0193] The F1 value comprehensively considers the precision and recall rate;
[0194] The average false positive rate is the mean of the model's false positive rate;
[0195] Youden index J is used to measure the overall recognition performance of the model, J = sensitivity + specificity - 1.
[0196] In the embodiment of the present invention, the model achieved high levels of indicators such as AUC, accuracy, and F1 value on the training set and the test set, and there was no statistically significant difference in performance between the training data set and the test set (DeLong test P values were all > 0.05), indicating that the model had no obvious overfitting and had good generalization ability, as shown in the following table:
[0197] Table 1. Performance comparison of models
[0198]
[0199] Step S3.5: Risk score-based intervention recommendation mechanism
[0200] This invention not only realizes the individualized risk assessment of CDI and rCDI, but also establishes a clear intervention guidance mechanism by setting the discrimination thresholds (τ1, τ2):
[0201] When P1 ≥ 0.718 (CDI risk score threshold), it is recommended to initiate preventive fecal microbiota transplantation, probiotic intervention, or close microecological monitoring for the Crohn's disease patient;
[0202] When P2 ≥ 0.771 (rCDI risk score threshold), the patient is at significant risk of recurrence and should receive intensive management during the recovery period, such as extended antimicrobial therapy, personalized dietary support, and interventions in the intestinal inflammatory environment. Continuously outputting risk scores and enabling quantifiable grading, these scores provide physicians with microbiome-based evidence to support their intervention decisions, enabling the transformation of microbiome-based precision medicine from "post-illness treatment" to "pre-illness warning."
[0203] According to another aspect of the present invention, a computer-readable storage medium is provided, on which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor is caused to perform the following steps:
[0204] Step S1: Acquisition module
[0205] The collection module is used to collect fresh stool samples from Crohn's disease patients. Through 16S rRNA high-throughput sequencing technology, the relative abundance data of the bacterial community at the genus level for each sample is obtained, and a standardized genus abundance matrix is constructed as input for subsequent analysis.
[0206] Step S2: Feature selection module
[0207] The feature selection module groups patients based on whether they have Clostridium difficile infection (CDI) or recurrent infection (rCDI). LEfSe analysis is used to screen out differential bacterial genera with LDA scores greater than 2 and p < 0.05. The Lasso algorithm is then used to further compress the feature dimensions, ultimately obtaining genus-level bacterial community features for CDI and rCDI prediction.
[0208] Step S3: Training module
[0209] The training module used the ElasticNet algorithm to construct a CDI infection risk prediction model and a rCDI recurrence risk model. The model inputs were standardized genus-level bacterial flora characteristics, and the outputs were continuous variable risk scores P1 and P2, ranging from 0 to 1, representing the risk levels of CDI infection and rCDI recurrence, respectively. During model training, 5-fold cross-validation was used to determine the minimum value of λ2.min to avoid overfitting.
[0210] Step S4: Identify module
[0211] The recognition module receives genus-level relative abundance data for Crohn's disease patients and inputs it into the trained model to calculate their CDI infection risk score (P1) and rCDI recurrence risk score (P2). A P1 score ≥ 0.718 indicates a high risk of CDI infection; a P2 score ≥ 0.771 indicates a high risk of recurrence and suggests intervention.
[0212] According to another aspect of the present invention, a computer-readable storage medium is provided, on which computer-executable instructions are stored. When the instructions are executed by a processor, the processor executes each module of any one of the above-mentioned systems.
[0213] According to another aspect of the present invention, there is also provided a product for identifying the risk of CDI infection and rCDI recurrence in Crohn's disease patients using genus-level features of intestinal flora. The genus-level features include at least the following:
[0214] CDI prediction characteristic bacterial genera include:
[0215] Veillonella; Bifidobacterium; Anaerostipes; Holdemanella; Romboutsia; Erysipelotrichaceae UCG-003; Clostridium innocuum group; Pyramidobacter; Parabacteroides; Weissella; Peptostreptococcus; Prevotella; Collinsella; Coprococcus; Bacteroides; Eubacterium eligens group; Eubacterium hallii group); Akkermansia; Ruminococcaceae UCG-013.
[0216] The rCDI prediction characteristic bacterial genera include:
[0217] Veillonella; Megamonas; Marvinbryantia; Dorea; Lachnoclostridium; Butyricicoccus; Romboutsia; Anaerotruncus; Clostridium sensu stricto 1; Staphylococcus.
[0218] According to another aspect of the present invention, a microecological detection kit for CDI and rCDI risk scoring is also provided, the detection kit comprising:
[0219] A stool collection assembly, comprising a stool sample collection tube and DNA stabilization solution for the individual to be tested;
[0220] DNA extraction kit for obtaining intestinal microbial DNA based on stool sample collection tubes and DNA stabilization solution;
[0221] Amplification and sequencing components are used to amplify and sequence the 16S rRNA V3–V4 region of intestinal microbial DNA to obtain genus-level relative abundance information for the first characteristic genus set or the second characteristic genus set. Here, the 16S rRNA V3-V4 region can be amplified and sequenced to obtain ASVs and map genus-level characteristics;
[0222] The data analysis module obtains the genus-level relative abundance information of the first characteristic bacterial genus set of the individual to be tested and inputs it into the trained CDI infection risk prediction model to output the infection risk score P1. The data analysis module obtains the genus-level relative abundance information of the second characteristic bacterial genus set of the individual to be tested and inputs it into the trained rCDI recurrence risk prediction model to output the recurrence risk score P2.
[0223] The result output module presents the infection risk score P1 or recurrence risk score P2 in text or visual form.
[0224] In summary, the present invention utilizes a machine learning approach based on genus-level abundance characteristics of the intestinal microbiome to construct a CDI infection risk scoring model and a rCDI recurrence risk scoring model, respectively, to quantitatively assess individual patients' infection risk and likelihood of recurrence. The system utilizes a standardized microbiome data processing pipeline, extracts key genus signatures using feature-based filtering algorithms such as Lasso and LEfSe, and generates continuous risk scores (P1 and P2) using ElasticNet modeling. Its predictive performance demonstrates excellent sensitivity, specificity, and accuracy in both training and independent test sets. Model performance has been validated using multidimensional metrics such as AUC, F1 value, and Youden index, demonstrating significantly superior predictive ability compared to traditional clinical indicators, enabling early identification of high-risk individuals.
[0225] This method requires only a small stool sample, and the detection process is compatible with the 16S sequencing platform. It is simple to operate, cost-effective, and readily applicable in clinical practice. The resulting risk score can be used to guide the development of personalized prevention and treatment strategies, such as microbiome intervention, probiotic prophylaxis, and fecal microbiota transplantation. It also has potential applications in CDI recurrence monitoring, antimicrobial regimen optimization, and precision microbiome management.
[0226] The system is highly versatile and integrated, and can be implemented through a software system or portable device. It can also be embedded in the hospital LIS system for rapid screening, providing key microecological evidence support for early screening, early diagnosis and early treatment of intestinal infectious diseases. It has good promotion value and industrial transformation prospects.
[0227] Obviously, those skilled in the art may make various changes and modifications to this application without departing from the spirit and scope of this application. Thus, if these modifications and variations of this application fall within the scope of the claims of this application and their equivalents, this application is intended to include these modifications and variations.
[0228] It should be noted that the present invention may be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of the present invention may be executed by a processor to implement the steps or functions described above. Similarly, the software program of the present invention (including related data structures) may be stored in a computer-readable recording medium, such as a RAM memory, a magnetic or optical drive, a floppy disk, or the like. In addition, some steps or functions of the present invention may be implemented in hardware, for example, as circuitry that cooperates with a processor to perform the various steps or functions.
[0229] In addition, a portion of the present invention may be applied as a computer program product, such as computer program instructions, which, when executed by a computer, can call or provide the method and / or technical solution according to the present invention through the operation of the computer. The program instructions for calling the method of the present invention may be stored in a fixed or removable recording medium, and / or transmitted through a data stream in a broadcast or other signal-carrying medium, and / or stored in a working memory of a computer device that operates according to the program instructions. Here, according to one embodiment of the present invention, a device is included, which includes a memory for storing computer program instructions and a processor for executing the program instructions, wherein, when the computer program instructions are executed by the processor, the device is triggered to operate based on the aforementioned methods and / or technical solutions according to multiple embodiments of the present invention.
[0230] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above and that the invention can be implemented in other specific forms without departing from the spirit or essential features of the invention. Therefore, from any point of view, the embodiments should be regarded as exemplary and non-restrictive, and the scope of the invention is defined by the appended claims rather than the above description, and it is intended that all changes that fall within the meaning and scope of the equivalents of the claims be encompassed within the present invention. Any figure marks in the claims should not be regarded as limiting the claims involved. In addition, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices stated in the device claim may also be implemented by one unit or device through software or hardware. Words such as first and second are used to indicate names and do not indicate any particular order.
Claims
1. A prediction model construction system for the occurrence and recurrent infection of Clostridium difficile infection, characterized in that: include: A collection module is used to collect stool samples from subjects with Crohn's disease and sequence the stool samples to obtain genus-level relative abundance information of the bacterial flora; The feature selection module is used to use the LEfSe and Lasso feature screening algorithms on the stool sample, and based on the genus-level relative abundance information of the bacterial community, screen out a characteristic genus set, wherein the characteristic genus set includes: a first characteristic genus set and a second characteristic genus set; the first characteristic genus set includes at least: Veillonella; Bifidobacterium; Anaerostipes; Holdemanella; Romboutsia; Erysipelotrichaceae UCG-003; Clostridium innocuum group); Pyramidobacter; Parabacteroides; Weissella; Peptostreptococcus; Prevotella; Collinsella; Coprococcus; Bacteroides; Eubacterium eligens group; Eubacterium hallii group; Akkermansia and Ruminococcaceae UCG-013); the second characteristic bacterial genus set includes at least: Veillonella; Megamonas; Marvinbryantia; Dorea; Lachnoclostridium; Butyricicoccus; Romboutsia; Anaerotruncus; Clostridium sensu stricto1 and Staphylococcus A training module is used to train a CDI infection risk prediction model and an rCDI recurrence risk prediction model using the ElasticNet method based on the genus-level relative abundance information of the characteristic bacterial genus set; The identification module is used to obtain the genus-level relative abundance information of the characteristic bacterial genus set of the individual to be tested, and input it into the trained CDI infection risk prediction model or rCDI recurrence risk prediction model to obtain the corresponding risk score.
2. The system according to claim 1, wherein The feature selection module is used to perform LEfSe analysis on fecal samples, calculate the linear discriminant analysis (LDA) score of each bacterial genus between groups, and obtain a preliminary screening set of bacterial genus based on the LDA score; based on the preliminary screening set of bacterial genus, use the Lasso feature selection algorithm to obtain a characteristic bacterial genus set.
3. The system according to claim 2, wherein: The stool samples include: stool samples of subjects with Clostridium difficile infection and without Crohn's disease during diagnosis or follow-up, as first samples; the collection module is used to obtain genus-level relative abundance information of the bacterial flora of the first sample; the first sample is split into a first training data set and a first test set; The feature selection module is used to perform LEfSe analysis on the fecal samples in the first training data set, calculate the linear discriminant analysis LDA score of each bacterial genus between groups, and select bacterial genera with LDA>2 or LDA<–2 and p<0.05 to enter the first preliminary screening set; based on the first genus-level relative abundance matrix of the first preliminary screening set, the Lasso feature selection algorithm is used to obtain the corresponding first feature bacterial genus set with non-zero coefficients.
4. The system according to claim 1, wherein: The training module is configured to perform a CDI infection risk prediction model based on the first training data set and output an infection risk score P1 representing the infection risk level; ; Among them, P 1,i represents the CDI infection risk score of the i-th subject; x ij is the relative abundance information of the jth characteristic genus in the first characteristic genus for the i-th subject in the first sample; wherein, the relative abundance information of the jth characteristic genus in the first characteristic genus for the i-th subject in the first sample is the relative abundance information after log10 transformation and Z-score normalization; β j is the weight coefficient of the jth characteristic bacterial genus in the first characteristic bacterial genus learned in the trained CDI infection risk prediction model; β0 is the bias term; The infection risk score P1 is a continuous variable between 0 and 1; The training module is also used to obtain the recognition performance of the infection risk score of the trained CDI infection risk prediction model on the first training data set and the first test set respectively; if the recognition performance of the trained CDI infection risk prediction model on the first training data set and the first test set is less than a preset difference threshold, the trained CDI infection risk prediction model will be used as the final CDI infection risk prediction model.
5. The system according to claim 2, wherein: The stool samples include: stool samples from subjects with recurrent Clostridium difficile infection or non-recurrent Crohn's disease after a previous Clostridium difficile infection, as second samples; the collection module is further used to obtain genus-level relative abundance information of the bacterial flora of the second sample; split the second sample into a second training data set and a second test set; The feature selection module is used to perform LEfSe analysis on the fecal samples in the second training data set, calculate the linear discriminant analysis LDA score of each bacterial genus between groups, and select bacterial genera with LDA>2 or LDA<–2 and p<0.05 to enter the second preliminary screening set; based on the second genus-level relative abundance matrix of the second preliminary screening set, a second characteristic bacterial genus set with corresponding non-zero coefficients is obtained.
6. The system according to claim 1, wherein: The training module is used to perform an rCDI recurrence risk prediction model based on the second training data set and output a recurrence risk score P2, which represents the probability of recurrence. The scoring formula is as follows: ; P 2,i represents the recurrence risk score of the i-th patient with previous CDI; x ik is the relative abundance information of the kth characteristic genus in the second characteristic genus for the i-th subject of the second sample; wherein, the relative abundance information of the kth characteristic genus in the second characteristic genus for the i-th subject of the second sample is the relative abundance information after log10 transformation and Z-score normalization; γ k is the weight coefficient of the kth characteristic bacterial genus in the second characteristic bacterial genus learned in the trained rCDI recurrence risk prediction model; σ() is the Sigmoid function, which maps the output to between 0 and 1; is the bias term; recurrence risk score P2, a continuous variable between 0 and 1; The training module is also used to obtain the recognition performance of the infection risk score of the trained rCDI recurrence risk prediction model on the second training data set and the second test set respectively; if the recognition performance of the trained rCDI infection risk prediction model on the second training data set and the second test set is less than a preset difference threshold, the trained rCDI recurrence risk prediction model is used as the final rCDI recurrence risk prediction model.
7. The system according to claim 1, wherein: The ElasticNet method integrates dual constraints of L1 regularization and L2 regularization, and the hyperparameters are set as follows: the α value is set to 0.5, that is, the L1:L2 ratio is 1:1; the λ2 value used to control the intensity of the penalty for the regression coefficient is determined based on the optimal λ2.min value based on 5-fold cross-validation.
8. The system according to any one of claims 1 to 7, characterized in that The threshold of the trained CDI infection risk prediction model is τ1 = 0.
718. If P1 ≥ τ1, the individual is judged to be at high risk of CDI. The threshold value of the trained rCDI recurrence risk prediction model is τ2 = 0.
771. If P2 ≥ τ2, it indicates that the individual to be tested has a high possibility of recurrence.
9. A computer-readable storage medium having computer-executable instructions stored thereon, wherein when the instructions are executed by a processor, the processor executes the modules of the system according to any one of claims 1 to 8.
10. A microecological detection kit, characterized in that: include: A stool collection assembly, comprising a stool sample collection tube and DNA stabilization solution for the individual to be tested; DNA extraction kit for obtaining intestinal microbial DNA based on stool sample collection tubes and DNA stabilization solution; The amplification and sequencing component is used to amplify and sequence the intestinal microbial DNA to obtain the genus-level relative abundance information of the first characteristic bacterial genus set or the second characteristic bacterial genus set; wherein, The first characteristic bacterial genus set at least includes: Veillonella; Bifidobacterium; Anaerostipes; Holdemanella; Romboutsia; Erysipelotrichaceae UCG-003; Clostridium innocuum group; Pyramidobacter; Parabacteroides; Weissella; Peptostreptococcus; Prevotella; Collinsella; Coprococcus; Bacteroides; Eubacterium eligens group; Eubacterium hallii group); Akkermansia and Ruminococcaceae UCG-013; The second characteristic bacterial genus set includes at least: Veillonella; Megamonas; Marvinbryantia; Dorea; Lachnoclostridium; Butyricicoccus; Romboutsia; Anaerotruncus; Clostridium sensu stricto 1; and Staphylococcus.
Citation Information
Patent Citations
Marking microorganisms for Crohn's disease in children and application of marking microorganisms
CN114085886A
Metabolic aging recognition model construction system, storage medium and kit
CN120280160A