Infection pathogen type detection system based on five markers
By combining machine learning algorithms with markers such as GSDMD, p-MLKL, IL-6, PCT, and CRP, a high-precision pathogen infection type detection system is built, which solves the problem of insufficient sensitivity and specificity of early diagnosis of blood flow infection in the prior art, and achieves high-precision pathogen type recognition.
Patent Information
- Application Number
- CN202511020965.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-24
- Publication Date
- 2025-08-22
- Estimated Expiration
- 2045-07-24
AI Technical Summary
The existing blood flow infection detection methods have insufficient sensitivity and specificity, making it difficult to accurately diagnose pathogen types, especially bacterial infections, and conventional biomarkers such as PCT have an increased lag in the early stage of infection, which cannot meet clinical diagnostic criteria.
A detection system based on five markers, including GSDMD, p-MLKL, IL-6, PCT, and CRP, is adopted to construct a random forest classification model or multi-layer feedforward neural network through machine learning algorithms, integrate the detection values of these markers, and establish a high-precision detection system for pathogen infection type.
It achieves high accuracy in the early diagnosis of blood flow infection, with a diagnostic accuracy of >95%, which is significantly better than traditional methods, and improves the identification accuracy and specificity of pathogen types.
Smart Images

Figure CN120526862A_ABST
Abstract
Description
Technical Field
[0001] The invention discloses a pathogen type detection system, which belongs to the technical field of microorganisms and artificial intelligence. Background Art
[0002] Early and accurate diagnosis of bloodstream infection (BSI) is crucial for improving prognosis. Conventional biomarkers currently tested in clinical practice include procalcitonin (PCT), acute phase protein (CRP), and interleukin-6 (IL-6). While PCT is currently recognized as the best biomarker for bacterial infection, with the highest diagnostic accuracy, its clinical utility remains limited: ① Its overall specificity for pathogen differentiation is only approximately 70% (meta-analysis data); ② It often exhibits a delayed increase within 24 hours of infection, making it more suitable for assessing sepsis severity rather than for early diagnosis of BSI; ③ Notably, non-infectious factors (such as burns, intestinal ischemia, and postoperative trauma) can also lead to abnormally elevated PCT levels. Therefore, the sensitivity (generally <85%) and specificity (usually <90%) of existing biomarkers do not meet the diagnostic technical standards for BSI recommended by the European Society for Clinical Microbiology and Infectious Diseases (ESCMID), urging the discovery and validation of new biomarkers.
[0003] The purpose of the present invention is to provide a more sensitive and specific pathogen type detection system. Summary of the Invention
[0004] Based on the above objectives, the present invention first provides a pathogen infection type detection system based on five markers, including GSDMD, p-MLKL, IL-6, PCT, and CRP. The system includes the following modules: (1) A training data set input module, which is used to input sample data and construct a training data set. Each sample data in the training data set contains five input features of the detection values of GSDMD, p-MLKL, IL-6, PCT and CRP, as well as a binary classification diagnostic label of bacterial infection and non-bacterial infection of the sample; (2) a training module, which is used to train the training data set constructed by the training data set input module through a machine learning method, and establish a binary classification prediction model for bacterial infection and non-bacterial infection based on five input features of the detection values of GSDMD, p-MLKL, IL-6, PCT and CRP; (3) The prediction module predicts bacterial infection or non-bacterial infection based on the five input features of the test values of GSDMD, p-MLKL, IL-6, PCT and CRP of the input sample based on the prediction model obtained by the training module.
[0005] In a preferred embodiment, the machine learning method described in the training module is performed using a random forest classification model, wherein the random forest classification model is composed of a base learner consisting of 100 decision trees, each of which is randomly sampled from the training data set to generate a training data subset through bootstrap sampling, and at each node split, 2 features are randomly selected from all features for optimal splitting, the minimum number of node split samples is set to 2, and each decision tree is trained independently; and the weight of each category is automatically adjusted to balance the influence of each category in the model training process; binary diagnostic labels are predicted by majority voting; and, In the prediction module, the five input features of the detection values of GSDMD, p-MLKL, IL-6, PCT and CRP of the sample to be tested are input, and the binary diagnosis label of the sample to be tested is predicted using the majority voting method consistent with the training module.
[0006] In a more preferred embodiment, in the training module, each decision tree uses the Gini Index as the splitting criterion when splitting a node. The Gini Index of node t is defined as: , Among them, p k is the proportion of samples of the kth category in node t, K is the total number of categories, and each split selects the features and split points that minimize the weighted average Gini index.
[0007] In another more preferred embodiment, in the training module, the weight category adjustment is set to: , Among them, n samples is the total number of samples, K is the number of categories, n c is the number of samples in category c.
[0008] In yet another more preferred embodiment, in the training module, the predicted class is determined by majority voting: , in, is the predicted category, C is the set of all possible categories, I(.) is the indicator function, when y i =c takes 1, otherwise takes 0.
[0009] In another preferred embodiment, the machine learning method in the training module is performed using a multi-layer feedforward neural network, the training module including an input layer of 5 nodes containing detection values of GSDMD, p-MLKL, IL-6, PCT and CRP, 1 to 2 hidden layers and 1 output layer, the output layer is provided with 1 node, and the activation function is a Sigmoid function to achieve binary classification probability output; the multi-layer feedforward neural network uses a stochastic gradient descent (SGD) method for back propagation, the learning rate is set to 0.001, and L2 regularization (weight decay) is selected to prevent overfitting; the cross entropy loss function (Categorical Cross-Entropy) is calculated, and when the loss function is observed to be stable, the model converges, the training is stopped, and an independent test data set is established; and In the prediction module, five input features (GSDMD, p-MLKL, IL-6, PCT, and CRP) are fed into a neural network. After calculations at each layer, the output is a probability value (ranging from 0 to 1). A threshold (e.g., 0.5) is set. If the output probability is greater than or equal to the threshold, the sample is classified as positive (1, bacterial infection); otherwise, it is classified as negative (0, non-bacterial infection).
[0010] In another preferred embodiment, each hidden layer comprises 10 to 32 neurons, and the activation function adopts the Sigmoid function: , The output of the neuron is: , in, is the output of the i-th neuron in the previous layer, is the weight, is the bias and σ is the activation function.
[0011] In another more preferred embodiment, the cross entropy loss function (Categorical Cross-Entropy) is: , Among them, y true is the true diagnosis label: 0 or 1, y pred Predict probabilities for the model.
[0012] In another more preferred embodiment, the prediction module performs feature normalization processing on the sample to be predicted consistent with the training stage, inputs the sample into the neural network, and after calculation at each layer, outputs a probability value ranging from 0 to 1. A judgment threshold is set. If the output probability is greater than or equal to the threshold, it is judged as a positive class (1, bacterial infection); otherwise, it is judged as a negative class (0, non-bacterial infection).
[0013] More preferably, the determination threshold is 0.5.
[0014] This study addresses the clinical challenge of early diagnosis of bloodstream infections by providing a five-biomarker panel. These include two novel markers: gasdermin D (GSDMD) and phosphorylated mixed-member kinase domain-like protein (p-MLKL), as well as three conventional markers: procalcitonin (PCT), CRP, and IL-6. GSDMD and p-MLKL are key proteins activated in the pyroptosis and necroptosis pathways of host cells following microbial infection, respectively. GSDMD is the core executor of pyroptosis. Bacterial infection activates caspase-4 / 5 / 11, which cleaves GSDMD and releases its N-terminal domain. The N-terminal fragments then oligomerize and form pores in the cell membrane, leading to ion imbalance, cell swelling and rupture, and the release of pro-inflammatory cytokines such as IL-1β and IL-18. These pores ultimately mediate cell membrane perforation, directly initiating pyroptosis. In the necroptosis pathway, bacteria stimulate death receptors (such as TNFR1) to activate MLKL through the RIPK1-RIPK3 complex, generating p-MLKL. This induces conformational changes and oligomerization, and p-MLKL oligomers translocate to the cell membrane, disrupting membrane integrity, leading to cell permeability necrosis, and the release of inflammatory factors. Pyroptosis and necroptosis play an important role in regulating the host cell's innate immune response to infection, but excessive activation can lead to inflammatory diseases such as sepsis and autoimmune diseases.
[0015] The technical solution of the present invention integrates five markers through a machine learning algorithm to establish a high-precision diagnostic model with an accuracy rate of >95%, which is significantly better than the current detection solution. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 Schematic diagram of the infectious pathogen type detection system based on five markers; Figure 2 Comparison of the receiver operating characteristic (ROC) curves of the five individual markers and the AUC values for each marker. ROC, receiver operating characteristic (ROC); AUC, area under the curve. Figure 3Comparison of receiver operating characteristic (ROC) curves and area under the curve (AUC) values for each marker across three marker combinations. The random forest model constructed with the five-marker combination had the highest AUC value for diagnosing bloodstream bacterial infection. (A) Characteristics of the model constructed with the combination of PCT, CRP, and IL-6; (B) Characteristics of the model constructed with the combination of GSDMD and p-MLK; (C) Characteristics of the model constructed with the five-marker combination. 1, Model error tree; 2, Rank order of variable importance; 3, ROC and AUC of the model. ROC, receiver operating characteristic (ROC); AUC, area under the curve. DETAILED DESCRIPTION
[0017] The present invention will be further described below with reference to specific embodiments, and the advantages and features of the present invention will become clearer as the description proceeds. However, these embodiments are merely exemplary and do not constitute any limitation on the scope of protection defined by the claims of the present invention.
[0018] Figure 1 A workflow diagram of the pathogen infection type detection system based on the random forest algorithm or the neural network algorithm and five markers of the present invention is given, and its specific implementation plan is introduced below.
[0019] Example 1. Sample collection and detection
[0020] 1. Collect data from 100 cases of bloodstream bacterial infection (including Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Enterococcus faecalis, etc.), 150 cases of non-bacterial infection, including 100 cases of simple viral infection (influenza A virus, herpes simplex virus, Epstein-Barr virus, etc.), and 50 cases without infection.
[0021] 2. Collect whole blood or plasma from patients with fever and use conventional ELISA to measure the levels of GSDMD (gasdermin D), p-MLKL (phosphorylated mixed series kinase domain-like protein), IL-6 (interleukin-6), PCT (procalcitonin), and CRP (acute phase reactant). The specific method is as follows: (1) Human GSDMD, p-MLKL, PCT, CRP, and IL-6 ELISA kits (CoAib Biotechnology Co., Ltd., Shanghai, China) were used; (2) Remove the kit from 4°C and equilibrate it at room temperature for 30 minutes (avoid direct exposure to light); (3) Dilute the concentrated washing solution with double distilled water to the working concentration (e.g., 20× concentrated solution needs to be diluted 1:19); (4) Plasma: After collection, centrifuge (2000 × g, 20 minutes), remove the precipitate, aliquot and store at -80°C; whole blood: 300 μL was collected and the total protein was extracted using a whole blood total protein extraction kit (Solar Biotechnology Co., Ltd., Beijing, China) and stored at -80°C; (5) Preparation of standard: Use standard diluent to dilute to different concentrations (e.g., 0, 50, 100, 200, 500, 1000 pg / mL) and establish a standard curve; (6) Incubation with coated antibodies: Place the 96-well plate pre-coated with anti-GSDMD antibodies at room temperature and tap gently to remove the protective solution; (7) Sample addition: Standard wells: add 100 μL of standard of different concentrations; Sample wells: add 100 μL of the sample to be tested (pre-dilution required); Blank wells: add 100 μL of sample diluent as a blank control; Negative / positive control wells: add according to the instructions; (8) Cover with film and incubate at 37°C for 90 minutes; (9) Wash, discard the liquid, add 300 μL of washing solution to each well, let it stand for 30 seconds and then discard, repeat 3 times (it is recommended to use a multi-channel pipette or automatic plate washer); (10) Add detection antibody (biotin-labeled), add 100 μL of biotinylated anti-human GSDMD / p-MLKL / PCT / CRP / IL-6 antibody working solution to each well; (11) Cover with film, incubate at 37°C for 60 minutes, and wash again three times; (12) Add streptavidin-HRP, add 100 μL enzyme conjugate (HRP-labeled streptavidin) to each well; (13) Cover with film and protect from light, incubate at 37°C for 30 minutes, and wash five times (to completely reduce nonspecific binding); (14) Color development (TMB): add 90 μL TMB substrate to each well and incubate at room temperature in the dark for 15-30 minutes (blue color will appear, and the time needs to be optimized to avoid too dark color); (15) Stop the reaction by adding 50 μL of stop solution (2M H2SO4) to each well. The color will change from blue to yellow. (16) Read the plate and read the absorbance (OD value) at 450 nm using a microplate reader within 30 minutes. Correct the difference between wells by referring to 630 nm. (17) Data analysis: The standard curve was drawn with the standard concentration as the abscissa (logarithmic coordinate) and the corresponding OD value as the ordinate, using four-parameter logistic (4PL) curve fitting; To calculate the sample concentration, substitute the sample OD value into the standard curve equation and multiply it by the dilution factor.
[0022] Example 2. Construction of random forest model Example 2 of the present invention uses a random forest classification model as the core prediction algorithm. The random forest model is composed of several decision tree base learners, and the specific structure is as follows: (1) Random Forest Model Structure: The random forest model contains N decision trees, denoted as {T1,T2,...,T N Each decision tree uses bootstrap sampling to randomly sample a training subset from the original training dataset. At each node split, a subset of features is randomly selected from all features for optimal splitting. Each decision tree is trained independently, and ultimately outputs a prediction using an ensemble voting method.
[0023] (2) Main technical parameter settings 1) The number of trees is set to 100, that is, the model contains 100 decision trees; 2) Maximum depth of the tree: set to 10; 3) Number of split features: set to the square root of the number of features, that is, each time a node splits, it randomly selects from all features features (M is the total number of features). In this embodiment, M=5, so about 2 features are randomly selected for each split; 4) Minimum number of node split samples: set to 2, that is, each internal node must contain at least 2 samples before it is allowed to continue splitting; 5) Category weight: Set to "balanced" to automatically adjust the weight of each category.
[0024] (3) Model training 1) Input features and processing A training dataset was constructed. Each sample in the training set contained five input features (GSDMD, p-MLKL, IL-6, PCT, and CRP) and a binary diagnostic label (bacterial infection vs. non-bacterial infection). The data was first cleaned. Outliers were detected using the interquartile range (IQR) method. Variable values exceeding Q3 + 1.5 × IQR were identified as outliers and removed. The training samples were then fed into each decision tree, and the classification results for each tree were obtained.
[0025] 2) Decision tree splitting criteria Each decision tree uses the Gini Index or Information Gain as the splitting criterion when splitting a node. Taking the Gini Index as an example, the Gini Index of node t is defined as: , Among them, p kis the proportion of samples of the kth class in node t, and K is the total number of classes (2 in this embodiment). Each split selects the feature and split point with the smallest weighted average Gini index.
[0026] 3) Weight category adjustment , Among them, n samples is the total number of samples, K is the number of categories (in this embodiment, the number of categories is 2), n c is the number of samples in category c (i.e., the number of samples in each category, e.g., n1=100, i.e., 100 cases in the bacterial infection group, n0=200, i.e., 200 cases in the non-bacterial infection group). This weight is used to balance the influence of each category during model training.
[0027] 4) Final prediction category determination The final prediction category is determined by majority voting as follows: , Where C is the set of all possible categories (including bacterial infection and non-bacterial infection in this embodiment), I(.) is the indicator function, when y i = c, the value is 1 (judged to be bacterial infection), otherwise it is 0 (judged to be non-bacterial infection).
[0028] (4) Model evaluation An independent test data set was established to verify the performance of the model on the test set, and the predictive performance of the model was comprehensively evaluated using specificity, sensitivity, positive predictive value (PPV), negative predictive value (NPV), and area under the ROC curve (AUC).
[0029] (5) Model prediction The prediction samples are processed in the same way as in the training phase. The model is input and the final prediction result is output using the majority voting method in the same way as in the training phase. i = c, the value is 1 (judged to be bacterial infection), otherwise it is 0 (judged to be non-bacterial infection).
[0030] Example 3. Construction of a multi-layer feedforward neural network model Embodiment 3 of the present invention further provides a prediction model based on a neural network (NN). The neural network model adopts a multi-layer feedforward neural network (MLP) structure, including an input layer, several hidden layers, and an output layer. It is described as follows: (1) Network structure and main parameter settings: 1) Input Layer: The number of nodes is equal to the number of input features. In this embodiment, there are 5 nodes, including GSDMD, p-MLKL, PCT, CRP, and IL-6.
[0031] 2) Hidden Layer: Set up 1 to 2 hidden layers, each containing 10 to 32 neurons. The activation function uses the Sigmoid function: , Where z is the input of the function, y is the output of the function, and e is the base of the natural logarithm.
[0032] The output of a neuron, taking the jth neuron in the lth layer as an example: , in, is the output of the i-th neuron in the previous layer, are weights (the weight parameters are randomly initialized and then updated through back-propagation learning), is the bias (the weight parameters are randomly initialized and then updated through back-propagation learning), and σ is the Sigmoid activation function.
[0033] 3) Output Layer: Set up one node, and the output layer activation function uses the Sigmoid function to achieve binary classification probability output.
[0034] (2) Model training: 1) Input Data Processing: A training dataset was constructed. Each sample in the training set contained five input features (GSDMD, p-MLKL, IL-6, PCT, and CRP) and a binary diagnostic label (bacterial infection vs. non-bacterial infection). The input features were normalized, and the diagnostic labels were encoded using a 0-1 scale.
[0035] 2) Loss Function: Categorical Cross-Entropy loss function: , Among them, y true is the true diagnostic label (0 or 1, 1 indicates bacterial infection, 0 indicates non-bacterial infection), y pred The predicted probability output by the model (the probability value is a floating point number between 0 and 1).
[0036] 3) Optimizer: Stochastic gradient descent (SGD) is used for back propagation. The SGD weight update formula is as follows: , Among them, θ is the parameter of the model, η is the learning rate, L i (θ) is the loss of the i-th sample, is the gradient operator. In this example, the learning rate is set to 0.001.
[0037] 4) Regularization method: L2 regularization (weight decay) is used to prevent overfitting. The loss function after L2 regularization is as follows: , Where L is the original loss function, θ is the model parameter, and λ is the regularization coefficient. In this example, the regularization coefficient λ can be set to 0.0001.
[0038] (3) Model convergence: When the loss function is observed to be stable, the model converges and training stops.
[0039] (4) Model evaluation An independent test data set was established to verify the performance of the model on the test set, and the predictive performance of the model was comprehensively evaluated using specificity, sensitivity, positive predictive value (PPV), negative predictive value (NPV), and area under the ROC curve (AUC).
[0040] (5) Model prediction The prediction samples are subjected to feature normalization consistent with the training phase. The samples are then fed into the neural network, and after calculations at each layer, the output is a probability value (ranging from 0 to 1). A threshold is set (e.g., 0.5). If the output probability is greater than or equal to the threshold, the sample is classified as positive (1, bacterial infection); otherwise, it is classified as negative (0, non-bacterial infection).
[0041] Example 4. Evaluation of the diagnostic ability of various marker combinations in the random forest model The present invention is characterized by a combination of five markers in plasma / whole blood for diagnosing bloodstream bacterial infection, whereas conventional markers for diagnosing bloodstream infection only use PCT, CRP, and IL-6.
[0042] 1. The key to the present invention is the combined use of five markers for diagnosing bloodstream bacterial infection, including GSDMD (gasdermin D), p-MLKL (phosphorylated mixed series kinase domain-like protein), IL-6 (interleukin-6), PCT (procalcitonin), and CRP (acute phase responder).
[0043] 2. Collect plasma from patients with fever and assay plasma / whole blood for levels of GSDMD (gasdermin D), p-MLKL (phosphorylated mixed series kinase domain-like protein), IL-6 (interleukin-6), PCT (procalcitonin), and CRP (acute phase reactant) using conventional ELISA.
[0044] 3. Model Development: A dataset consisting of plasma / whole blood GSDMD, p-MLKL, IL-6, PCT, and CRP values from patients with bacterial infections and non-bacterial infection controls (viral infections and no infections) was divided into a training set (210 cases) and a test set (90 cases). The training set accounted for 70% of the total dataset. A diagnostic model was constructed using a random forest algorithm and a neural network algorithm. The test set accounted for 30% of the total dataset and was used to evaluate and validate the established model.
[0045] 4. Diagnosis of new cases: Plasma was collected from patients with fever and assayed for GSDMD, p-MLKL, IL-6, PCT, and CRP levels using ELISA.
[0046] Random forest model diagnostic method: The five variable values of the sample to be predicted are input into the model, and the majority voting method consistent with the training phase is used to output the final prediction result. is the predicted category, C is the set of all possible categories (2 in this embodiment), when y i = c, the value is 1 (judged to be bacterial infection), otherwise it is 0 (judged to be non-bacterial infection).
[0047] Neural network model diagnosis method: The sample is input into the neural network, and after calculations at each layer, the output is a probability value (ranging from 0 to 1). If the output probability is greater than or equal to the threshold (0.5), it is judged as a positive class (1, bacterial infection); otherwise, it is judged as a negative class (0, non-bacterial infection).
[0048] 5. The inventors constructed a five-variable random forest model, using positive blood culture results and the detection of bacterial nucleic acid sequences in high-throughput sequencing screening for blood pathogens as the diagnostic criteria. The diagnostic results for new cases (test set) are shown in Table 1. The five-marker combination achieved a diagnostic accuracy of 0.909 (95% confidence interval: 0.757, 0.981), with sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) of 0.95, 0.85, 0.90, and 0.92, respectively. In contrast, the three-traditional marker combination achieved a diagnostic accuracy of only 0.758 (95% confidence interval: 0.577, 0.889), with sensitivity, specificity, positive predictive value (PPV), and negative predictive value (NPV) of 0.75, 0.77, 0.83, and 0.67, respectively.
[0049] Table 1. Systematic evaluation of the diagnostic power of various marker combinations in the random forest model. .
[0050] 95%CI, 95% confidence interval; PPV, positive predictive value; NPV, negative predictive value.
[0051] Figure 2 and Figure 3 The receiver operating characteristic (ROC) and area under the curve (ACU) results of various detection methods are also shown. Figure 2 The ROC curves of the five individual markers and the AUC value results of each marker were compared. Figure 3 Comparison of ROC curves for three marker combinations and the AUC values for each marker; (A) Characteristics of the model constructed by the combination of PCT, CRP, and IL-6; (B) Characteristics of the model constructed by the combination of GSDMD and p-MLK; (C) Characteristics of the model constructed by the combination of five markers; 1, Model error tree; 2, Rank of importance of each variable; 3, ROC and AUC of the model. Figure 2 and Figure 3 It can be seen that the present invention uses a combination of five markers to diagnose bloodstream bacterial infection, and its diagnostic AUC value is 0.788 ( Figure 2 ) increased to 0.962 ( Figure 3 C-3), with an accuracy of >95%, significantly improving the diagnostic accuracy of bloodstream bacterial infections. Figure 3 Among them, the random forest model constructed by the combination of five markers had the largest AUC value for diagnosing bloodstream bacterial infection (0.962), which was significantly higher than that of the PCT, CRP and IL-6 group (0, 765) and the GSDMD and p-MLK group (0.9).
[0052] The following examples illustrate that the five-marker combination detection method is superior to traditional single marker or traditional marker combination detection methods: Case 1. A patient with myelodysplastic syndrome, hospitalized for bone marrow transplantation, developed a sudden onset of high fever. Plasma PCT, CRP, and IL-6 levels were 0.4 ng / mL (mildly elevated), 1.8 mg / L (normal), and 10 pg / mL (mildly elevated), respectively. These three variables were entered into the model, suggesting a nonbacterial infection. However, plasma GSDMD and p-MLKL levels were 1593.907 ng / mL and 630.666 ng / mL, respectively. Simultaneously entering all five markers into the model suggested a bacterial infection. Whole-blood pathogen high-throughput sequencing revealed Pseudomonas aeruginosa infection, with 610 bacterial nucleic acid sequences.
[0053] Case 2: A multiple myeloma patient was admitted to the hospital with a fever of 38.8°C. Plasma PCT, CRP, and IL-6 levels were 0.36 ng / mL (mildly elevated), 3.1 mg / L (mildly elevated), and 14.4 pg / mL (elevated), respectively. These three variables were entered into the model, leading to a bacterial infection diagnosis. Plasma GSDMD and p-MLKL levels were 68.217 ng / mL and 16.407 ng / mL, respectively. The model also concluded that all five markers were non-bacterial. Whole-blood pathogen high-throughput sequencing confirmed the absence of pathogen nucleic acid sequences in the patient's blood.
[0054] Case 3: A patient with acute lymphoblastic leukemia maintained a low-grade fever (38.0-38.5°C) during hospitalization. Plasma PCT, CRP, and IL-6 levels were 0.09 ng / mL (mildly elevated), 1.3 mg / L (normal), and 4.6 pg / mL (normal), respectively. These three variables were input into the model, leading to a diagnosis of nonbacterial infection. However, plasma GSDMD and p-MLKL levels were 3188.466 ng / mL and 761.948 ng / mL, respectively. Simultaneously inputting all five markers into the model led to a diagnosis of bacterial infection. Subsequently, whole-blood pathogen high-throughput sequencing revealed Klebsiella pneumoniae infection, with 870 bacterial nucleic acid sequences.
Claims
1. A pathogen infection type detection system based on five markers, characterized in that: The five markers include: GSDMD, p-MLKL, IL-6, PCT, and CRP. The system includes the following modules: (1) A training data set input module, which is used to input sample data and construct a training data set. Each sample data in the training data set contains five input features of the detection values of GSDMD, p-MLKL, IL-6, PCT and CRP, as well as a binary classification diagnostic label of bacterial infection and non-bacterial infection of the sample; (2) a training module, which is used to train the training data set constructed by the training data set input module through a machine learning method, and establish a binary classification prediction model for bacterial infection and non-bacterial infection based on five input features of the detection values of GSDMD, p-MLKL, IL-6, PCT and CRP; (3) The prediction module predicts bacterial infection or non-bacterial infection based on the five input features of the test values of GSDMD, p-MLKL, IL-6, PCT and CRP of the input sample based on the prediction model obtained by the training module.
2. The pathogen infection type detection system based on five markers according to claim 1, characterized in that: The machine learning method described in the training module is performed using a random forest classification model. The random forest classification model is composed of 100 decision trees as a base learner. Each decision tree randomly samples from the training data set to generate a training data subset through a bootstrap sampling method. At each node split, two features are randomly selected from all features for optimal splitting. The minimum number of node split samples is set to 2. Each decision tree is trained independently; in addition, the weight of each category is automatically adjusted to balance the influence of each category in the model training process; a binary diagnosis label is predicted by majority voting; and, In the prediction module, the five input features of the detection values of GSDMD, p-MLKL, IL-6, PCT and CRP of the sample to be tested are input, and the binary diagnosis label of the sample to be tested is predicted using the majority voting method consistent with the training module.
3. The pathogen infection type detection system based on five markers according to claim 2, characterized in that: In the training module, each decision tree uses the Gini index as the splitting criterion when splitting a node. The Gini index of node t is defined as: , Among them, P k is the proportion of samples of the kth category in node t, K is the total number of categories, and each split selects the features and split points that minimize the weighted average Gini index.
4. The pathogen infection type detection system based on five markers according to claim 2, characterized in that: In the training module, the weight category adjustment is set as: , Among them, n samples is the total number of samples, K is the number of categories, n c is the number of samples in category c.
5. The pathogen infection type detection system based on five markers according to claim 2, characterized in that: In the training module, the predicted class is determined by majority voting: , in, is the predicted category, C is the set of all possible categories, I(.) is the indicator function, when y i =c takes 1, otherwise takes 0.
6. The pathogen infection type detection system based on five markers according to claim 1, characterized in that: The machine learning method described in the training module is performed using a multi-layer feedforward neural network. The training module includes an input layer with 5 nodes containing detection values of GSDMD, p-MLKL, IL-6, PCT and CRP, 1 to 2 hidden layers, and 1 output layer. The output layer is provided with 1 node, and the activation function is a sigmoid function to achieve binary classification probability output; the multi-layer feedforward neural network uses a stochastic gradient descent method for back propagation, the learning rate is set to 0.001, and L2 regularization is selected to prevent overfitting; the cross entropy loss function is calculated, and when the loss function is observed to be stable, the model converges, the training is stopped, and an independent test data set is established; and In the prediction module, five input features (GSDMD, p-MLKL, IL-6, PCT, and CRP) of the sample to be tested are input into the neural network. After calculations at each layer, the output is a probability between 0 and 1. A judgment threshold is set. If the output probability is greater than or equal to the threshold, it is judged as a positive class: 1, that is, bacterial infection; otherwise, it is judged as a negative class: 0, that is, non-bacterial infection.
7. The pathogen infection type detection system based on five markers according to claim 6, characterized in that: Each hidden layer contains 10 to 32 neurons, and the activation function uses the Sigmoid function: , The output of the neuron is: , in, is the output of the i-th neuron in the previous layer, is the weight, is the bias and σ is the activation function.
8. The pathogen infection type detection system based on five markers according to claim 6, characterized in that: The cross entropy loss function is: , Among them, y true is the true diagnosis label: 0 or 1, y pred Predict probabilities for the model.
9. The pathogen infection type detection system based on five markers according to claim 6, characterized in that: The prediction module performs feature normalization processing on the predicted samples consistent with the training stage, inputs the samples into the neural network, and outputs a probability value ranging from 0 to 1 after calculations at each layer. A judgment threshold is set. If the output probability is greater than or equal to the threshold, it is judged as a positive class: 1, that is, bacterial infection; otherwise, it is judged as a negative class: 0, that is, non-bacterial infection.
10. The pathogen infection type detection system based on five markers according to claim 9, characterized in that: The determination threshold is 0.5.
Citation Information
Patent Citations
Early detection of septicemia
CN101076806A
Early diagnosis of infections
CN109804245A
Marker and method for detecting inflammation-related diseases
CN115819548A
Optimal combination of early biomarkers for diagnosis of infection and sepsis in emergency department
CN116261663A
Marker application of blood GSDMD in early stage, differential diagnosis of blood flow infection and curative effect evaluation
CN117452000A