Rapid detection method for bloodstream infection pathogenic bacteria based on rapid mass spectrometry technology

By constructing a pathogen identification model using UVP-TOFMS and machine learning stacking algorithms, the problem of time-consuming bloodstream infection detection is solved, enabling rapid and accurate pathogen identification. This model is applicable to the identification of various common bloodstream infection pathogens and supports early clinical diagnosis and treatment.

CN121122425APending Publication Date: 2025-12-12SICHUAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202511210505.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-12-12

AI Technical Summary

Technical Problem

Existing methods for detecting bloodstream pathogens are time-consuming and difficult to use for rapid diagnosis, leading to delays in antibiotic treatment.

Method used

Volatile metabolites in bloodstream infection samples were analyzed using ultraviolet photoionization time-of-flight mass spectrometry (UVP-TOFMS). A pathogen identification model was constructed by combining the model with a machine learning stacking algorithm. The model's effectiveness was verified through training and testing sets, and characteristic metabolic biomarkers for each pathogen were screened out.

Benefits of technology

It enables rapid and accurate identification of pathogens, with a single sample testing time of only 20 seconds. It is applicable to the identification of a variety of common bloodstream infection pathogens, suitable for early clinical diagnosis and treatment, and provides a basis for precise antibiotic treatment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121122425A_ABST
    Figure CN121122425A_ABST
Patent Text Reader

Abstract

The invention discloses a rapid detection method for bloodstream infection pathogenic bacteria based on a rapid mass spectrometry technology, and belongs to the technical field of medical detection. Analyzing the metabolic gas of the blood flow infected pathogenic bacteria by using UVP-TOFMS to obtain a pathogenic bacteria metabolic profile spectrogram; preprocessing the obtained pathogenic bacterium metabolism profile spectrogram data; integrating the preprocessed pathogenic bacterium metabolism profile spectrogram data by using a stacking algorithm, and constructing a pathogenic bacterium identification model; evaluating the effectiveness of the pathogenic bacterium identification model on a training set and a test set through accuracy, precision, a confusion matrix and a subject operation curve; and comparing and screening characteristic metabolic markers of the pathogenic bacteria by combining the importance and significance of model characteristics, and indicating the types of the pathogenic bacteria. The method does not need sample pretreatment, is simple and convenient to operate, short in detection time (20 seconds per sample), high in accuracy and suitable for clinical popularization, and provides technical support for early diagnosis and treatment of bloodstream infection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of medical testing technology, specifically relating to a rapid detection method for bloodstream infectious pathogens based on rapid mass spectrometry technology. Background Technology

[0002] Bloodstream infection (BSI) is an infectious disease with high morbidity and mortality caused by pathogenic microorganisms (bacteria, fungi, and viruses, etc.) invading the bloodstream. It includes bacteremia, sepsis, and septicemia. Studies show that for sepsis patients, the mortality rate increases by 7.6% for every hour of delayed targeted antibiotic treatment; a 6-hour delay increases the mortality rate by 58%. Therefore, early and rapid diagnosis of the causative agent is crucial for patient treatment. Currently, the gold standard for detecting BSI pathogens is Blood Culture Identification (BCID), which involves enriching the infected blood sample in a blood culture bottle until approximately 10... 9 The process of obtaining a positive culture of pathogenic bacteria at CFU / mL takes an average of 12 hours; followed by purification culture to obtain single colonies, which typically takes 12–24 hours; then, pathogen identification is performed, followed by drug susceptibility testing based on the identification results (12–24 hours). It is evident that traditional blood culture testing is time-consuming and labor-intensive, making it difficult to achieve rapid detection of pathogenic bacterial infections, thus delaying precise antibiotic treatment. New rapid and sensitive methods for detecting pathogenic bacteria are of great significance for reducing the emergence of drug-resistant bacteria and lowering the mortality rate of pathogenic bacterial infections.

[0003] Currently, some mass spectrometry or molecular biology techniques can shorten the detection time of pathogens to some extent, but they still have their own shortcomings in clinical use. Mass spectrometry techniques, such as matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOFMS), can directly detect pathogenic peptides or protein molecules in centrifuged positive blood cultures for species identification. However, it is easily interfered with by matrix protein information, usually requiring bacterial purification culture to obtain single colonies for re-detection. Furthermore, MALDI cannot identify multiple infections of pathogens, and its application in identifying drug resistance is still in the research stage. Molecular biology techniques, such as polymerase chain reaction (PCR), utilize automated platforms to rapidly identify pathogens and their drug resistance genes in whole blood or positive blood cultures through processes such as microbial nucleic acid extraction, purification, and DNA amplification analysis. However, PCR technology is prone to false negative results due to primer specificity or PCR inhibition, and false positive results due to cell-free DNA.

[0004] In recent years, metabolomics analysis has gradually developed into a new technology for detecting pathogenic bacteria. Its principle is to use pathogen-specific metabolites as identification criteria. Pathogenic bacteria produce abundant volatile metabolites (BVMs) during their growth. Different pathogenic bacteria produce BVMs with significant differences in their nutrient synthesis and transformation pathways. Studying the composition and content differences of BVMs among different pathogenic bacteria allows for the identification of pathogenic species. For example, furan and 2-furan carbendazim can be detected in the feces of patients with Clostridium difficile diarrhea; high levels of isobutane, 2-butanone, and ethyl acetate are found in the exhaled breath of patients infected with Helicobacter pylori; and elevated levels of pyrrole and 2,3-butanedione can be detected in the blood of patients infected with Klebsiella pneumoniae. Traditional volatile compound detection techniques mainly involve gas chromatography-mass spectrometry (GC-MS). However, GC-MS analysis of trace components requires pre-enrichment, and the chromatographic separation is time-consuming and the analytical process is cumbersome, making it inconvenient for rapid clinical detection.

[0005] In summary, all existing BSI pathogen detection methods suffer from the problem of long detection times. Summary of the Invention

[0006] In order to overcome the shortcomings of the prior art, the present invention aims to provide a rapid detection method for bloodstream pathogens based on rapid mass spectrometry technology, so as to solve the technical problem of long detection time in existing BSI pathogen detection methods.

[0007] To achieve the above objectives, the present invention employs the following technical solution:

[0008] The first aspect of this invention discloses a rapid detection method for bloodstream infectious pathogens based on rapid mass spectrometry technology, comprising the following steps:

[0009] Step 1: Analyze the metabolic gases of bloodstream pathogens using UVP-TOFMS to obtain pathogen metabolic profile data;

[0010] Step 2: Preprocess the obtained pathogenic bacteria metabolic profile data, and randomly divide the preprocessed pathogenic bacteria metabolic profile data into training set and test set;

[0011] Step 3: Integrate the training set using a stacking algorithm to construct a pathogen identification model;

[0012] Step 4: Evaluate the effectiveness of the pathogen identification model on the test set using accuracy, precision, confusion matrix, and receiver operating characteristic curve;

[0013] Step 5: Based on the importance and significance of the pathogen identification model features, compare and screen the characteristic metabolic markers of each pathogen to indicate the type of pathogen.

[0014] Preferably, in step 1, the UVP-TOFMS settings are as follows: sample injection tube temperature is 100℃, ion source temperature is 70℃, test integration time is 20s, and the detection mass number range is a mass-to-charge ratio of 10-400.

[0015] Preferably, in step 2, the preprocessing involves signal correction, missing value filling, and standardization of the obtained pathogenic bacteria metabolic spectrum data.

[0016] More preferably, the signal correction is performed by correcting the signal fluctuation of UVP-TOFMS by measuring the change in the signal response intensity of UVP-TOFMS to benzene standard gas.

[0017] Preferably, missing value filling is performed using the KNN algorithm to calculate and fill missing values.

[0018] Preferably, the standardization process involves normalizing the area of ​​the signal peaks in the pathogenic bacteria metabolic spectrum data to the (0, 1) interval.

[0019] Preferably, in step 3, the stacking algorithm selects linear discriminant analysis, random forest, multilayer perceptron, and extreme gradient boosting tree as base classifiers, and logistic regression as a meta classifier; the classification result of each base classifier is used as an input variable to the meta classifier for further classification, and the output result of the meta classifier is used as the result of the stacking algorithm model.

[0020] Preferably, in step 4, the predictive ability of the pathogen identification model is quantified by the precision, recall, F1 score, and accuracy of the pathogen identification model in the test set prediction, and the predictive ability of the pathogen identification model for each type of pathogen is represented by the confusion matrix and the area under the receiver operating system curve.

[0021] Preferably, in step 5, the characteristic metabolic markers of the pathogenic bacteria include:

[0022] Escherichia coli: including m / z 117, m / z 59, m / z 45 and m / z 48; among them, m / z 117 corresponds to indole, m / z 59 corresponds to acetone, and m / z 45 corresponds to acetaldehyde;

[0023] Klebsiella pneumoniae: including m / z 73, m / z 147, m / z 89 and m / z 138; among them, m / z 73 corresponds to 2-butanone, m / z 147 corresponds to N,N-dimethylformamide, m / z 89 corresponds to ethyl acetate, and m / z 138 corresponds to decanal;

[0024] Bacteroides fragilis: including m / z 37 and m / z 55; where m / z 37 corresponds to ethanethiol and m / z 55 corresponds to butyraldehyde;

[0025] Acinetobacter baumannii: including m / z 51 and m / z 95; where m / z 95 corresponds to undecylaldehyde;

[0026] Pseudomonas aeruginosa: including m / z 170 and m / z 209; where m / z 170 corresponds to dodecane;

[0027] Staphylococcus aureus: including m / z 159, m / z 177 and m / z 197; among which, m / z 177 corresponds to 3-pentanol;

[0028] Enterococcus faecium: including m / z 68, m / z 171, m / z 45 and m / z 61; among them, m / z 68 corresponds to furan, m / z 45 corresponds to acetaldehyde, and m / z 61 corresponds to 2,3-butanedione.

[0029] A second aspect of the present invention discloses a rapid detection device for bloodstream infectious pathogens based on rapid mass spectrometry technology, comprising:

[0030] The pathogenic bacteria metabolic profile construction module is used to analyze the metabolic gases of bloodstream pathogenic bacteria using UVP-TOFMS to obtain the pathogenic bacteria metabolic profile.

[0031] The preprocessing module is used to preprocess the obtained pathogenic bacteria metabolic profile data and randomly divide the preprocessed pathogenic bacteria metabolic profile data into training set and test set;

[0032] The pathogen identification model building module is used to integrate the training set using a stacking algorithm to build a pathogen identification model.

[0033] The model performance evaluation module is used to evaluate the performance of the pathogen identification model on the training and test sets through accuracy, precision, confusion matrix, and receiver operating characteristic curve.

[0034] The detection module combines the importance and significance of pathogen identification model features to compare and screen characteristic metabolic markers of each pathogen to indicate the type of pathogen.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] This invention provides a rapid detection method for bloodstream pathogens based on UVP-TOFMS. It analyzes volatile metabolites produced by pathogens in bloodstream infection samples using ultraviolet photoionization time-of-flight mass spectrometry (UVP-TOFMS); combines a machine learning stacking algorithm to construct a pathogen identification model using a training set; verifies the effectiveness of the pathogen identification model using a test set; and finally, identifies characteristic markers for each pathogen by comparing the differences in the content of volatile metabolites from different pathogens. Experiments have demonstrated that this method has the following advantages: 1) The constructed pathogen identification model achieved an average accuracy ≥0.80 on the training set after five-fold cross-validation, and a comprehensive accuracy ≥0.85 on the test set. The area under the ROC curve for all categories was 0.93 or higher, indicating that the pathogen identification model has good predictive performance; 2) Rapid detection: Metabolite analysis using UVP-TOFMS requires only 20 seconds per sample, significantly shortening the diagnostic time; 3) Simple operation: No sample pretreatment is required, and headspace gas in blood culture bottles can be analyzed directly; 4) High accuracy: Combined with machine learning algorithms, the classification accuracy is high (94% under anaerobic conditions and 86% under aerobic conditions); 5) Wide clinical application: Applicable to the rapid identification of various common bloodstream infection pathogens (including Escherichia coli, Klebsiella pneumoniae, Acinetobacter baumannii, Pseudomonas aeruginosa, Staphylococcus aureus, Enterococcus urealyticum, and Bacteroides fragilis), suitable for clinical promotion, providing technical support for the early diagnosis and treatment of bloodstream infections, and providing a basis for precision antibiotic treatment. Attached Figure Description

[0037] Figure 1 This is a schematic diagram of the detection process for bloodstream pathogens based on UVP-TOFMS according to the present invention;

[0038] Figure 2 This is a schematic diagram of the multi-gas path injection system of the present invention; wherein, 1-constant temperature water bath device; 2-blood culture bottle; 3-double-hole rubber stopper; 4-inlet needle valve; 5-outlet needle valve; 6-zero-level air bottle; 7-gas path valve;

[0039] Figure 3 The following is a graph showing the evaluation results of the Gram bacteria classification test set samples in Example 1 of the present invention; wherein, (a) is a confusion matrix prediction graph of the Gram bacteria classification test set samples in Example 1, and (b) is the ROC curve of the Gram bacteria classification test set samples in Example 1.

[0040] Figure 4The following is a graph showing the evaluation results of the pathogenic bacteria classification test set samples in Example 1 of the present invention; wherein, (a) is a confusion matrix prediction graph of the pathogenic bacteria classification test set samples in Example 1, and (b) is the ROC curve of the pathogenic bacteria classification test set samples in Example 1;

[0041] Figure 5 This is a graph showing the detection results of characteristic markers of various pathogens under aerobic conditions in Example 1 of the present invention;

[0042] Figure 6 The following is a graph showing the evaluation results of the Gram bacteria classification test set samples in Example 2 of the present invention; wherein, (a) is a confusion matrix prediction graph of the Gram bacteria classification test set samples in Example 2, and (b) is the ROC curve of the Gram bacteria classification test set samples in Example 2;

[0043] Figure 7 The following is a graph showing the evaluation results of the pathogenic bacteria classification test set samples in Example 2 of the present invention; wherein, (a) is a confusion matrix prediction graph of the pathogenic bacteria classification test set samples in Example 2, and (b) is the ROC curve of the pathogenic bacteria classification test set samples in Example 2;

[0044] Figure 8 These are characteristic markers of various pathogenic bacteria under anaerobic conditions in Example 2 of the present invention. Detailed Implementation

[0045] To enable those skilled in the art to understand the features and effects of the present invention, the following descriptions and definitions are only general descriptions of the terms and expressions mentioned in the specification and claims. Unless otherwise specified, all technical and scientific terms used herein have the ordinary meaning understood by those skilled in the art regarding the present invention, and in the event of any conflict, the definitions in this specification shall prevail.

[0046] The theories or mechanisms described and disclosed herein, whether right or wrong, should not in any way limit the scope of the invention, that is, the contents of the invention can be implemented without being limited by any particular theory or mechanism.

[0047] In this document, all features defined by numerical ranges or percentage ranges, such as numerical values, quantities, contents, and concentrations, are for the sake of brevity and convenience only. Accordingly, descriptions of numerical ranges or percentage ranges should be considered as covering and specifically disclosing all possible sub-ranges and individual numerical values ​​(including integers and fractions) within those ranges.

[0048] In this article, unless otherwise specified, “contains,” “includes,” “containing,” “has,” or similar terms cover the meanings of “composed of” and “mainly composed of,” for example, “A contains a” covers the meanings of “A contains a and others” and “A contains only a.”

[0049] For the sake of brevity, not all possible combinations of the technical features in each implementation scheme or embodiment are described herein. Therefore, as long as there is no contradiction in the combination of these technical features, the technical features in each implementation scheme or embodiment can be combined arbitrarily, and all possible combinations should be considered within the scope of this specification.

[0050] This article utilizes ultraviolet photoionization time-of-flight mass spectrometry (UVP-TOFMS), an online detection mass spectrometry technique. It generates 10.6 eV photons using a krypton lamp. The analyte molecules absorb the photon energy and are ionized. Photon ionization is a low-energy, soft ionization method, resulting in simple and easily interpretable spectra. Furthermore, common gases in the air (such as N2, O2, and CO2) have high ionization energies and will not be ionized, thus avoiding detection interference. UVP-TOFMS features high sensitivity (volatile substance detection limit <1 ppbv), high throughput (simultaneous analysis at 10-450 m / z), and rapid online detection. Its simple structure makes it easy to maintain and its low operating cost, making it suitable for rapid and convenient analytical detection in clinical settings.

[0051] In this article, *Bacteroides fragilis* (BF) is an obligate anaerobe, *Acinetobacter baumannii* (AB) and *Pseudomonas aeruginosa* (PA) are obligate aerobes, and the other bacteria (*Escherichia coli* (EC), *Klebsiella pneumoniae* (KP), *Staphylococcus aureus* (SA), and *Enterococcus faecalis* (EF)) are facultative anaerobes.

[0052] The multi-channel sample introduction system used in this paper includes a constant temperature water bath device 1, a zero-grade air bottle 6, and a pipeline control unit. The pipeline control unit consists of an inlet pipe and an outlet pipe, both equipped with gas valves 7. These gas valves 7 are used to control the switching of different sample channels, thereby enabling rapid detection of different samples. The constant temperature water bath device 1 is filled with 37°C water and contains several blood culture bottles 2. The blood culture bottles 2 are used to hold the samples to be tested. A carbon dioxide colorimetric sensor is installed at the bottom of the blood culture bottles 2 to indicate pathogenic bacteria in the samples. A double-hole rubber stopper 3 is detachably connected to the top of the blood culture bottles 2. An inlet needle valve 4 and an outlet needle valve 5 are respectively installed on the two holes of the double-hole rubber stopper 3. One end of the inlet needle valve 4 extends into the blood culture bottle 2, and the other end is connected to the zero-grade air bottle 6 containing high-purity zero-grade air through the inlet pipe. One end of the outlet needle valve 5 extends into the blood culture bottle 2, and the other end is connected to the UVP-TOFMS system through the outlet pipe. The working process of the multi-channel sampling system is as follows: A blood culture bottle 2 containing the sample to be tested is placed in a constant temperature water bath 1 and incubated for 10 minutes. One end of the inlet needle valve 4 is inserted into the double-hole rubber stopper 3 of the blood culture bottle 2, and the other end is connected to a zero-grade air bottle 6 via an inlet pipe. One end of the outlet needle valve 5 is inserted into the double-hole rubber stopper 3 of the blood culture bottle 2, and the other end is connected to the mass spectrometer injection tube of the UVP-TOFMS system via an outlet pipe. High-purity zero-grade air is then introduced through the inlet pipe and inlet needle valve 4 into the sample liquid below the surface of the blood culture bottle 2. Under atmospheric pressure difference, the headspace gas metabolized by pathogenic bacteria in the blood culture bottle 2 is sent out through the outlet needle valve 5 and outlet pipe to the mass spectrometer injection tube of the UVP-TOFMS system for detection by UVP-TOFMS. After each sample injection, zero-grade air is introduced to flush the injection pipeline.

[0053] In this paper, the stacking algorithm used to build the model is an ensemble learning algorithm that improves the overall prediction performance by integrating the prediction results of multiple classification algorithms, which can significantly improve model efficiency.

[0054] This paper employs the k-Nearest Neighbor (KNN) algorithm, a supervised learning method based on distance metrics. It finds the k nearest training samples in the training set to a given test sample and uses these k training samples to predict the test sample. This algorithm is simple, has strong generalization ability, and can estimate missing values ​​by calculating the average of the selected k training samples.

[0055] The Receiver Operating Characteristic Curve (ROC Curve) used in this paper is a core tool for evaluating the performance of binary classification models, and is particularly suitable for judging the balance between the sensitivity and specificity of a detection method.

[0056] The five-fold cross-validation method used in this paper is a core method for evaluating the generalization ability and stability of machine learning models.

[0057] The present invention provides a rapid detection method for bloodstream pathogens based on UVP-TOFMS, such as... Figure 1 As shown, it includes the following steps:

[0058] 1. Sample processing

[0059] Bloodstream infection samples are enriched in blood culture bottle 2 (aerobic or anaerobic culture). After blood culture bottle 2 reports a positive result (i.e., the carbon dioxide color sensor in blood culture bottle 2 changes color, indicating the presence of pathogens), blood culture bottle 2 is placed in a water bath for a certain period of time to maintain gas-liquid balance.

[0060] 2. Construction of metabolic profile of pathogenic bacteria

[0061] The headspace gas (i.e., pathogenic metabolic gas) in blood culture bottle 2 was introduced into the UVP-TOFMS system through a multi-channel sampling system for full-spectrum analysis, and the pathogenic metabolic spectrum data of each sample were obtained.

[0062] Furthermore, the UVP-TOFMS settings for each sample detection include: injection tube temperature of 100℃, ion source temperature of 70℃, test integration time of 20s, and the detection mass number range of mass-to-charge ratio of 10-400.

[0063] 3. Data Preprocessing

[0064] The obtained pathogenic bacterial metabolic profile data were subjected to signal correction, missing value imputation, and standardization.

[0065] Furthermore, signal correction refers to correcting instrument signal fluctuations by measuring changes in the signal response intensity of UVP-TOFMS to standard gases. Specifically, after measuring 20 samples, the signal intensity of 50 ppbv benzene standard gas is measured using UVP-TOFMS. Then, the signal peak intensity of each sample detected each day is divided by the standard gas signal intensity at the corresponding time to obtain the corrected detection data. Missing value imputation uses the KNN algorithm to calculate and fill missing values; standardization normalizes the signal peak areas in the pathogenic bacteria metabolic spectrum data to the (0, 1) interval.

[0066] 4. Construction of pathogen identification model

[0067] A stacking algorithm was used to model bloodstream infection samples from aerobic and anaerobic cultures, respectively.

[0068] 4.1 The processed pathogenic bacteria metabolic spectrum data (aerobic culture / anaerobic culture) matrix was randomly divided into a training set and a test set in a 7:3 ratio. The training set was used for model building and parameter optimization, and the test set was used to independently verify the model performance.

[0069] 4.2. The model is constructed using a stacked algorithm framework of "multi-base classifier + meta-classifier". The specific steps are as follows:

[0070] Base classifier selection: Four algorithms are fixed: Linear Discriminant Analysis (LDA), Random Forest (RF), Multi-layer Perceptron (MLP), and Extreme Gradient Boosting (XGBoost). The characteristics of different algorithms are used to achieve feature complementarity.

[0071] Meta-classifier selection: Logistic Regression (LR) is selected as the meta-classifier to integrate the output results of all base classifiers, reduce the limitations of a single algorithm, and improve the model's generalization ability;

[0072] 4.3. The classification results of each base classifier are used as input variables to the meta-classifier for further classification. The output of the meta-classifier is used as the result of the stacked algorithm model. The specific steps are as follows: Using the training set obtained in 4.1 as input, the base classifier training and parameter optimization are completed sequentially: The training set is input into the four base classifiers respectively. For the core hyperparameters of each base classifier, five-fold cross-validation is used for iterative optimization. The training set is randomly divided into five mutually exclusive subsets. Each time, four subsets are used for training and one subset is used for validation. After five iterations, the parameter combination with the best average performance is taken as the optimal parameters of the base classifier. The classification results of the four base classifiers on the training set are collected and used as "meta-features" to be input into the meta-classifier. The hyperparameters of the meta-classifier are optimized through five-fold cross-validation to obtain the pathogen identification model.

[0073] 5. Model Performance Evaluation

[0074] Input the test set obtained in 4.1 into the pathogen identification model obtained in 4.3, and obtain the pathogen classification results output by the model. Compare the results with the clinical real classification results of the samples. Quantify the predictive ability of the pathogen identification model by the precision, recall, F1 score, and accuracy of the model in the test set prediction. The predictive ability of the pathogen identification model for each type of pathogen is represented by the confusion matrix and the area under the curve (AUC), and the effectiveness of the pathogen identification model is evaluated.

[0075] 6. Rapid detection of bloodstream pathogens

[0076] By combining the importance and significance of pathogen identification model features, characteristic metabolic markers of each pathogen are compared and screened to indicate the type of pathogen.

[0077] The present invention is further illustrated below with specific examples based on clinical data. It should be understood that these examples are for illustrative purposes only and are not intended to limit the scope of the invention. Furthermore, it should be understood that after reading this description, those skilled in the art can make various alterations or modifications to the invention, and these equivalent forms also fall within the scope defined by the appended claims.

[0078] The following examples use instruments and equipment conventional in the art. Experimental methods in the following examples, unless otherwise specified, are generally performed under standard conditions or as recommended by the manufacturer. All raw materials used in the following examples are conventional commercially available products with specifications in the art, unless otherwise stated.

[0079] Example 1

[0080] A rapid detection method for bloodstream pathogens based on UVP-TOFMS was established for the identification of pathogens in aerobic culture samples, including the following steps:

[0081] 1. Sample collection

[0082] We collected 176 Gram-positive bacterial infection samples, 393 Gram-negative bacterial infection samples, and 92 negative samples from clinical settings. Specifically, these included Escherichia coli (SA, 134 cases), Klebsiella pneumoniae (KP, 102 cases), Acinetobacter baumannii (AB, 114 cases), Pseudomonas aeruginosa (PA, 43 cases), Staphylococcus aureus (SA, 95 cases), and Enterococcus faecalis (EF, 81 cases).

[0083] 2. Sample processing

[0084] Clinical bloodstream infection samples collected in step 1 were cultured in BacT / ALERT FN / FA Plus blood culture bottles (purchased from Bio Mérieux) at 37°C until the carbon dioxide sensor at the bottom of the BacT / ALERT FN / FA Plus blood culture bottle changed color, indicating the presence of pathogens in the sample, i.e., a clinical positive report. Simultaneously, samples that did not report a positive result after 7 days of culture in BacT / ALERT FN / FA Plus blood culture bottles were considered negative blood culture bottles and used as a control study.

[0085] The aerobic blood culture bottles that reported positive results and the 92 aerobic negative blood culture bottles were placed in a 37°C water bath for 10 minutes to allow the gas-liquid state in the BacT / ALERT FN / FA Plus blood culture bottles to reach equilibrium, and then the samples were injected through a multi-channel injection system.

[0086] 3. Mass spectrometry detection

[0087] like Figure 2 As shown, inlet needle valve 4 and outlet needle valve 5 are simultaneously inserted into the BacT / ALERT FN / FA Plus blood culture bottle. Under atmospheric pressure difference, the headspace gas inside the bottle is delivered to the sample injection pipeline of the UVP-TOFMS system by zero-level air. The UVP-TOF MS instrument parameters are set as follows: ion source pressure 500 Pa, ion source temperature 70℃, sample tube temperature 100℃, integration time 20 s. The mass number detection range is 10-400, and the mass number is detected by the detector after ionization by the UV lamp and separation by the time-of-flight mass analyzer. After each sample injection, zero-level air is introduced to flush the pipeline until the spectral background is clean.

[0088] 4. Quality Control

[0089] A standard gas analysis was performed after every 20 samples. The standard gas consisted of benzene (m / z 78). The change in UV lamp sensitivity and instrument status were assessed by observing the standard gas spectrum. When observing the standard gas spectrum, if the intensity of the detected standard substance was found to be below 1 × 10⁻⁶, [further analysis was needed]. 5 If the mass-to-charge ratio of the detected standard substance differs from the theoretical value by more than 0.1 Da, it is necessary to check whether the mass axis has drifted and to perform mass axis correction in a timely manner.

[0090] 5. Data Preprocessing

[0091] The obtained pathogenic bacterial metabolic spectrum data underwent signal correction, missing value imputation, and standardization. The specific steps were as follows: First, the signal fluctuation of the instrument was corrected by measuring the signal response intensity changes of the UVP-TOFMS to the standard gas. Specifically, after measuring 20 samples, the signal intensity of 50 ppbv of benzene standard gas was measured using UVP-TOFMS. Then, the signal peak intensity of each sample detected each day was divided by the standard gas signal intensity at the corresponding time to obtain the corrected detection data. Second, the K-nearest neighbor algorithm (k=10) was used to impute missing values ​​in the original data matrix. Finally, the data matrix was scaled using standardization to make it conform to a distribution with a mean of 0 and a standard deviation of 1, thereby eliminating the influence of differences in magnitude between different features on model building.

[0092] 6. Constructing an aerobic culture model for identifying pathogenic bacteria

[0093] 6.1 Data Processing of Aerobic Culture Samples

[0094] Take the pathogenic bacterial metabolic spectrum data matrix of the aerobic culture bloodstream infection samples obtained in step 5. For the Gram tri-classification (Gram positive bacteria, Gram negative bacteria and negative samples) scenario of pathogenic bacteria, in order to balance the sample size of each category and avoid model bias, 92 cases were randomly selected from the Gram positive bacteria sample library and 92 cases were randomly selected from the Gram negative bacteria sample library to form an aerobic culture modeling sample set (a total of 276 cases) with 92 negative samples.

[0095] 6.2 Splitting the Training and Test Sets in the Dataset

[0096] Of the 276 aerobic culture pretreatment samples in section 6.1, 193 samples were used for training set learning, and 83 samples were used for test set evaluation.

[0097] 6.3 Constructing the Overlay Algorithm Framework

[0098] Base classifiers: Four algorithms were selected: Linear Discriminant Analysis (LDA), Random Forest (RF), Multilayer Perceptron (MLP), and Extreme Gradient Boosting Tree (XGBoost);

[0099] Meta-classifier: Logistic Regression (LR) is selected to integrate the output results of the base classifiers and output the final identification result.

[0100] 6.4 Construction and Parameter Optimization of the Aerobic Culture Model

[0101] 6.4.1 Training of Base Classifiers

[0102] The aerobic training set (193 cases) was input into four base classifiers, and independent training was performed for each base classifier: for LDA: the regularization parameter was optimized; for RF: the number of decision trees (n_estimators) and the maximum tree depth (max_depth) were optimized; for MLP: the number of hidden layer nodes, the learning rate, and the number of iterations were optimized; for XGBoost: the learning rate (learning_rate), the number of trees, and the maximum depth were optimized.

[0103] 6.4.2 Five-fold cross-validation parameter optimization: For each base classifier, five-fold cross-validation is used to iteratively optimize the hyperparameters. The aerobic training set is randomly divided into 5 mutually exclusive subsets. Each time, 4 subsets are used for training and 1 subset is used for validation. After 5 iterations, the parameter combination with the highest average accuracy is taken as the optimal parameters of the base classifier.

[0104] 6.4.3 Meta-classifier training: The classification results of the four base classifiers on the aerobic culture training set (the class probability of each sample to the output of the four base classifiers) are used as "meta-features" and input into the meta-classifier (LR) for secondary training. The regularization parameters of LR are optimized by five-fold cross-validation to obtain the aerobic culture pathogen identification model.

[0105] 7. Model Performance Evaluation

[0106] Eighty-three aerobic culture test set samples were input into an aerobic culture pathogen identification model. The Gram classification results of the pathogens output by the model were obtained and compared with the clinical blood culture identification results of the samples. The fitting effect of the aerobic culture pathogen identification model was evaluated by the average accuracy of the training set samples after five-fold cross-validation. The model's generalization ability was evaluated by the predictive accuracy of the model on the test set, with evaluation metrics including precision, recall, F1 score, and overall accuracy. Precision represents the proportion of samples predicted as belonging to the correct category; recall represents the proportion of samples with a true result correctly predicted by the model; the F1 score is the harmonic value of precision and recall; all three values ​​are better when closer to 1; and overall accuracy measures the overall predictive performance of the model, representing the proportion of correctly predicted samples out of the total number of samples. The formulas for calculating precision, recall, F1 score, and overall accuracy are shown in (1), (2), (3), and (4), respectively. TP (True Positives) represents the number of samples correctly predicted as positive by the model; TN (True Negatives) represents the number of samples correctly predicted as negative by the model; FP (False Positives) represents the number of samples incorrectly predicted as positive by the model; and FN (False Negatives) represents the number of samples incorrectly predicted as negative by the model.

[0107]

[0108] The classification prediction performance of the model was evaluated by plotting the confusion matrix, ROC curve, and calculating the area under the curve. The ROC curve was constructed using a one-to-many, micro-average, and macro-average strategy. The one-to-many strategy converts the multi-class problem into a binary classification problem for prediction, namely: (1) the target class is regarded as the positive class and all other classes are regarded as the negative class, and a threshold for distinguishing between positive and negative classes is set; (2) the True Positive Rate (TPR) and False Positive Rate (FPR) of each class at different thresholds are calculated, and then the ROC curve is plotted with FPR as the horizontal axis and TPR as the vertical axis. Micro-average is to combine the true and false positives of all classes and calculate the overall TPR and FPR to plot the ROC curve; Macro-average is to calculate the TPR and FPR of each class separately and then take the average to plot the ROC curve. Macro-average gives the same weight to the number of samples of each class in the model during analysis.

[0109] After multiple random training iterations, the classification results, confusion matrix, and ROC curves are shown in Table 1 and Table 2, respectively. Figure 3 As shown, the model achieved an average accuracy of 0.88 on the training set after five-fold cross-validation, and an overall accuracy of 0.93 on the test set. The area under the ROC curve for all classes was 0.96 or higher, indicating that the model has excellent predictive performance.

[0110] Furthermore, multi-class classification modeling analysis was performed on positive samples of the six pathogenic bacteria. A total of 569 samples were included in the model fitting, of which 398 samples were used for training and 171 samples were used for testing. The classification results of the six pathogenic bacteria are shown in Table 1. Figure 4 As shown, its average accuracy after five-fold cross-validation on the training set is 0.80, and its classification prediction accuracy on the test set is 0.85. The area under the ROC curve is above 0.93, indicating that the model's classification results are reliable.

[0111] Table 1. Results of pathogen classification using the stacking model (aerobic).

[0112]

[0113] Furthermore, through significance testing and box plot analysis, characteristic metabolic markers of various pathogens under aerobic culture conditions were obtained by comparison, such as... Figure 5As shown, under aerobic culture conditions, a significant increase in mass-to-charge ratio (M / C ratio) of 48 or 59 (acetone) suggests the possible presence of *Escherichia coli*; a significant increase in M / C ratio of 73 (2-butanone) or 147 (N,N-dimethylformamide) suggests the possible presence of *Klebsiella pneumoniae*; a significant increase in M / C ratio of 51 or 95 (undecaldehyde) suggests the possible presence of *Acinetobacter baumannii*; a significant increase in M / C ratio of 170 (dodecane) or 209 suggests the possible presence of *Pseudomonas aeruginosa*; a significant increase in M / C ratio of 159 or 177 (3-pentanol) suggests the possible presence of *Staphylococcus aureus*; and a significant increase in M / C ratio of 45 (acetaldehyde) or 61 (2,3-butanedione) suggests the possible presence of *Enterococcus faecalis*.

[0114] Example 2

[0115] A rapid detection method for bloodstream pathogens based on UVP-TOFMS was established for the identification of pathogens in anaerobic culture samples, including the following steps:

[0116] 1. Sample collection

[0117] We collected 196 Gram-positive bacterial infection samples, 274 Gram-negative bacterial infection samples, and 94 negative samples from clinical settings. Specifically, these included Escherichia coli (SA, 142 cases), Klebsiella pneumoniae (KP, 103 cases), Bacteroides fragilis (BF, 29 cases), Staphylococcus aureus (SA, 116 cases), and Enterococcus faecalis (EF, 80 cases).

[0118] 2. Sample processing

[0119] Clinical bloodstream infection samples collected in step 1 were cultured and identified using BacT / ALERT FN / FA Plus blood culture bottles (Bio Mérieux). If pathogenic bacteria were present in the bloodstream infection sample, the carbon dioxide sensor at the bottom of the BacT / ALERT FN / FA Plus blood culture bottle would change color after a certain period of incubation, indicating a positive clinical result. Samples that did not report a positive result after 5 days of incubation in BacT / ALERT FN / FA Plus blood culture bottles were considered negative blood culture bottles and used as a control.

[0120] The anaerobic blood culture bottles that reported positive results and the 94 anaerobic negative blood culture bottles were placed in a 37°C water bath for 10 minutes to allow the gas-liquid state in the BacT / ALERT FN / FAPlus blood culture bottles to reach equilibrium, and then the samples were injected through a multi-channel injection system.

[0121] 3. Mass spectrometry detection

[0122] The testing steps are the same as in Example 1.

[0123] 4. Quality Control

[0124] The quality control steps are the same as in Example 1.

[0125] 5. Data Preprocessing

[0126] The data preprocessing steps are the same as in Example 1.

[0127] 6. Construct a model for identifying pathogenic bacteria in anaerobic culture.

[0128] 6.1 Data Processing of Anaerobic Culture Samples

[0129] Take the pathogenic bacteria metabolic spectrum data matrix of the anaerobic culture bloodstream infection samples obtained in step 5. For the Gram tri-classification scenario of pathogenic bacteria (Gram positive bacteria, Gram negative bacteria and negative samples), in order to balance the sample size of each category and avoid model bias, 92 cases were randomly selected from the Gram positive bacteria sample library and 92 cases were randomly selected from the Gram negative bacteria sample library to form the anaerobic culture modeling sample set (a total of 276 cases) together with 92 negative samples.

[0130] 6.2 Splitting the Training and Test Sets in the Dataset

[0131] Of the 276 anaerobic culture pretreatment samples in section 6.1, 197 samples were used for training set learning, and 85 samples were used for test set evaluation.

[0132] 6.3 Constructing the Overlay Algorithm Framework

[0133] Base classifiers: Four algorithms were selected: Linear Discriminant Analysis (LDA), Random Forest (RF), Multilayer Perceptron (MLP), and Extreme Gradient Boosting Tree (XGBoost);

[0134] Meta-classifier: Logistic Regression (LR) is selected to integrate the output results of the base classifiers and output the final identification result.

[0135] 6.4 Construction and Parameter Optimization of Anaerobic Culture Model

[0136] 6.4.1 Training of Base Classifiers

[0137] The anaerobic culture training set (197 cases) was input into four base classifiers, and independent training was performed for each base classifier: for LDA: the regularization parameter was optimized; for RF: the number of decision trees (n_estimators) and the maximum tree depth (max_depth) were optimized; for MLP: the number of hidden layer nodes, the learning rate, and the number of iterations were optimized; for XGBoost: the learning rate (learning_rate), the number of trees, and the maximum depth were optimized.

[0138] 6.4.2 Five-fold cross-validation parameter optimization: For each base classifier, five-fold cross-validation is used to iteratively optimize the hyperparameters. The anaerobic culture training set is randomly divided into 5 mutually exclusive subsets. Each time, 4 subsets are used for training and 1 subset is used for validation. After 5 iterations, the parameter combination with the highest average accuracy is taken as the optimal parameters of the base classifier.

[0139] 6.4.3 Meta-classifier training: The classification results of the four base classifiers on the anaerobic culture training set (the probability of each sample to the class output of the four base classifiers) are used as "meta-features" and input into the meta-classifier (LR) for secondary training. The regularization parameters of LR are optimized by five-fold cross-validation to obtain the anaerobic culture pathogen identification model.

[0140] 7. Model Performance Evaluation

[0141] The steps are the same as in Example 1.

[0142] After multiple random training iterations, the classification results, confusion matrix, and ROC curves are shown in Table 2 and Table 3, respectively. Figure 6 As shown, the model achieved an average accuracy of 0.85 on the training set after five-fold cross-validation, and an overall accuracy of 0.95 on the test set. The area under the ROC curve for all classes was 0.97 or higher, indicating that the model has excellent predictive performance.

[0143] Furthermore, multi-class classification modeling analysis was performed on positive samples of the five pathogenic bacteria. A total of 470 samples were included in the model fitting, of which 328 samples were used for training and 142 samples were used for testing. The classification results of the five pathogenic bacteria are shown in Table 2. Figure 7 As shown, its average accuracy after five-fold cross-validation on the training set is 0.89, and its classification prediction accuracy on the test set is 0.90. The area under the ROC curve is above 0.97, indicating that the model's classification results are very good.

[0144] Table 2 shows the results of the stacked model for classifying pathogenic bacteria (anaerobic).

[0145]

[0146] Furthermore, through significance testing and box plot analysis, characteristic metabolic markers of various pathogens under anaerobic culture conditions were obtained by comparison, such as... Figure 8As shown, under anaerobic culture conditions, a significant increase in mass-to-charge ratio (M / C ratio) of 45 (acetaldehyde) and 117 suggests the possible presence of *Escherichia coli*; a significant increase in M / C ratio of 89 (ethyl acetate) and 138 (decanal) suggests the possible presence of *Klebsiella pneumoniae*; a significant increase in M / C ratio of 37 (ethanethiol) and 55 (butyraldehyde) suggests the possible presence of *Bacteroides fragilis*; a significant increase in M / C ratio of 177 and 197 suggests the possible presence of *Staphylococcus aureus*; and a significant increase in M / C ratio of 68 (furan) and 171 suggests the possible presence of *Enterococcus faecalis*.

[0147] The above content is only for illustrating the technical concept of the present invention and should not be construed as limiting the scope of protection of the present invention. Any modifications made to the technical solution based on the technical concept proposed in this invention shall fall within the scope of protection of the claims of this invention.

Claims

1. A rapid detection method of pathogenic bacteria of blood stream infection based on a rapid mass spectrometry technique, characterized in that, The method comprises the following steps: Step 1: analyzing blood stream infection pathogen metabolic gas by UVP-TOFMS to obtain pathogen metabolic profile data; Step 2: preprocessing the obtained pathogen metabolic profile data, and randomly dividing the preprocessed pathogen metabolic profile data into a training set and a test set; Step 3: integrating the training set by using a stacking algorithm to construct a pathogen identification model; Step 4: evaluating the performance of the pathogen identification model on the test set by accuracy, precision, confusion matrix and receiver operating curve; Step 5: combining the feature importance and significance of the pathogen identification model, comparing and screening the characteristic metabolic markers of each pathogen to indicate the pathogen species.

2. The method for rapid detection of pathogenic bacteria of bloodstream infection based on fast mass spectrometry technology according to claim 1, characterized in that, In step 1, the setting parameters of UVP-TOFMS are as follows: the temperature of the sample inlet tube is 100 DEG C, the temperature of the ion source is 70 DEG C, the test integration time is 20 s, and the detection mass number range is mass-to-charge ratio 10-400.

3. The method according to claim 1, wherein the method is characterized by: In step 2, the preprocessing is signal correction, missing value filling and standardization processing of the obtained pathogen metabolic profile data.

4. The method according to claim 3, wherein the method is characterized by, The signal correction is to correct the signal fluctuation of UVP-TOFMS by the signal response intensity change of benzene standard gas of UVP-TOFMS.

5. The method according to claim 3, wherein the method is characterized by, The missing value filling adopts KNN algorithm to calculate the missing values for filling.

6. The method according to claim 3, wherein the method is characterized by, The standardization processing is to normalize the signal peak area in the pathogen metabolic profile data to the interval (0, 1).

7. The method according to claim 1, wherein the method is characterized by: In step 3, the stacking algorithm selects linear discriminant analysis, random forest, multilayer perception and extreme gradient boosting tree as the base classifier, and uses logistic regression as the meta-classifier; the classification results of each base classifier on the data are input into the meta-classifier for further classification, and the output results of the meta-classifier are the results of the stacking algorithm model.

8. The method for rapid detection of pathogenic bacteria of bloodstream infection based on fast mass spectrometry technology according to claim 1, characterized in that, In step 4, the prediction ability of the pathogen identification model is quantified by the precision, recall, F1 score and accuracy of the pathogen identification model in the test set prediction, and the prediction ability of the pathogen identification model for each type of pathogen is represented by the confusion matrix and the area under the receiver operating curve.

9. The method according to claim 1, wherein the method is characterized by, In step 5, the characteristic metabolic markers of the pathogen include: Escherichia coli includes: m / z 117, m / z 59, m / z 45 and m / z 48; wherein m / z 117 corresponds to indole, m / z 59 corresponds to acetone, and m / z 45 corresponds to acetaldehyde; Klebsiella pneumoniae includes: m / z 73, m / z 147, m / z 89 and m / z 138; wherein m / z 73 corresponds to 2-butanone, m / z 147 corresponds to N,N-dimethylformamide, m / z 89 corresponds to ethyl acetate, and m / z 138 corresponds to decanal; Bacteroides fragilis includes: m / z 37 and m / z 55; wherein m / z 37 corresponds to ethanethiol, and m / z 55 corresponds to butyraldehyde; Acinetobacter baumannii includes: m / z 51 and m / z 95; wherein m / z 95 corresponds to undecanal; Pseudomonas aeruginosa includes: m / z 170 and m / z 209; wherein m / z 170 corresponds to dodecane. Staphylococcus aureus includes m / z 159, m / z 177 and m / z 197; among them, m / z 177 corresponds to 3-pentanol; Enterococci include m / z 68, m / z 171, m / z 45 and m / z 61; among them, m / z 68 corresponds to furan, m / z 45 corresponds to acetaldehyde, and m / z 61 corresponds to 2,3-butanedione.

10. A rapid detection device for pathogenic bacteria of blood stream infection based on a rapid mass spectrometry technique, characterized in that, include: The pathogenic bacteria metabolic profile construction module is used to analyze the metabolic gases of bloodstream pathogenic bacteria using UVP-TOFMS to obtain the pathogenic bacteria metabolic profile. The preprocessing module is used to preprocess the obtained pathogenic bacteria metabolic profile data and randomly divide the preprocessed pathogenic bacteria metabolic profile data into training set and test set; The pathogen identification model building module is used to integrate the training set using a stacking algorithm to build a pathogen identification model. The model performance evaluation module is used to evaluate the performance of the pathogen identification model on the training and test sets through accuracy, precision, confusion matrix, and receiver operating characteristic curve. The detection module combines the importance and significance of pathogen identification model features to compare and screen characteristic metabolic markers of each pathogen to indicate the type of pathogen.

Citation Information

Cited By

  • Noninvasive detection method and system for yin and yang deficiency syndrome based on EESI-MS

    CN122238463A