A rapid, AI-assisted identification method for common childhood pathogens and their drug resistance based on Raman spectroscopy.
By combining Raman spectroscopy with deep learning algorithms, a deep learning classification model is established to quickly and accurately identify common childhood pathogens and their drug resistance. This solves the identification difficulties in existing technologies, provides an efficient diagnostic tool, and reduces the risk of treatment failure and the spread of drug resistance.
Patent Information
- Application Number
- CN202510764934.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-10-28
- Estimated Expiration
- 2045-06-10
AI Technical Summary
Current technologies make it difficult to quickly and accurately identify common childhood pathogens and their drug resistance, leading to difficulties in clinical treatment and the spread of drug resistance.
A deep learning classification model was established by combining Raman spectroscopy with deep learning algorithms. Spectral data of common pathogens in children were collected using a laser confocal Raman spectrometer, and data processing and classification were performed using a 1D convolutional neural network to achieve rapid identification.
It enables rapid and accurate identification of bacteria and fungi, Gram-positive and Gram-negative bacteria, common pathogenic bacteria in children, and carbapenem-resistant and carbapenem-sensitive Acinetobacter baumannii, providing an efficient and convenient diagnostic tool and reducing the risk of treatment failure and the spread of drug resistance.
Smart Images

Figure CN120293947B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the field of artificial intelligence technology, specifically relating to an artificial intelligence-assisted rapid identification method for common pathogenic bacteria in children and their drug resistance based on Raman spectroscopy. Background Technology
[0002] Different pathogens can present with many similarities in clinical manifestations, making them difficult to distinguish. However, the appropriate treatment varies depending on the pathogen. Bacterial infections require antibiotics, while fungal infections require antifungal drugs. Early identification of whether an infection is bacterial or fungal helps clinicians develop targeted treatment plans and provides a reference for selecting effective antibiotics.
[0003] Bacteria are classified into Gram-positive and Gram-negative bacteria based on their Gram staining characteristics. Their pathogenic mechanisms differ: Gram-positive bacteria primarily rely on exotoxins and invasive enzymes, while Gram-negative bacteria mainly rely on endotoxins. Furthermore, there are differences in their susceptibility to antimicrobial agents. Gram-positive bacteria are more sensitive to penicillins and vancomycin, while Gram-negative bacteria typically require a combination of third-generation cephalosporins, β-lactamase inhibitors, or carbapenems. Once a bacterial infection is confirmed, early identification of Gram-positive or Gram-negative bacteria can provide a basis for clinical selection of antimicrobial agents.
[0004] Furthermore, different bacterial species exhibit significant differences in their sensitivity to antimicrobial drugs; identifying the specific species can prevent indiscriminate medication. Determining the exact species of common pathogens allows for precise treatment, reduces the risk of treatment failure, and improves patient prognosis.
[0005] Antimicrobial resistance caused by the overuse and abuse of antibiotics has become a global concern, especially carbapenem-resistant Acinetobacter baumannii (CRAB). CRAB is resistant to most β-lactam antibiotics, requiring second-line treatment with polymyxins and tigecycline, but these have limited efficacy and high toxicity. Furthermore, because these strains can reside in the respiratory tract and other sites for months, they can spread in hospitals, ultimately causing nosocomial infections and posing a significant challenge to clinical treatment. Early identification of CRAB can prevent the clinical use of ineffective carbapenems, reduce the risk of resistance gene spread, and curb the spread of resistance. CRAB is easily transmitted through contact, leading to nosocomial outbreaks; early identification of its resistance helps in implementing isolation and control measures as soon as possible to break the chain of transmission.
[0006] These classifications are directly related to precision clinical treatment, drug resistance control, and patient prognosis. Rapid identification of pathogen types and drug resistance characteristics allows for the development of individualized treatment plans, thereby improving efficacy, shortening disease duration, and curbing drug resistance. Summary of the Invention
[0007] To overcome the shortcomings of existing pathogen identification technologies, this invention provides an artificial intelligence-assisted rapid identification method for common childhood pathogens and their drug resistance based on Raman spectroscopy.
[0008] By combining in-situ Raman spectroscopy with deep learning algorithms, the problems of cumbersome identification process, long time consumption, and low positive rate of pathogens can be solved.
[0009] This invention utilizes deep learning to train Raman spectra of common childhood pathogens, thereby establishing an identification model capable of accurately identifying these pathogens. A Raman spectral database is created by collecting Raman spectra of common childhood pathogens. The data is divided into three groups: training, testing, and validation. This data is then input into a proposed neural network model for processing. Deep learning is used to train the model on the training group data, selecting the best-performing model. The model is then validated using the validation group data, and the model parameters are adjusted. Finally, the identification effectiveness of the model is tested using the target bacteria.
[0010] A method for rapid identification of common childhood pathogens and their drug resistance based on Raman spectroscopy with artificial intelligence assistance, comprising the following steps:
[0011] (1) Collection of pathogenic bacteria samples: Twelve common pathogenic bacteria were isolated from clinical specimens. The twelve common pathogenic bacteria included bacteria and fungi. The bacteria were Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, and Staphylococcus aureus. The fungi were Candida albicans, Candida glabrata, and Saccharomyces cerevisiae. The twelve common pathogenic bacteria were identified by matrix-assisted laser desorption / ionization time-of-flight mass spectrometry.
[0012] (2) The strains identified as Acinetobacter baumannii were classified into carbapenem-resistant Acinetobacter baumannii and carbapenem-sensitive Acinetobacter baumannii according to their minimum inhibitory concentration of antimicrobial drugs.
[0013] (3) Raman spectral database establishment: The 12 common pathogens described in step (1) are detected using a laser confocal Raman spectrometer, and the resulting Raman spectral database is collected and saved;
[0014] (4) Establish a Raman spectroscopy deep learning classification model and pre-train weights on 30 categories of bacterial Raman datasets available online;
[0015] The design incorporates a 1D convolutional neural network architecture, consisting of an initial convolutional block, multi-scale residual blocks, mean pooling, and fully connected layers. The acquired Raman spectra, after self-normalization, are randomly divided into three sets: training, testing, and validation, which are then input into the neural network model. Gaussian random noise is added to augment the data during training. The ADAMW optimizer is used for model training, and the model is trained on a publicly available bacterial Raman dataset with 30 classes. The optimal learning rate, channels, and other hyperparameters are selected based on the model's performance on the validation set to determine the best-trained model. The trained model is then tested using the test set with accuracy, precision, recall, F1 score, and confusion matrix to evaluate its classification performance.
[0016] (5) By modifying the output channel of the last fully connected layer of the model, deep learning classification models for bacteria and fungi, Gram-positive and Gram-negative bacteria, the 12 common pathogenic bacteria, and carbapenem-resistant and carbapenem-sensitive Acinetobacter baumannii were established respectively. The pre-trained model was used as the initial weight, and the corresponding model was fine-tuned on the training set of the corresponding dataset. The classification effect of the model was tested using the test set with accuracy, precision, recall, F1 score and confusion matrix.
[0017] As a preferred embodiment of the method for rapid identification of common childhood pathogens and their drug resistance based on Raman spectroscopy using artificial intelligence, as described in this invention: In step (3), a laser confocal Raman spectrometer is selected, with a laser wavelength of 532 nm, a laser power of 0.1-2%, a spectrometer grating of 1200 g / mm, and a spectral range of 300-2000 cm⁻¹. -1 The acquisition time for a single spectrum is 30-60 seconds.
[0018] As a preferred embodiment of the method for rapid identification of common pathogenic bacteria and their drug resistance in children based on Raman spectroscopy described in this invention: In step (3), the pathogenic bacteria are suspended in sterile water with a turbidity of 0.5 McFarland degrees, dropped onto the surface of an aluminum sheet, and detected using a laser confocal Raman spectrometer.
[0019] As a preferred embodiment of the method for rapid identification of common pathogenic bacteria and their drug resistance in children based on Raman spectroscopy as described in this invention: in step (4), the proportion of each set of data is 70% for the training set, 10% for the validation set, and 20% for the test set.
[0020] As a preferred embodiment of the method for rapid identification of common pathogenic bacteria and their drug resistance in children based on Raman spectroscopy as described in this invention: in step (4), the mean of the Gaussian random noise is 0 and the variance is 0.1.
[0021] As a preferred embodiment of the method for rapid identification of common pathogenic bacteria and their drug resistance in children based on Raman spectroscopy as described in this invention: In step (4), a 1D convolutional neural network architecture is used, wherein each convolutional layer is followed by a 1D BatchNorm normalization layer and a ReLU activation function layer; the multi-scale residual blocks reduce the signal resolution through convolution with stride=2, and extract features at 6 scales respectively. The feature extraction at each scale is composed of 2 residual block network structures; the input data format for building the model is: number of channels, spectral length.
[0022] As a preferred embodiment of the method for rapid identification of common pathogenic bacteria and their drug resistance in children based on Raman spectroscopy according to the present invention: in step (5), the sample data to be tested is input into a pre-trained convolutional neural network model, and the model will output the confidence level of the sample to be tested as bacteria or fungi, and select the category with the highest confidence level to identify the sample as bacteria or fungi.
[0023] As a preferred embodiment of the method for rapid identification of common pathogenic bacteria and their drug resistance in children based on Raman spectroscopy as described in this invention: In step (5), the sample data to be tested is input into a pre-trained convolutional neural network model, and the model outputs the confidence level of the sample to be tested as Gram-positive bacteria or Gram-negative bacteria, and the category with the highest confidence level is selected to identify the sample as Gram-positive bacteria or Gram-negative bacteria.
[0024] As a preferred embodiment of the method for rapid identification of common pathogenic bacteria and their drug resistance in children based on Raman spectroscopy as described in this invention: In step (5), the sample data to be tested is input into a pre-trained convolutional neural network model. The model outputs the confidence level of the sample being identified as one of twelve strains: Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, Staphylococcus aureus, Candida albicans, Candida glabrata, and Saccharomyces cerevisiae. The category with the highest confidence level is selected to identify the sample as the corresponding bacterial species.
[0025] As a preferred embodiment of the method for rapid identification of common pathogenic bacteria and their drug resistance in children based on Raman spectroscopy as described in this invention: In step (5), the Acinetobacter baumannii data is input into a pre-trained convolutional neural network model, and the model outputs the confidence level of the sample being identified as carbapenem-resistant Acinetobacter baumannii or carbapenem-sensitive Acinetobacter baumannii. The category with the highest confidence level is selected to identify the sample as carbapenem-resistant Acinetobacter baumannii or carbapenem-sensitive Acinetobacter baumannii.
[0026] By employing the above technical solutions, this invention has the following advantages and positive effects compared with the prior art:
[0027] 1. This invention innovatively utilizes laser confocal Raman spectroscopy to construct four deep learning typing models that efficiently and accurately distinguish between bacteria and fungi, Gram-positive bacteria and Gram-negative bacteria, common pathogenic bacteria in children, and carbapenem-resistant and susceptible Acinetobacter baumannii. This enables rapid identification of bacteria and fungi, Gram-positive bacteria and Gram-negative bacteria, common pathogenic bacteria in children, and carbapenem-resistant and susceptible Acinetobacter baumannii.
[0028] 2. This invention utilizes deep learning methods based on convolutional neural networks to overcome the shortcomings of traditional analysis methods that require complex preprocessing of Raman spectra. It establishes a Raman spectral database of common pathogenic bacteria in children and innovatively builds four deep learning typing models based on this database: bacteria and fungi, Gram-positive and Gram-negative bacteria, common pathogenic bacteria in children, and carbapenem-resistant and susceptible Acinetobacter baumannii. This provides an efficient and convenient rapid diagnostic tool for identifying common pathogenic bacteria and drug resistance in children. Attached Figure Description
[0029] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below, wherein:
[0030] Figure 1 Raman spectra of six Gram-negative bacteria: Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, and Haemophilus influenzae.
[0031] Figure 2 Raman spectra of three Gram-positive bacteria: Streptococcus agalactiae, Streptococcus pneumoniae, and Staphylococcus aureus.
[0032] Figure 3 Raman spectra of three fungi: Candida albicans, Saccharomyces cerevisiae, and Candida glabrata.
[0033] Figure 4 This is a PCA analysis diagram of two types of data in Example 1 of the present invention: bacteria (including nine types of bacteria: Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, and Staphylococcus aureus) and fungi (including three types of fungi: Candida albicans, Saccharomyces cerevisiae, and Candida glabrata).
[0034] Figure 5 This is the identification result of the two types of data, bacteria and fungi, in Example 1 of the present invention.
[0035] Figure 6 This is a PCA analysis diagram of the two types of data, Gram-negative bacteria and Gram-positive bacteria, in Example 1 of the present invention.
[0036] Figure 7 This is the identification result of Gram-negative and Gram-positive bacteria in Example 1 of the present invention.
[0037] Figure 8 This is a PCA analysis chart of 12 categories of common pathogenic bacteria in children in Example 1 of the present invention.
[0038] Figure 9 This is the identification result of 12 categories of common pathogenic bacteria in children in Example 1 of the present invention.
[0039] Figure 10 Raman spectra of carbapenem-resistant Acinetobacter baumannii and carbapenem-sensitive Acinetobacter baumannii.
[0040] Figure 11 This is a PCA analysis diagram of two types of data in Example 1 of the present invention: carbapenem-resistant and susceptible Acinetobacter baumannii.
[0041] Figure 12 The results of the identification of two types of data in Example 1 of the present invention are carbapenem-resistant and susceptible Acinetobacter baumannii. Detailed Implementation
[0042] To make the above-mentioned objectives, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to specific examples.
[0043] A method for rapid identification of common childhood pathogens and their drug resistance based on Raman spectroscopy with artificial intelligence assistance, comprising the following steps:
[0044] Step 1: Establish a Raman spectral database of common pathogenic bacteria in children; classify and pre-train the classification model;
[0045] Step 2: Fine-tune the bacterial and fungal classification models;
[0046] Step 3: Fine-tune the classification models for Gram-positive and Gram-negative bacteria;
[0047] Step 4: Fine-tune the classification model of common pathogenic bacteria in children;
[0048] Step 5: Fine-tune the classification model of carbapenem-resistant Acinetobacter baumannii and carbapenem-sensitive Acinetobacter baumannii;
[0049] Step 6: Evaluate the performance of each classification model.
[0050] Step 1 includes the following:
[0051] Step 1.1 Strain identification: After resuscitation, the clinical isolates were cultured overnight on culture medium. A suitable amount of single colonies were selected and spread evenly onto the target site of the target plate to form a thin layer. 1 μL of α-cyano-4-hydroxycinnamic acid (HCCA) standard solvent was added and the plate was allowed to air dry at room temperature. The target plate was then placed in a matrix-assisted laser desorption / ionization time-of-flight mass spectrometer (MALDI-TOF-MS) and bacterial identification was performed according to the automated identification process to confirm the strain species (identification result score ≥2.0, and consistency classification A).
[0052] Step 1.2 Antimicrobial susceptibility testing: For strains identified as Acinetobacter baumannii, the minimum inhibitory concentration (MIC) of the antimicrobial agent was determined using a fully automated microbial identification and antimicrobial susceptibility analysis system. Acinetobacter baumannii resistant to any one of the three drugs (ertapenem, imipenem, and meropenem) is called carbapenem-resistant Acinetobacter baumannii, while those sensitive to all three are called carbapenem-sensitive Acinetobacter baumannii.
[0053] Step 1.3 Strain processing procedure: Process Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, Staphylococcus aureus, Candida albicans, Candida parapsilosis, and Saccharomyces cerevisiae. Pick colonies and suspend them in a 1.5 mL centrifuge tube containing 500 μL of sterile deionized water. Repeat the washing 3 times (10000 rpm / min, centrifugation for 5 min) until the turbidity reaches 0.5 McFarland. Set aside 200 μL of pathogenic bacteria sample solution for later use.
[0054] Step 1.4 Raman Spectroscopy Database Establishment: Prepare the pathogenic bacterial sample solutions mentioned in Step 1.3, and drop 3 μL onto an aluminum plate to form droplets approximately 3 mm in diameter. Allow to air dry. Perform Raman spectroscopy detection (laser wavelength: 532 nm; microscope objective: 100x (numerical aperture NA = 0.9); laser power: 0.1-2%; grating line density: 1200 gr / mm; signal acquisition time: 30-60 seconds, single acquisition; signal collection frequency range: 300 cm⁻¹). -1 -2000 cm -1 They collected and preserved the Raman spectra formed by each individual.
[0055] Step 2 includes the following:
[0056] Dataset partitioning: The 30-class classification data of the publicly available common bacteria Raman spectroscopy dataset (HYPERLINK "https: / / www.nature.com / articles / s41467-019-12898-9") were partitioned into a training set of 70%, a validation set of 10%, and a test set of 20%.
[0057] Dataset processing: The Raman spectra were all normalized by themselves as a preprocessing step.
[0058] Data augmentation method: During the model training process, Gaussian random noise is added to augment the data. The mean of the Gaussian noise is 0 and the variance is 0.1.
[0059] Classification Model Building: A 1D convolutional neural network architecture was designed, consisting of an initial convolutional block, multi-scale residual blocks, mean pooling, and fully connected layers. Each convolutional layer is followed by a 1D BatchNorm normalization layer and a ReLU activation function layer. The multi-scale residual blocks reduce signal resolution through convolutions with stride=2, extracting features at six scales. Feature extraction at each scale consists of two residual block network structures. The input data format for the model is (number of channels, spectral length). The fully connected layers output 30 channels, corresponding to the confidence scores of the 30 classes.
[0060] Model training: The ADAMW optimizer was used, with cross-entropy as the loss function. The weights of the model were trained in the training set, and the classification categories were encoded using a one-hot encoding scheme. The initial learning rate was set to 0.01, and the OneCycle learning rate adjustment scheme was used to adjust the learning rate during the optimization process.
[0061] Model testing: Then, based on the model's results on the validation set, the optimal hyperparameters such as the number of network layers and the learning rate are selected to determine the trained model. For the trained novel convolutional neural network model, the classification performance of the model is tested using the test set with accuracy, precision, recall, F1 score, and confusion matrix.
[0062] Step 2.1: The Raman spectroscopy dataset collected in Step 1 is grouped as follows: bacteria (including Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, and Staphylococcus aureus), and fungi (including Candida albicans, Candida glabrata, and Saccharomyces cerevisiae). The proportion of each group is 70% for training, 10% for validation, and 20% for testing.
[0063] Step 2.2 Dataset processing: The Raman spectra in Step 2.1 are all normalized by themselves as preprocessing.
[0064] Step 2.3 Data Augmentation Method: During the model training process, Gaussian random noise is added to augment the data. The mean of the Gaussian noise is 0 and the variance is 0.1.
[0065] Step 2.4 Classification Model Modeling: Modify the number of output channels of the final fully connected layer of the established classification model to 2, corresponding to the confidence scores of the two categories of bacteria and fungi.
[0066] Step 2.5 Fine-tuning the bacterial and fungal classification model: The trained model weights are used as the initial values of the model, and the ADAMW optimizer is used with cross-entropy as the loss function. The weights of the model are trained in the training set, and the classification categories are encoded using a one-hot encoding scheme. The initial learning rate is set to 0.001, and the OneCycle learning rate adjustment scheme is used to adjust the learning rate during the optimization process.
[0067] Step 2.6 Model Testing: For the trained convolutional neural network model, use the test set to test the model's classification performance using accuracy, precision, recall, F1 score, and confusion matrix.
[0068] Step 3 includes the following:
[0069] Step 3.1 The Raman spectroscopy dataset collected in Step 1 is grouped as follows: Gram-negative bacteria (including Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, and Haemophilus influenzae); Gram-positive bacteria (including Streptococcus agalactiae, Streptococcus pneumoniae, and Staphylococcus aureus). The proportion of each group is 70% for training, 10% for validation, and 20% for testing.
[0070] Step 3.2 Dataset processing: The Raman spectra in Step 2.1 are all normalized by themselves as preprocessing.
[0071] Step 3.3 Data Augmentation Method: During the model training process, Gaussian random noise is added to augment the data. The mean of the Gaussian noise is 0 and the variance is 0.1.
[0072] Step 3.4 Classification Model Modeling: Modify the number of output channels of the final fully connected layer of the established classification model to 2, corresponding to the confidence scores of the two categories of Gram-negative bacteria and Gram-positive bacteria.
[0073] Step 3.5 Fine-tuning the Gram-negative and Gram-positive bacteria classification model: The model weights obtained from training are used as the initial values of the model, and the ADAMW optimizer is used with cross-entropy as the loss function. The weights of the model are trained in the training set, and the classification categories are encoded using a one-hot encoding scheme. The initial learning rate is set to 0.001, and the OneCycle learning rate adjustment scheme is used to adjust the learning rate during the optimization process.
[0074] Step 3.6 Model Testing: For the trained convolutional neural network model, use the test set to test the model's classification performance using accuracy, precision, recall, F1 score, and confusion matrix.
[0075] Step 4 includes the following:
[0076] Step 4.1 The Raman spectroscopy dataset collected in Step 1 is grouped as follows: Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, Staphylococcus aureus, Candida albicans, Candida glabrata, and Saccharomyces cerevisiae. The proportion of each group is 70% for training, 10% for validation, and 20% for testing.
[0077] Step 4.2 Dataset processing: The Raman spectra in Step 2.1 are all normalized by themselves as preprocessing.
[0078] Step 4.3 Data Augmentation Method: During the model training process, Gaussian random noise is added to augment the data. The mean of the Gaussian noise is 0 and the variance is 0.1.
[0079] Step 4.4 Classification Model Modeling: Modify the number of output channels of the final fully connected layer of the established classification model to 12, corresponding to the confidence scores of the 12 categories: Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, Staphylococcus aureus, Candida albicans, Candida glabrata, and Saccharomyces cerevisiae.
[0080] Step 4.5 Fine-tuning the classification model of common pathogenic bacteria in children: The model weights obtained from training are used as the initial values of the model, and the ADAMW optimizer is used with cross-entropy as the loss function. The weights of the model are trained in the training set, and the classification categories are encoded using a one-hot encoding scheme. The initial learning rate is set to 0.001, and the OneCycle learning rate adjustment scheme is used to adjust the learning rate during the optimization process.
[0081] Step 4.6 Model Testing: For the trained convolutional neural network model, use the test set to test the model's classification performance using accuracy, precision, recall, F1 score, and confusion matrix.
[0082] Step 5 includes the following:
[0083] Step 5.1 The Raman spectroscopy dataset collected in Step 1 is grouped as follows: carbapenem-resistant and susceptible Acinetobacter baumannii. The proportion of data in each group is 70% for training set, 10% for validation set, and 20% for test set;
[0084] Step 5.2 Dataset processing: The Raman spectra in Step 2.1 are all normalized by themselves as preprocessing;
[0085] Step 5.3 Data Augmentation Method: During the model training process, Gaussian random noise is added to augment the data. The mean of the Gaussian noise is 0 and the variance is 0.1.
[0086] Step 5.4 Classification Model Modeling: Modify the number of output channels of the final fully connected layer of the established classification model to 2, corresponding to the confidence scores of the two categories of carbapenem-resistant and susceptible Acinetobacter baumannii.
[0087] Step 5.5 Fine-tuning the classification model of carbapenem-resistant Acinetobacter baumannii and carbapenem-sensitive Acinetobacter baumannii: The model weights obtained from training are used as the initial values of the model, and the ADAMW optimizer is used with cross-entropy as the loss function. The weights of the model are trained in the training set, and the classification categories are encoded using a one-hot encoding scheme. The initial learning rate is set to 0.001, and the OneCycle learning rate adjustment scheme is used to adjust the learning rate during the optimization process.
[0088] Step 5.6 Model Testing: For the trained convolutional neural network model, use the test set to test the model's classification performance using accuracy, precision, recall, F1 score, and confusion matrix.
[0089] Step 6 includes the following:
[0090] The constructed typing model was evaluated for laboratory specificity and accuracy using subsequently collected pathogenic bacteria and a blank control group. The model was then applied to 100 clinical bacterial strains to assess its sensitivity, specificity, and accuracy. MALDI-TOF MS identification was used as the standard for evaluation, and divergence results were validated using whole-genome sequencing.
[0091] Example 1:
[0092] like Figure 1 As shown, this invention discloses a method for rapid identification of common childhood pathogens and their drug resistance based on Raman spectroscopy with artificial intelligence assistance, comprising the following steps:
[0093] Step 1: Establish a Raman spectral database of common pathogenic bacteria in children
[0094] Step 1 includes the following:
[0095] Step 1.1 Strain identification: After resuscitating the clinical isolates, incubate them overnight on culture medium. Select an appropriate amount of single colonies and spread them evenly onto the target site of the target plate to form a thin layer. Cover with 1 μL of HCCA standard solvent and allow to air dry at room temperature. Place the target plate into a matrix-assisted laser desorption / ionization time-of-flight mass spectrometer (MALDI-TOF-MS) and perform bacterial identification according to the automated identification process to confirm the strain species (identification result score ≥2.0, and consistency classification A).
[0096] Step 1.2 Antimicrobial susceptibility testing: For strains identified as Acinetobacter baumannii, the minimum inhibitory concentration (MIC) of the antimicrobial agent was determined using a fully automated microbial identification and antimicrobial susceptibility analysis system. Acinetobacter baumannii resistant to any one of the three drugs (ertapenem, imipenem, and meropenem) is called carbapenem-resistant Acinetobacter baumannii, while those sensitive to all three are carbapenem-sensitive Acinetobacter baumannii.
[0097] Step 1.3 Strain processing procedure: Process Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, Staphylococcus aureus, Candida albicans, Candida parapsilosis, and Saccharomyces cerevisiae. Pick colonies and suspend them in a 1.5 mL centrifuge tube containing 500 μL of sterile deionized water. Repeat the washing 3 times (12000 rpm / min, centrifuge for 3 min) until the turbidity reaches 0.5 McFarland. Finally, reserve 200 μL for later use.
[0098] Step 1.4 Raman Spectroscopy Database Establishment: Prepare sample solutions of the pathogenic bacteria mentioned in Step 1.3, and drop 3 μL onto an aluminum plate to form droplets approximately 3 mm in diameter. Allow to air dry. Perform Raman spectroscopy detection (laser wavelength: 532 nm; microscope objective: 100x (numerical aperture NA = 0.9); laser power: 1%; grating line density: 1200 gr / mm; signal acquisition time: 30-60 seconds, single acquisition; signal acquisition frequency range: 300 cm⁻¹). -1 -2000 cm -1 They collected and preserved the Raman spectra formed by each individual.
[0099] Step 2: Establish identification models for common bacteria and fungi in children.
[0100] Step 2 includes the following:
[0101] Step 2.1 The Raman spectroscopy dataset collected in Step 1 is grouped as follows: bacteria (including Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, and Staphylococcus aureus), and fungi (including Candida albicans, Candida glabrata, and Saccharomyces cerevisiae). The proportion of each group is 70% for training, 10% for validation, and 20% for testing.
[0102] Step 2.2 Dataset processing: The Raman spectra collected in Step 2.1 are all normalized as preprocessing.
[0103] Step 2.3 Data Augmentation Method: During the model training process, Gaussian random noise is added to augment the data. The mean of the Gaussian noise is 0 and the variance is 0.1.
[0104] Step 2.4 Model Structure Design: Improve the ResNet18 residual convolutional neural network structure by replacing the 2D convolution operation with a 1D convolution operation, replacing the 2D BatchNorm operation with a 1D BatchNorm operation, and changing the input data format to (number of channels, spectral length).
[0105] Step 2.5 Model Training: The ADAM optimizer is used, with cross-entropy as the loss function and the L2 norm of the weights added as a regularization term. The weights of the model are trained on the training set, and then the optimal hyperparameters such as the number of network layers and the learning rate are selected based on the model's results on the validation set to determine the trained model.
[0106] Step 2.6 Model Testing: For the trained novel convolutional neural network model, use the test set to test the model's classification performance using accuracy, precision, recall, F1 score, and confusion matrix.
[0107] Step 3: Establish an identification model for common Gram-positive and Gram-negative bacteria in children.
[0108] Step 3 includes the following:
[0109] Step 3.1 The Raman spectroscopy dataset collected in Step 1 is grouped as follows: Gram-negative bacteria (including Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, and Haemophilus influenzae); Gram-positive bacteria (including Streptococcus agalactiae, Streptococcus pneumoniae, and Staphylococcus aureus). The proportion of each group is 70% for training, 10% for validation, and 20% for testing.
[0110] Step 3.2 Dataset processing: The Raman spectra collected in Step 3.1 are all normalized as preprocessing.
[0111] Step 3.3 Data Augmentation Method: During the model training process, Gaussian random noise is added to augment the data. The mean of the Gaussian noise is 0 and the variance is 0.1.
[0112] Step 3.4 Model Structure Design: Improve the ResNet18 residual convolutional neural network structure by replacing the 2D convolution operation with a 1D convolution operation, replacing the 2D BatchNorm operation with a 1D BatchNorm operation, and changing the input data format to (number of channels, spectral length).
[0113] Step 3.5 Model Training: The ADAM optimizer is used, with cross-entropy as the loss function and the L2 norm of the weights added as a regularization term. The model weights are trained on the training set, and then the optimal hyperparameters such as the number of network layers and the learning rate are selected based on the model's performance on the validation set to determine the trained model.
[0114] Step 3.6 Model Testing: For the trained novel convolutional neural network model, use the test set to test the model's classification performance using accuracy, precision, recall, F1 score, and confusion matrix.
[0115] Step 4: Establish identification models for 12 common pathogenic bacteria species in children.
[0116] Step 4 includes the following:
[0117] Step 4.1 The Raman spectroscopy dataset collected in Step 1 is grouped as follows: Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, Staphylococcus aureus, Candida albicans, Candida glabrata, and Saccharomyces cerevisiae. The proportion of each group is 70% for training, 10% for validation, and 20% for testing.
[0118] Step 4.2 Dataset processing: The Raman spectra collected in Step 1 are all normalized by themselves as preprocessing.
[0119] Step 4.3 Data Augmentation Method: During the model training process, Gaussian random noise is added to augment the data. The mean of the Gaussian noise is 0 and the variance is 0.1.
[0120] Step 4.4 Model Structure Design: Improve the ResNet18 residual convolutional neural network structure by replacing the 2D convolution operation with a 1D convolution operation, replacing the 2D BatchNorm operation with a 1D BatchNorm operation, and changing the input data format to (number of channels, spectral length).
[0121] Step 4.5 Model Training: The ADAM optimizer is used, with cross-entropy as the loss function and the L2 norm of the weights added as a regularization term. The model weights are trained on the training set, and then the optimal hyperparameters such as the number of network layers and learning rate are selected based on the model's performance on the validation set to determine the trained model.
[0122] Step 4.6 Model Testing: For the trained novel convolutional neural network model, use the test set to test the model's classification performance using accuracy, precision, recall, F1 score, and confusion matrix.
[0123] Step 5: Establish an identification model for carbapenem-resistant and susceptible Acinetobacter baumannii.
[0124] Step 5 includes the following:
[0125] Step 5.1 The Raman spectroscopy dataset collected in Step 1 is grouped as follows: carbapenem-resistant and susceptible Acinetobacter baumannii. The proportion of data in each group is 70% for training set, 10% for validation set, and 20% for test set.
[0126] Step 5.2 Dataset processing: The Raman spectra collected in Step 1 are all normalized by themselves as preprocessing.
[0127] Step 5.3 Data Augmentation Method: During the model training process, Gaussian random noise is added to augment the data. The mean of the Gaussian noise is 0 and the variance is 0.1.
[0128] Step 5.4 Model Structure Design: Improve the ResNet18 residual convolutional neural network structure by replacing the 2D convolution operation with a 1D convolution operation, replacing the 2D BatchNorm operation with a 1D BatchNorm operation, and changing the input data format to (number of channels, spectral length).
[0129] Step 5.5 Model Training: The ADAM optimizer is used, with cross-entropy as the loss function and the L2 norm of the weights added as a regularization term. The model weights are trained on the training set, and then the optimal hyperparameters such as the number of network layers and the learning rate are selected based on the model's performance on the validation set to determine the trained model.
[0130] Step 5.6 Model Testing: For the trained novel convolutional neural network model, use the test set to test the model's classification performance using accuracy, precision, recall, F1 score, and confusion matrix.
[0131] Step 6 includes the following:
[0132] The constructed typing model was evaluated for laboratory specificity and accuracy using subsequently collected pathogenic bacteria and a blank control group. The identification model was then evaluated using 100 clinical bacterial strains to assess its sensitivity, specificity, and accuracy. MALDI-TOF MS identification was used as the standard for evaluation, and divergence results were validated using whole-genome sequencing.
[0133] The sample data is input into a pre-trained bacterial and fungal identification model. Deep learning analysis is then used to analyze the test group data to determine whether it is bacteria or fungi. When testing the data, the model outputs the confidence level of identifying the data as bacteria or fungi, and the category with the highest confidence level is selected to identify the sample as either bacteria or fungi. The identification accuracy reaches 99.85%. Figure 5 As shown.
[0134] The test sample data is input into a pre-trained convolutional neural network model. The model outputs the confidence level at which the test sample is identified as Gram-positive or Gram-negative bacteria. The category with the highest confidence level is selected to identify the sample as either Gram-positive or Gram-negative bacteria. The identification accuracy reaches 94.70%. Figure 7 As shown.
[0135] The sample data is input into a pre-trained identification model for 12 common childhood pathogens. The model outputs the confidence level of the sample being identified as one of twelve strains: Escherichia coli (Eco), Klebsiella pneumoniae (Kpn), Pseudomonas aeruginosa (Pa), Acinetobacter baumannii (Ab), Salmonella (Sal), Haemophilus influenzae (Hin), Streptococcus agalactiae (Sag), Streptococcus pneumoniae (Sp), Staphylococcus aureus (Sa), Candida albicans (Cal), Candida parapsilosis (Cpa), and Saccharomyces cerevisiae (Sce). The category with the highest confidence level is selected to identify the sample as the corresponding bacterial species. The identification accuracy reaches 95.08%. Figure 9 As shown.
[0136] Input the sample data into a pre-trained model for identifying carbapenem-resistant and carbapenem-sensitive Acinetobacter baumannii. The model outputs the confidence level for identifying the sample as either carbapenem-resistant or carbapenem-sensitive Acinetobacter baumannii. The category with the highest confidence level is selected to identify the sample as either carbapenem-resistant Acinetobacter baumannii (AB IPM R) or carbapenem-sensitive Acinetobacter baumannii (AB IPM S). The identification accuracy reaches 96.20%. Figure 12 As shown.
[0137] Comparative Example 1:
[0138] Table 1 compares the accuracy of different classification methods in identifying 12 categories of common pathogenic bacteria in children.
[0139] Table 1 Comparison of accuracy rates of different classification methods
[0140]
[0141] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A method for rapid identification of common childhood pathogens and their drug resistance based on Raman spectroscopy with artificial intelligence assistance, characterized in that: Includes the following steps: (1) Collection of pathogenic bacteria samples: Twelve common pathogenic bacteria were isolated from clinical specimens. The twelve common pathogenic bacteria included bacteria and fungi. The bacteria were Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, and Staphylococcus aureus. The fungi were Candida albicans, Candida glabrata, and Saccharomyces cerevisiae. The twelve common pathogenic bacteria were identified by matrix-assisted laser desorption / ionization time-of-flight mass spectrometry. (2) The strains identified as Acinetobacter baumannii were classified into carbapenem-resistant Acinetobacter baumannii and carbapenem-sensitive Acinetobacter baumannii according to their minimum inhibitory concentration of antimicrobial drugs. (3) Establish a Raman spectroscopy database; (4) Establish a Raman spectroscopy deep learning classification model and pre-train the weights on the bacterial Raman spectroscopy dataset; The design of the 1D convolutional neural network model consists of an initial convolutional block, multi-scale residual blocks, mean pooling, and fully connected layers. The collected Raman spectral data, after being normalized, is randomly divided into three sets: a training set, a test set, and a validation set, which are then input into the 1D convolutional neural network model for processing. Gaussian random noise is added to augment the data during the training process of the 1D convolutional neural network model. (5) By modifying the output channel of the last fully connected layer of the 1D convolutional neural network model, deep learning classification models for bacteria and fungi, Gram-positive and Gram-negative bacteria, the 12 common pathogenic bacteria, and carbapenem-resistant and carbapenem-sensitive Acinetobacter baumannii were established respectively. The pre-trained deep learning classification models were used as initial weights, and the corresponding deep learning classification models were fine-tuned on the training set of the corresponding dataset. The classification effect of the deep learning classification models was tested using the test set with accuracy, precision, recall, F1 score and confusion matrix.
2. The method for rapid identification of common childhood pathogens and their drug resistance based on Raman spectroscopy with artificial intelligence assistance, as described in claim 1, is characterized in that: Step (3) establishing the Raman spectral database involves detecting the 12 common pathogenic bacteria mentioned in step (1) using a laser confocal Raman spectrometer, collecting and storing the resulting Raman spectral database; wherein, a laser confocal Raman spectrometer is selected, with a laser wavelength of 532 nm, a laser power of 0.1-2%, a spectrometer grating of 1200 g / mm, and a spectral range of 300-2000 cm⁻¹. -1 The acquisition time for a single spectrum is 30-60 seconds.
3. The method for rapid identification of common childhood pathogens and their drug resistance based on Raman spectroscopy with artificial intelligence assistance, as described in claim 1 or 2, is characterized in that: In step (3), the pathogenic bacteria are suspended in sterile water with a turbidity of 0.5 McFarland degrees and dropped onto the surface of an aluminum sheet for detection using a laser confocal Raman spectrometer.
4. The method for rapid identification of common childhood pathogens and their drug resistance based on Raman spectroscopy with artificial intelligence assistance, as described in claim 1, is characterized in that: In step (4), the 1D convolutional neural network model is trained using the ADAMW optimizer. The 1D convolutional neural network model is trained on the bacterial Raman spectroscopy dataset. Then, the optimal hyperparameters are selected based on the results of the 1D convolutional neural network model on the validation set, and the trained 1D convolutional neural network model is determined. For the trained 1D convolutional neural network model, the classification performance of the 1D convolutional neural network model is tested using the test set with accuracy, precision, recall, F1 score and confusion matrix.
5. The method for rapid identification of common childhood pathogens and their drug resistance based on Raman spectroscopy with artificial intelligence assistance, as described in claim 1 or 2, is characterized in that: In step (4), the proportion of normalized Raman spectral data in each group is 70% for the training set, 10% for the validation set, and 20% for the test set; the mean of Gaussian random noise is 0 and the variance is 0.
1.
6. The method for rapid identification of common childhood pathogens and their drug resistance based on Raman spectroscopy with artificial intelligence assistance, as described in claim 1 or 2, is characterized in that: In step (4), the 1D convolutional neural network model has a 1D BatchNorm normalization layer and a ReLU activation function layer after each convolutional layer; the multi-scale residual blocks reduce the signal resolution through convolution with stride=2, and extract features at 6 scales respectively. The feature extraction at each scale is composed of 2 residual block network structures; the input data format for constructing the 1D convolutional neural network model is: number of channels, spectral length.
7. The method for rapid identification of common childhood pathogens and their drug resistance based on Raman spectroscopy with artificial intelligence assistance, as described in claim 1, is characterized in that: In step (5), the Raman spectral data of the sample to be tested is input into the pre-trained 1D convolutional neural network model. The 1D convolutional neural network model will output the confidence level of the sample to be tested as bacteria or fungi. The category with the highest confidence level is selected to identify the sample as bacteria or fungi. The data of the sample to be tested is input into the pre-trained 1D convolutional neural network model. The 1D convolutional neural network model will output the confidence level of the sample to be tested as Gram-positive bacteria or Gram-negative bacteria. The category with the highest confidence level is selected to identify the sample as Gram-positive bacteria or Gram-negative bacteria.
8. The method for rapid identification of common childhood pathogens and their drug resistance based on Raman spectroscopy with artificial intelligence assistance according to claim 1, characterized in that: In step (5), the test sample data is input into the pre-trained 1D convolutional neural network model. The 1D convolutional neural network model will output the confidence level of the test sample as one of twelve strains: Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, Staphylococcus aureus, Candida albicans, Candida glabrata, and Saccharomyces cerevisiae. The category with the highest confidence level is selected to identify the sample as the corresponding bacterial species.
9. The method for rapid identification of common childhood pathogens and their drug resistance based on Raman spectroscopy with artificial intelligence assistance, as described in claim 1, is characterized in that: In step (5), the test sample data of Acinetobacter baumannii is input into the trained 1D convolutional neural network model. The 1D convolutional neural network model will output the confidence level of the test sample being identified as carbapenem-resistant Acinetobacter baumannii or carbapenem-sensitive Acinetobacter baumannii. The category with the highest confidence level is selected to identify the test sample as carbapenem-resistant Acinetobacter baumannii or carbapenem-sensitive Acinetobacter baumannii.
Citation Information
Patent Citations
Rapid identification method for carbapenem drug susceptibility, based on Raman spectra technology
CN107586823A
Training method of food borne pathogenic bacteria Raman spectrum identification method established based on PCA-Stacking
CN109781706A