Method for rapidly identifying common pathogenic bacteria of children and drug resistance thereof by artificial intelligence assistance based on Raman spectrum

By combining Raman spectroscopy and deep learning algorithms, a 1D convolutional neural network model is established to quickly and accurately identify common pathogenic bacteria in children and their drug resistance, solving the problem of long-term identification and low positive rate in the existing technology, and improving the accuracy of treatment and drug resistance prevention and control capabilities.

CN120293947AActive Publication Date: 2025-07-11ZHEJIANG UNIV
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
CN202510764934.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-10
Publication Date
2025-07-11
Estimated Expiration
2045-06-10

AI Technical Summary

Technical Problem

The prior art is difficult to quickly and accurately identify common pathogenic bacteria and their drug resistance in children, resulting in inaccurate clinical treatment and spread of drug resistance. The existing methods take a long time and have a low positive rate.

Method used

Using a method based on the combination of Raman spectroscopy and deep learning algorithm, a 1D convolutional neural network model was established, and a laser confocal Raman spectrometer was used to collect spectral data of common pathogenic bacteria in children, and conducted deep learning training to construct deep learning classification models of bacteria and fungi, Gram-positive and negative bacteria, common pathogenic bacteria strains in children, and carbapenem-resistant and sensitive Acinetobacter baumannii.

Benefits of technology

It has achieved rapid and accurate identification of common pathogenic bacteria and their drug resistance in children, improved the accuracy of treatment, reduced the risk of drug resistance spread, and provided an efficient and convenient rapid diagnostic tool.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120293947A_ABST
    Figure CN120293947A_ABST
Patent Text Reader

Abstract

The invention discloses a method for rapidly identifying common pathogenic bacteria and drug resistance of children based on artificial intelligence assistance of Raman spectrum. The method comprises the following steps: separating the common pathogenic bacteria; rechecking 12 common pathogenic bacteria strains by using matrix-assisted laser desorption time-of-flight mass spectrometry; determining the carbapenem drug-resistant and sensitive acinetobacter baumannii according to a drug sensitivity result; constructing a Raman spectrum database of common pathogenic bacteria of children; and four deep learning typing models are established. According to the invention, rapid detection of classification of bacteria and fungi, rapid detection of classification of gram-positive bacteria and gram-negative bacteria, rapid detection of classification of 12 common pathogenic bacteria strains for children and rapid detection of classification of carbapenem drug-resistant and sensitive acinetobacter baumannii can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of artificial intelligence, and particularly relates to a method for rapidly identifying common pathogenic bacteria in children and their drug resistance by artificial intelligence assistance based on Raman spectroscopy. Background Art

[0002] There are many similarities in clinical manifestations of different pathogen infections, and it is often difficult to distinguish them. However, the drugs used for treatment vary depending on the bacterial species. Bacterial infections require the use of antibacterial drugs, while fungal infections require antifungal drugs. Identifying whether it is a bacterial or fungal infection in the early stage of infection can help clinicians formulate targeted treatment plans and provide reference for clinically selecting effective antibacterial drugs.

[0003] Bacteria are classified into Gram-positive bacteria and Gram-negative bacteria according to Gram staining characteristics. Their pathogenic mechanisms are different. Gram-positive bacteria mainly rely on exotoxins and invasive enzymes, while Gram-negative bacteria mainly rely on endotoxins. And there are differences in antibacterial drug sensitivity between them. Gram-positive bacteria are more sensitive to penicillins and vancomycin, while Gram-negative bacteria usually require third-generation cephalosporins, β-lactamase inhibitor combinations, or carbapenems. On the premise of determining that it is a bacterial infection, early determination of Gram-positive and Gram-negative bacterial infections can provide a basis for clinically selecting the types of antibacterial drugs.

[0004] In addition, the sensitivity of different species of pathogens to antibacterial drugs varies significantly. Identifying the species can avoid blind drug use. Determining the specific species of common pathogenic bacteria can enable precise treatment, reduce the risk of treatment failure, and improve the prognosis of patients.

[0005] The problem of bacterial drug resistance caused by the abuse and overuse of antibacterial drugs has attracted global attention, especially Acinetobacter baumannii resistant to carbapenems. Carbapenem-resistant Acinetobacter baumannii (CRAB) is resistant to most β-lactam antibacterial drugs and requires the selection of second-line drugs such as polymyxins and tigecycline, but the efficacy is limited and the toxicity is relatively high. And because such strains can parasitize in parts such as the respiratory tract for up to several months, it leads to the spread of drug-resistant bacteria in the hospital and ultimately causes nosocomial infections, bringing great challenges to clinical treatment. Early identification of CRAB can avoid the clinical use of ineffective carbapenem drugs, reduce the risk of spread of drug-resistant genes, and contain the spread of drug resistance. CRAB is easily transmitted through contact, leading to nosocomial outbreaks. Early differentiation of its drug resistance helps to implement isolation and prevention and control as early as possible and block the transmission chain.

[0006] The above classifications are directly related to clinical precision treatment, drug resistance prevention and control, and the prognosis of patients. By rapidly identifying the pathogen type and drug resistance characteristics, individualized treatment plans can be formulated, thereby improving the curative effect, shortening the course of disease, and containing drug resistance. Summary of the Invention

[0007] To make up for the deficiencies in existing pathogenic bacteria identification technologies, the present invention provides a method for rapidly identifying common pathogenic bacteria in children and their drug resistance with the assistance of artificial intelligence based on Raman spectroscopy.

[0008] By combining in-situ Raman spectroscopy with deep learning algorithms, problems such as the cumbersome process, long time consumption, and low positive rate in the identification of pathogenic bacteria are solved.

[0009] The present invention uses deep learning methods to train the Raman spectra of common pathogenic bacteria in children, thereby establishing an identification model that can correctly identify pathogenic bacteria. By collecting the Raman spectra of common pathogenic bacteria in children, a corresponding Raman spectroscopy database is established. The data is divided into three categories: a training group, a test group, and a validation group. The data is input into the proposed neural network model for processing. Deep learning training is performed on the training group data, the model with the best performance is selected, the validation group data is used for validation, the model parameters are adjusted, and finally the test bacteria are used to test the identification effect of the model.

[0010] A method for rapidly identifying common pathogenic bacteria in children and their drug resistance with the assistance of artificial intelligence based on Raman spectroscopy, which includes the following steps: (1) Collect pathogenic bacteria samples: Isolate 12 common pathogenic bacteria from clinical specimens. The 12 common pathogenic bacteria include bacteria and fungi. The bacteria are Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, and Staphylococcus aureus; the fungi are Candida albicans, Candida parapsilosis, and Saccharomyces cerevisiae; Determine the 12 common pathogenic bacteria by matrix-assisted laser desorption / ionization time-of-flight mass spectrometry; (2) For the strains identified as Acinetobacter baumannii, divide them into carbapenem-resistant Acinetobacter baumannii and carbapenem-susceptible Acinetobacter baumannii according to their minimum inhibitory concentration against antibacterial drugs; (3) Establish a Raman spectroscopy database: Detect the 12 common pathogenic bacteria in step (1) using a confocal Raman spectrometer, and collect and save the generated Raman spectroscopy database; (4) Establish a deep learning classification model for Raman spectroscopy and pre-train the weights on 30 classification problems of the bacterial Raman dataset publicly available on the Internet; Among them, a 1D convolutional neural network architecture is designed. The neural network consists of an initial convolutional block, multi-scale residual blocks, average pooling, and a fully connected layer. After the collected Raman spectra are normalized by themselves, the data is randomly divided into three groups: a training set, a test set, and a validation set, and then input into the neural network model for processing. During the model training process, the method of adding Gaussian random noise is used to enhance the data. The ADAMW optimizer is used for model training. The model is trained on 30-class classification problems of the publicly available bacterial Raman dataset on the Internet. Then, the optimal hyperparameters such as the learning rate and channels are selected based on the results of the model on the validation set to determine the trained model. For the trained model, the test set is used to test the classification effect of the model with accuracy, precision, recall, F1-score, and confusion matrix. (5) By modifying the output channels of the last fully connected layer of the model, a deep learning classification model for bacteria and fungi, a deep learning classification model for Gram-positive and Gram-negative bacteria, a deep learning classification model for the 12 common pathogenic bacteria, and a deep learning classification model for carbapenem-resistant Acinetobacter baumannii and carbapenem-sensitive Acinetobacter baumannii are respectively established. Using the pre-trained model as the initial weight, the corresponding model is fine-tuned on the training set of the corresponding dataset, and the test set is used to test the classification effect of the model with accuracy, precision, recall, F1-score, and confusion matrix.

[0011] As a preferred embodiment of the method for artificial intelligence-assisted rapid identification of common pathogenic bacteria and their drug resistance in children based on Raman spectroscopy according to the present invention: In step (3), a laser confocal Raman spectrometer is selected, the laser wavelength is 532 nm, the laser power is set to 0.1 - 2%, the spectrometer grating is selected as 1200 g / mm, and the spectral range is 300 - 2000 cm -1 , and the single-spectrum acquisition time is 30 - 60 s.

[0012] As a preferred embodiment of the method for artificial intelligence-assisted rapid identification of common pathogenic bacteria and their drug resistance in children based on Raman spectroscopy according to the present invention: In step (3), the pathogenic bacteria are suspended with sterile water to a turbidity of 0.5 McFarland degrees, dropped on the surface of an aluminum sheet, and detected using a laser confocal Raman spectrometer.

[0013] As a preferred embodiment of the method for artificial intelligence-assisted rapid identification of common pathogenic bacteria and their drug resistance in children based on Raman spectroscopy according to the present invention: In step (4), the proportion of each group of data is 70% for the training set, 10% for the validation set, and 20% for the test set.

[0014] As a preferred embodiment of the method for artificial intelligence-assisted rapid identification of common pathogenic bacteria and their drug resistance in children based on Raman spectroscopy according to the present invention: In step (4), the mean of the Gaussian random noise is 0, and the variance is 0.1.

[0015] As a preferred embodiment of the method for rapid identification of common pathogenic bacteria and their drug resistance in children assisted by artificial intelligence based on Raman spectroscopy according to the present invention: In step (4), a 1D convolutional neural network architecture is used, where there is a 1D BatchNorm normalization layer and a ReLU activation function layer after each convolutional layer; the multi-scale residual block reduces the signal resolution through a convolution with a stride of 2, extracts features at 6 scales respectively, and the feature extraction at each scale is composed of 2 residual block network structures; the input data format for constructing the model is: number of channels, spectral length.

[0016] As a preferred embodiment of the method for rapid identification of common pathogenic bacteria and their drug resistance in children assisted by artificial intelligence based on Raman spectroscopy according to the present invention: In step (5), the data of the sample to be tested is input into the trained convolutional neural network model, and the model will output the confidence level that the sample to be tested is identified as a bacterium or a fungus, and the category with the maximum confidence level is selected to identify the sample as a bacterium or a fungus.

[0017] As a preferred embodiment of the method for rapid identification of common pathogenic bacteria and their drug resistance in children assisted by artificial intelligence based on Raman spectroscopy according to the present invention: In step (5), the data of the sample to be tested is input into the trained convolutional neural network model, and the model will output the confidence level that the sample to be tested is identified as a Gram-positive bacterium or a Gram-negative bacterium, and the category with the maximum confidence level is selected to identify the sample as a Gram-positive bacterium or a Gram-negative bacterium.

[0018] As a preferred embodiment of the method for rapid identification of common pathogenic bacteria and their drug resistance in children assisted by artificial intelligence based on Raman spectroscopy according to the present invention: In step (5), the data of the sample to be tested is input into the trained convolutional neural network model, and the model will output the confidence levels that the sample to be tested is identified as twelve strains including Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, Staphylococcus aureus, Candida albicans, Candida parapsilosis, and Saccharomyces cerevisiae, and the category with the maximum confidence level is selected to identify the sample as the corresponding bacterial species.

[0019] As a preferred embodiment of the method for rapid identification of common pathogenic bacteria and their drug resistance in children assisted by artificial intelligence based on Raman spectroscopy according to the present invention: In step (5), the data of Acinetobacter baumannii is input into the trained convolutional neural network model, and the model will output the confidence levels that the sample to be tested is identified as carbapenem-resistant Acinetobacter baumannii and carbapenem-sensitive Acinetobacter baumannii, and the category with the maximum confidence level is selected to identify the sample as carbapenem-resistant Acinetobacter baumannii or carbapenem-sensitive Acinetobacter baumannii.

[0020] Due to the adoption of the above technical solutions, the present invention has the following advantages and positive effects compared with the prior art: 1. The present invention innovatively uses a laser confocal Raman spectrometer to construct four deep learning classification models for efficiently and accurately distinguishing bacteria and fungi, Gram-positive bacteria and Gram-negative bacteria, common pathogenic bacteria species in children, and carbapenem-resistant and sensitive Acinetobacter baumannii, realizing the rapid identification of bacteria and fungi, Gram-positive bacteria and Gram-negative bacteria, common pathogenic bacteria species in children, and carbapenem-resistant and sensitive Acinetobacter baumannii; 2. The present invention uses the deep learning method of convolutional neural network, overcomes the disadvantage that traditional analysis methods require complex preprocessing of Raman spectra, establishes a Raman spectrum database of common pathogenic bacteria in children, and innovatively based on this database, establishes four deep learning classification models for bacteria and fungi, Gram-positive bacteria and Gram-negative bacteria, common pathogenic bacteria species in children, and carbapenem-resistant and sensitive Acinetobacter baumannii, providing an efficient and convenient rapid diagnosis tool for identifying common pathogenic bacteria and drug resistance in children. Description of the Drawings

[0021] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required for the description of the embodiments will be briefly introduced below, where: Figure 1 Are Raman spectra of six Gram-negative bacteria: Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, and Haemophilus influenzae.

[0022] Figure 2 Are Raman spectra of three Gram-positive bacteria: Streptococcus agalactiae, Streptococcus pneumoniae, and Staphylococcus aureus.

[0023] Figure 3 Are Raman spectra of three fungi: Candida albicans, Saccharomyces cerevisiae, and Candida parapsilosis.

[0024] Figure 4 Is the PCA analysis chart of two types of data of bacteria (including nine bacteria: Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, and Staphylococcus aureus) and fungi (including three fungi: Candida albicans, Saccharomyces cerevisiae, and Candida parapsilosis) in Example 1 of the present invention.

[0025] Figure 5 Is the discrimination result of two types of data of bacteria and fungi in Example 1 of the present invention.

[0026] Figure 6 Is the PCA analysis chart of two types of data of Gram-negative bacteria and Gram-positive bacteria in Example 1 of the present invention.

[0027] Figure 7The identification results of two types of data, Gram-negative bacteria and Gram-positive bacteria, in Example 1 of the present invention.

[0028] Figure 8 The PCA analysis chart of 12 types of data of common pathogenic bacteria in children in Example 1 of the present invention.

[0029] Figure 9 The identification results of 12 types of data of common pathogenic bacteria in children in Example 1 of the present invention.

[0030] Figure 10 The Raman spectra of carbapenem-resistant Acinetobacter baumannii and carbapenem-sensitive Acinetobacter baumannii.

[0031] Figure 11 The PCA analysis chart of two types of data of carbapenem-resistant and carbapenem-sensitive Acinetobacter baumannii in Example 1 of the present invention.

[0032] Figure 12 The identification results of two types of data of carbapenem-resistant and carbapenem-sensitive Acinetobacter baumannii in Example 1 of the present invention. Detailed implementation manners

[0033] To make the above objects, features, and advantages of the present invention more obvious and understandable, the following detailed description of the specific implementation manners of the present invention will be given in combination with specific embodiments.

[0034] A method for rapid identification of common pathogenic bacteria in children and their drug resistance based on Raman spectroscopy assisted by artificial intelligence, comprising the following steps: Step 1: Establish a Raman spectroscopy database of common pathogenic bacteria in children; build and pre-train a classification model; Step 2: Fine-tune the classification model of bacteria and fungi; Step 3: Fine-tune the classification models of Gram-positive bacteria and Gram-negative bacteria; Step 4: Fine-tune the classification model of common pathogenic bacteria species in children; Step 5: Fine-tune the classification model of carbapenem-resistant Acinetobacter baumannii and carbapenem-sensitive Acinetobacter baumannii; Step 6: Evaluate the performance of each classification model.

[0035] Step 1 includes the following content: Step 1.1 Strain identification: After clinical isolates are revived, they are cultured on culture medium overnight. An appropriate number of single colonies are selected and spread evenly on the target points of the target plate to form a thin layer. The target plate is covered with 1 μL of α-cyano-4-hydroxycinnamic acid (HCCA) standard solvent and allowed to dry naturally at room temperature. The target plate is placed in matrix-assisted laser desorption ionization time-of-flight mass spectrometry (MALDI-TOF-MS) and bacterial identification is performed according to the automatic identification process to confirm the strain species (identification result score ≥ 2.0, and consistency classification is A).

[0036] Step 1.2 Drug sensitivity identification: The strains were identified as Acinetobacter baumannii, and the minimum inhibitory concentration of antimicrobial drugs was detected by a fully automatic microbial identification and drug sensitivity analysis system. Acinetobacter baumannii that was resistant to any one of the three drugs, ertapenem, imipenem, and meropenem, was called carbapenem-resistant Acinetobacter baumannii, and those that were sensitive to all three drugs were called carbapenem-sensitive Acinetobacter baumannii.

[0037] Step 1.3 Strain processing flow: For strains of Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, Staphylococcus aureus, Candida albicans, Candida parapsilosis and Saccharomyces cerevisiae, pick colonies and suspend them in a 1.5 mL centrifuge tube containing 500 μL of sterile deionized water. Repeat washing 3 times (10,000 rpm / min, centrifugation for 5 min) until the turbidity reaches 0.5 McLaren degrees. Keep 200 μL of pathogenic bacteria sample solution for later use.

[0038] Step 1.4 Raman spectroscopy database establishment: Prepare the pathogenic bacteria sample liquid mentioned in step 1.3 respectively, draw 3 μL and drop it on the aluminum sheet to form a droplet with a diameter of about 3 mm, and dry it naturally. Perform Raman spectroscopy detection (laser wavelength: 532 nm; microscope objective magnification: 100 times (numerical aperture NA = 0.9); laser power: 0.1-2%; grating line density: 1200 gr / mm; signal acquisition time: 30-60 seconds, single time; signal collection frequency range: 300 cm -1 -2000 cm -1 ), collect and save the Raman spectra formed by each.

[0039] Step 2 includes the following: Dataset division: The 30-category classification data of the publicly available common bacterial Raman spectroscopy dataset (HYPERLINK "https: / / www.nature.com / articles / s41467-019-12898-9) is divided into 70% training set, 10% validation set, and 20% test set.

[0040] Dataset processing: The Raman spectra are preprocessed by normalizing themselves.

[0041] Data augmentation method: During model training, Gaussian random noise is added to augment the data. The mean of the Gaussian noise is 0 and the variance is 0.1.

[0042] Classification model building: Design a 1D convolutional neural network architecture. The neural network consists of an initial convolutional block, multi-scale residual blocks, average pooling, and a fully connected layer. After each convolutional layer, there is a 1D BatchNorm normalization layer and a ReLU activation function layer. The multi-scale residual blocks reduce the signal resolution through convolutions with a stride of 2 and extract features at 6 scales. The feature extraction at each scale is composed of the network structure of 2 residual blocks. The input data format for building the model is (number of channels, spectral length). The fully connected layer finally outputs 30 channels, corresponding to the confidence levels of 30 categories.

[0043] Model training: The ADAMW optimizer is used, and cross-entropy is used as the loss function to train the weights of the model in the training set. The classification categories are encoded using the one-hot encoding scheme. The initial learning rate is set to 0.01, and the OneCycle learning rate adjustment scheme is used to adjust the learning rate during the optimization process.

[0044] Model testing: Then, the optimal hyperparameters such as the number of network layers and the learning rate are selected based on the results of the model on the validation set to determine the trained model. For the trained new convolutional neural network model, the test set is used to test the classification effect of the model with accuracy, precision, recall, F1 score, and confusion matrix.

[0045] Step 2.1: Group the Raman spectroscopy dataset collected in Step 1 as follows: bacteria (including Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, Staphylococcus aureus), fungi (including Candida albicans, Candida parapsilosis, and Saccharomyces cerevisiae). The proportion of each group of data is divided as 70% for the training set, 10% for the validation set, and 20% for the test set.

[0046] Step 2.2 Dataset processing: The Raman spectra in Step 2.1 are preprocessed by normalizing themselves.

[0047] Step 2.3 Data augmentation method: During model training, Gaussian random noise is added to augment the data. The mean of the Gaussian noise is 0 and the variance is 0.1.

[0048] Step 2.4 Classification model building: Modify the output channel number of the fully connected layer of the established classification model to 2, corresponding to the confidence levels of 2 categories of bacteria and fungi.

[0049] Step 2.5 Fine-tune the bacterial and fungal classification model: Use the weights of the trained model as the initial values of the model, and adopt the ADAMW optimizer, with cross-entropy as the loss function, to train the weights of the model in the training set. Use the one-hot encoding scheme to encode the classification categories, set the initial learning rate to 0.001, and adopt the OneCycle learning rate adjustment scheme to adjust the learning rate during the optimization process.

[0050] Step 2.6 Model testing: For the trained convolutional neural network model, use the test set to test the classification effect of the model with accuracy, precision, recall, F1 score, and confusion matrix.

[0051] Step 3 includes the following: Step 3.1 Group the Raman spectroscopy dataset collected in Step 1 as follows: Gram-negative bacteria (including Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, and Haemophilus influenzae); Gram-positive bacteria (including Streptococcus agalactiae, Streptococcus pneumoniae, and Staphylococcus aureus). The proportion of each group of data is divided into 70% for the training set, 10% for the validation set, and 20% for the test set.

[0052] Step 3.2 Dataset processing: Normalize each Raman spectrum in Step 2.1 by itself as preprocessing.

[0053] Step 3.3 Data augmentation method: During the model training process, use the method of adding Gaussian random noise to augment the data. The mean of the Gaussian noise is 0, and the variance is 0.1.

[0054] Step 3.4 Classification model modeling: Modify the number of output channels of the last fully connected layer of the established classification model to 2, corresponding to the confidence levels of the two categories of Gram-negative bacteria and Gram-positive bacteria.

[0055] Step 3.5 Fine-tune the Gram-negative and Gram-positive bacteria classification model: Use the weights of the trained model as the initial values of the model, and adopt the ADAMW optimizer, with cross-entropy as the loss function, to train the weights of the model in the training set. Use the one-hot encoding scheme to encode the classification categories, set the initial learning rate to 0.001, and adopt the OneCycle learning rate adjustment scheme to adjust the learning rate during the optimization process.

[0056] Step 3.6 Model testing: For the trained convolutional neural network model, use the test set to test the classification effect of the model with accuracy, precision, recall, F1 score, and confusion matrix.

[0057] Step 4 includes the following: Step 4.1 Group the Raman spectroscopy dataset collected in Step 1 as follows: Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, Staphylococcus aureus, Candida albicans, Candida parapsilosis, and Saccharomyces cerevisiae. The proportion of each group of data is divided into 70% for the training set, 10% for the validation set, and 20% for the test set.

[0058] Step 4.2 Dataset processing: Normalize each Raman spectrum in Step 2.1 using itself as preprocessing.

[0059] Step 4.3 Data augmentation method: During model training, add Gaussian random noise to augment the data. The mean of the Gaussian noise is 0, and the variance is 0.1.

[0060] Step 4.4 Classification model construction: Modify the number of output channels of the last fully connected layer of the established classification model to 12, corresponding to the confidence levels of the 12 categories of Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, Staphylococcus aureus, Candida albicans, Candida parapsilosis, and Saccharomyces cerevisiae.

[0061] Step 4.5 Fine-tune the classification model for common pathogenic bacteria in children: Use the weights of the trained model as the initial values of the model, and adopt the ADAMW optimizer and cross-entropy as the loss function to train the weights of the model in the training set. Encode the classification categories using the one-hot encoding scheme, set the initial learning rate to 0.001, and adopt the OneCycle learning rate adjustment scheme to adjust the learning rate during the optimization process.

[0062] Step 4.6 Model testing: For the trained convolutional neural network model, use the test set to test the classification effect of the model with accuracy, precision, recall, F1 score, and confusion matrix.

[0063] Step 5 includes the following content: Step 5.1 Group the Raman spectroscopy dataset collected in Step 1 as follows: Carbapenem-resistant and -sensitive Acinetobacter baumannii. The proportion of each group of data is divided into 70% for the training set, 10% for the validation set, and 20% for the test set; Step 5.2 Dataset processing: Normalize each Raman spectrum in Step 2.1 using itself as preprocessing; Step 5.3 Data augmentation method: During model training, add Gaussian random noise to augment the data. The mean of the Gaussian noise is 0, and the variance is 0.1.

[0064] Step 5.4 Classification model construction: Modify the number of output channels of the fully connected layer at the end of the established classification model to 2, corresponding to the confidence levels of the two categories of carbapenem-resistant and carbapenem-sensitive Acinetobacter baumannii.

[0065] Step 5.5 Fine-tuning the classification model for carbapenem-resistant Acinetobacter baumannii and carbapenem-sensitive Acinetobacter baumannii: Use the weights of the trained model as the initial values of the model, and adopt the ADAMW optimizer, with cross-entropy as the loss function. Train the weights of the model in the training set, use the one-hot encoding scheme to encode the classification categories, set the initial learning rate to 0.001, and adopt the OneCycle learning rate adjustment scheme to adjust the learning rate during the optimization process.

[0066] Step 5.6 Model testing: For the trained convolutional neural network model, use the test set to test the classification effect of the model with accuracy, precision, recall, F1 score, and confusion matrix.

[0067] Step 6 includes the following content: Evaluate the laboratory specificity and accuracy of the constructed typing model using the subsequently collected pathogenic bacteria and the blank control group. Evaluate the sensitivity, specificity, and accuracy of the constructed typing model using 100 clinical strains. Evaluate with MALDI-TOF MS identification as the standard, and verify the discrepant results with whole genome sequencing.

[0068] Example 1: As Figure 1 shown, the present invention discloses a method for rapid identification of common pathogenic bacteria and their drug resistance in children assisted by artificial intelligence based on Raman spectroscopy, including the following steps: Step 1: Establish a Raman spectroscopy database of common pathogenic bacteria in children Step 1 includes the following content: Step 1.1 Strain identification: After resuscitating the clinically isolated strains, culture them overnight on a culture medium. Select an appropriate amount of single colonies and smear them evenly on the target points of the target plate to form a thin layer. Cover with 1 μL of HCCA standard solvent and air dry at room temperature naturally. Place the target plate into a matrix-assisted laser desorption / ionization time-of-flight mass spectrometry (MALDI-TOF-MS) and perform bacterial identification according to the automatic identification process to confirm the strain species (the identification result score ≥ 2.0 and the consistency classification is A).

[0069] Step 1.2 Drug sensitivity identification: For the strains identified as Acinetobacter baumannii, use an automatic microbial identification and drug sensitivity analysis system to detect the minimum inhibitory concentration of the strains against antibacterial drugs. If Acinetobacter baumannii is resistant to any one of ertapenem, imipenem, and meropenem, it is called carbapenem-resistant Acinetobacter baumannii, and if it is sensitive to all three, it is carbapenem-sensitive Acinetobacter baumannii.

[0070] Step 1.3 Strain processing procedure: Process strains of Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, Staphylococcus aureus, Candida albicans, Candida parapsilosis, and Saccharomyces cerevisiae. Pick colonies and suspend them in a 1.5 mL centrifuge tube containing 500 μL of sterile deionized water. Repeat the washing 3 times (12,000 rpm / min, centrifuge for 3 min) until the turbidity reaches 0.5 McFarland units. Finally, reserve 200 μL for use.

[0071] Step 1.4 Raman spectroscopy database establishment: Prepare the pathogenic bacteria sample solutions mentioned in Step 1.3 respectively. Pipette 3 μL and drop it on an aluminum sheet to form a liquid droplet with a diameter of about 3 mm, and let it dry naturally. Conduct Raman spectroscopy detection (laser wavelength: 532 nm; microscope objective magnification: 100x (numerical aperture N.A. = 0.9); laser power: 1%; grating groove density: 1200 gr / mm; signal acquisition time: 30 - 60 seconds, single time; signal collection frequency band range: 300 cm -1 -2000 cm -1 ), and collect and save the Raman spectra formed respectively.

[0072] Step 2: Establish an identification model for common bacteria and fungi in children Step 2 includes the following contents: Step 2.1 Group the Raman spectroscopy dataset collected in Step 1 as follows: bacteria (including Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, Staphylococcus aureus), fungi (including Candida albicans, Candida parapsilosis, and Saccharomyces cerevisiae). Divide the proportion of each group of data into 70% for the training set, 10% for the validation set, and 20% for the test set.

[0073] Step 2.2 Dataset processing: Normalize each Raman spectrum collected in Step 2.1 with itself as preprocessing.

[0074] Step 2.3 Data augmentation method: During the model training process, use the method of adding Gaussian random noise to augment the data. The mean of the Gaussian noise is 0, and the variance is 0.1.

[0075] Step 2.4 Model structure design: Improve the ResNet18 residual convolutional neural network structure, replace the 2D convolution operation with a 1D convolution operation, replace the 2D BatchNorm operation with a 1D BatchNorm operation, and change the input data format to (number of channels, spectral length).

[0076] Step 2.5 Model Training: Use the ADAM optimizer, cross-entropy as the loss function, add the L2 norm of the weights as a regularization term, train the weights of the model on the training set, and then select the optimal hyperparameters such as the number of network layers and learning rate based on the results of the model on the validation set to determine the trained model; Step 2.6 Model Testing: For the trained novel convolutional neural network model, use the test set to test the classification effect of the model with accuracy, precision, recall, F1 score, and confusion matrix.

[0077] Step 3: Establish an identification model for common Gram-positive and Gram-negative bacteria in children Step 3 includes the following: Step 3.1 Group the Raman spectroscopy dataset collected in Step 1 as follows: Gram-negative bacteria (including Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, and Haemophilus influenzae); Gram-positive bacteria (including Streptococcus agalactiae, Streptococcus pneumoniae, and Staphylococcus aureus). The proportion of each group of data is divided into 70% for the training set, 10% for the validation set, and 20% for the test set.

[0078] Step 3.2 Dataset Processing: Normalize each Raman spectrum collected in Step 3.1 by itself as preprocessing.

[0079] Step 3.3 Data Augmentation Method: During the model training process, use the method of adding Gaussian random noise to augment the data. The mean of the Gaussian noise is 0 and the variance is 0.1.

[0080] Step 3.4 Model Structure Design: Improve the ResNet18 residual convolutional neural network structure, replace the 2D convolution operation with a 1D convolution operation, replace the 2D BatchNorm operation with a 1D BatchNorm operation, and change the input data format to (number of channels, spectral length).

[0081] Step 3.5 Model Training: Use the ADAM optimizer, cross-entropy as the loss function, add the L2 norm of the weights as a regularization term, train the weights of the model on the training set, and then select the optimal hyperparameters such as the number of network layers and learning rate based on the results of the model on the validation set to determine the trained model.

[0082] Step 3.6 Model Testing: For the trained novel convolutional neural network model, use the test set to test the classification effect of the model with accuracy, precision, recall, F1 score, and confusion matrix.

[0083] Step 4: Establish an identification model for 12 common pathogenic bacteria species in children Step 4 includes the following: Step 4.1 Group the Raman spectrum dataset collected in Step 1 as follows: Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, Staphylococcus aureus, Candida albicans, Candida parapsilosis, and Saccharomyces cerevisiae. The proportion of each group of data is divided into 70% for the training set, 10% for the validation set, and 20% for the test set.

[0084] Step 4.2 Dataset processing: Normalize each Raman spectrum collected in Step 1 using itself as preprocessing.

[0085] Step 4.3 Data augmentation method: During the model training process, add Gaussian random noise to augment the data. The mean of the Gaussian noise is 0 and the variance is 0.1.

[0086] Step 4.4 Model structure design: Improve the ResNet18 residual convolutional neural network structure by replacing the 2D convolution operation with a 1D convolution operation, the 2D BatchNorm operation with a 1D BatchNorm operation, and changing the input data format to (number of channels, spectral length).

[0087] Step 4.5 Model training: Use the ADAM optimizer, cross-entropy as the loss function, add the L2 norm of the weights as a regularization term, train the weights of the model in the training set, and then select the optimal hyperparameters such as the number of network layers and learning rate based on the results of the model on the validation set to determine the trained model. Step 4.6 Model testing: For the trained new convolutional neural network model, use the test set to test the classification effect of the model with accuracy, precision, recall, F1 score, and confusion matrix.

[0088] Step 5: Establish an identification model for carbapenem-resistant and carbapenem-sensitive Acinetobacter baumannii Step 5 includes the following: Step 5.1 Group the Raman spectrum dataset collected in Step 1 as follows: Carbapenem-resistant and carbapenem-sensitive Acinetobacter baumannii. The proportion of each group of data is divided into 70% for the training set, 10% for the validation set, and 20% for the test set.

[0089] Step 5.2 Dataset processing: Normalize each Raman spectrum collected in Step 1 using itself as preprocessing.

[0090] Step 5.3 Data augmentation method: During the model training process, add Gaussian random noise to augment the data. The mean of the Gaussian noise is 0 and the variance is 0.1.

[0091] Step 5.4 Model Structure Design: Improve the ResNet18 residual convolutional neural network structure, replace the 2D convolution operation with a 1D convolution operation, replace the 2D BatchNorm operation with a 1D BatchNorm operation, and change the input data format to (number of channels, spectral length).

[0092] Step 5.5 Model Training: Use the ADAM optimizer, cross-entropy as the loss function, add the second norm of the weights as a regularization term, train the weights of the model on the training set, and then select the optimal hyperparameters such as the number of network layers and learning rate based on the results of the model on the validation set to determine the trained model.

[0093] Step 5.6 Model Testing: For the trained novel convolutional neural network model, use the test set to test the classification effect of the model with accuracy, precision, recall, F1 score, and confusion matrix.

[0094] Step 6 includes the following content: Evaluate the laboratory specificity and accuracy of the constructed typing model using the subsequently collected pathogenic bacteria and blank control group. Evaluate the application of the constructed identification model using 100 clinical strains, and evaluate the sensitivity, specificity, and accuracy of the typing model. Evaluate using MALDI-TOF MS identification as the standard, and verify the discrepant results using whole-genome sequencing.

[0095] Input the data of the sample to be tested into the trained bacteria and fungi identification model, and perform deep learning analysis on the test group data to identify whether it is bacteria or fungi. When testing the test data, the model outputs the confidence levels of the test data being identified as bacteria or fungi respectively, and selects the category with the maximum confidence level to identify the sample as bacteria or fungi. The identification accuracy rate reaches 99.85%, as Figure 5 shown.

[0096] Input the data of the sample to be tested into the trained convolutional neural network model, and the model will output the confidence levels of the sample to be tested being identified as Gram-positive bacteria or Gram-negative bacteria, and select the category with the maximum confidence level to identify the sample as Gram-positive bacteria or Gram-negative bacteria. The identification accuracy rate reaches 94.70%, as Figure 7 shown.

[0097] Input the data of the sample to be tested into the trained identification models for 12 common pathogenic bacteria in children. The models will output the confidence levels of the sample to be tested being identified as Escherichia coli (Eco), Klebsiella pneumoniae (Kpn), Pseudomonas aeruginosa (Pa), Acinetobacter baumannii (Ab), Salmonella (Sal), Haemophilus influenzae (Hin), Streptococcus agalactiae (Sag), Streptococcus pneumoniae (Sp), Staphylococcus aureus (Sa), Candida albicans (Cal), Candida parapsilosis (Cpa), and Saccharomyces cerevisiae (Sce). Select the category with the highest confidence level to identify the sample as the corresponding strain of bacteria. The identification accuracy rate reaches 95.08%, as Figure 9 shown.

[0098] Input the data of the sample to be tested into the trained identification model for carbapenem-resistant and carbapenem-sensitive Acinetobacter baumannii. The model will output the confidence levels of the sample to be tested being identified as carbapenem-resistant Acinetobacter baumannii and carbapenem-sensitive Acinetobacter baumannii. Select the category with the highest confidence level to identify the sample as carbapenem-resistant Acinetobacter baumannii (AB IPM R) or carbapenem-sensitive Acinetobacter baumannii (AB IPM S). The identification accuracy rate reaches 96.20%, as Figure 12 shown.

[0099] Comparative Example 1: Table 1 shows the comparison of the correct rates of identifying 12 types of data of common pathogenic bacteria in children by different classification methods.

[0100] Table 1 Comparison of Correct Rates of Different Classification Methods It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not restrictive. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit and scope of the technical solutions of the present invention, and they should all be covered by the scope of the claims of the present invention.

Claims

1. A method for the rapid identification of common pathogenic bacteria and their drug resistance in children assisted by artificial intelligence based on Raman spectroscopy, characterized in that: Including the following steps: (1) Collect pathogenic bacteria samples: Isolate 12 common pathogenic bacteria from clinical specimens. The 12 common pathogenic bacteria include bacteria and fungi. The bacteria are Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, and Staphylococcus aureus; the fungi are Candida albicans, Candida parapsilosis, and Saccharomyces cerevisiae. Determine the 12 common pathogenic bacteria by matrix-assisted laser desorption / ionization time-of-flight mass spectrometry; (2) Divide the strains identified as Acinetobacter baumannii into carbapenem-resistant Acinetobacter baumannii and carbapenem-sensitive Acinetobacter baumannii according to their minimum inhibitory concentration against antibacterial drugs; (3) Establish a Raman spectroscopy database; (4) Establish a Raman spectroscopy deep learning classification model and pre-train the weights on the bacterial Raman dataset; Among them, design a 1D convolutional neural network architecture. The neural network consists of an initial convolutional block, multi-scale residual blocks, average pooling, and a fully connected layer. After the collected Raman spectra are normalized by themselves, the data is randomly divided into three groups: a training set, a test set, and a validation set, and input into the neural network model for processing; The method of adding Gaussian random noise is used to enhance the data during the model training process; (5) Establish deep learning classification models by modifying the output channels of the last fully connected layer of the model.

2. The method for the artificial intelligence-assisted rapid identification of common pathogenic bacteria in children and their drug resistance based on Raman spectroscopy according to claim 1, wherein: Step (3) of establishing a Raman spectroscopy database is to detect the 12 common pathogenic bacteria described in step (1) using a laser confocal Raman spectrometer, and collect and save the generated Raman spectroscopy database; among them, a laser confocal Raman spectrometer is selected, the laser wavelength is 532 nm, the laser power is set to 0.1-2%, the spectrometer grating is selected as 1200 g / mm, and the spectral range is 300-2000 cm -1 , and the single-spectrum acquisition time is 30-60 s.

3. The method for rapid identification of common pathogenic bacteria and their drug resistance in children based on Raman spectroscopy assisted by artificial intelligence according to claim 1 or 2, characterized in that: In step (3), suspend the pathogenic bacteria with sterile water to a turbidity of 0.5 McFarland units, drop it on the surface of an aluminum sheet, and detect it using a confocal Raman spectrometer.

4. The method for rapidly identifying common pathogenic bacteria in children and their drug resistance based on Raman spectroscopy and assisted by artificial intelligence according to claim 1, characterized in that: In step (4), the model is trained using the ADAMW optimizer. The model is trained on the bacterial Raman dataset, and then the optimal hyperparameters are selected based on the results of the model on the validation set to determine the trained model. For the trained model, use the test set to test the classification effect of the model with accuracy, precision, recall, F1 score, and confusion matrix.

5. The method for rapidly identifying common pathogenic bacteria in children and their drug resistance based on Raman spectroscopy assisted by artificial intelligence according to claim 1 or 2, characterized in that: In step (4), the proportion of each group of data is 70% for the training set, 10% for the validation set, and 20% for the test set; the mean of the Gaussian random noise is 0 and the variance is 0.

1.

6. The method for rapidly identifying common pathogenic bacteria in children and their drug resistance based on Raman spectroscopy with artificial intelligence assistance according to claim 1 or 2, characterized in that: In step (4), the 1D convolutional neural network architecture, where there is a 1D BatchNorm normalization layer and a ReLU activation function layer after each convolutional layer; the multi-scale residual blocks reduce the signal resolution through convolutions with stride = 2, extract features at 6 scales respectively, and the feature extraction at each scale consists of 2 residual block network structures; the input data format of the constructed model is: number of channels, spectral length.

7. The method for rapidly identifying common pathogenic bacteria in children and their drug resistance based on Raman spectroscopy assisted by artificial intelligence according to claim 1 or 2, characterized in that: Step (5) is to establish deep learning classification models for bacteria and fungi, Gram-positive and Gram-negative bacteria, the 12 common pathogenic bacteria, and carbapenem-resistant Acinetobacter baumannii and carbapenem-sensitive Acinetobacter baumannii by modifying the output channels of the last fully connected layer of the model; use the pre-trained model as the initial weight, fine-tune the corresponding model on the training set of the corresponding dataset, and use the test set to test the classification effect of the model with accuracy, precision, recall, F1 score, and confusion matrix.

8. The method for rapid identification of common pathogenic bacteria in children and their drug resistance based on Raman spectroscopy assisted by artificial intelligence according to claim 7, characterized in that: In step (5), the data of the sample to be tested is input into the trained convolutional neural network model, and the model will output the confidence levels of the sample to be tested being identified as bacteria or fungi. Select the category with the highest confidence level to identify the sample as bacteria or fungi; input the data of the sample to be tested into the trained convolutional neural network model, and the model will output the confidence levels of the sample to be tested being identified as Gram-positive bacteria or Gram-negative bacteria. Select the category with the highest confidence level to identify the sample as Gram-positive bacteria or Gram-negative bacteria.

9. The method for rapidly identifying common pathogenic bacteria in children and their drug resistance with the assistance of artificial intelligence based on Raman spectroscopy according to claim 7, characterized in that: In step (5), the data of the sample to be tested is input into the trained convolutional neural network model, and the model will output the confidence levels of the sample to be tested being identified as twelve strains, namely Escherichia coli, Klebsiella pneumoniae, Pseudomonas aeruginosa, Acinetobacter baumannii, Salmonella, Haemophilus influenzae, Streptococcus agalactiae, Streptococcus pneumoniae, Staphylococcus aureus, Candida albicans, Candida parapsilosis, and Saccharomyces cerevisiae. Select the category with the highest confidence level to identify the sample as the corresponding bacterial species.

10. The method for rapidly identifying common pathogenic bacteria in children and their drug resistance with the assistance of artificial intelligence based on Raman spectroscopy according to claim 7, wherein: In step (5), the data of the sample to be tested of Acinetobacter baumannii is input into the trained convolutional neural network model, and the model will output the confidence levels of the sample to be tested being identified as carbapenem-resistant Acinetobacter baumannii and carbapenem-susceptible Acinetobacter baumannii. Select the category with the highest confidence level to identify the sample to be tested as carbapenem-resistant Acinetobacter baumannii or carbapenem-susceptible Acinetobacter baumannii.

Citation Information

Patent Citations

  • Rapid identification method for carbapenem drug susceptibility, based on Raman spectra technology

    CN107586823A

  • Training method of food borne pathogenic bacteria Raman spectrum identification method established based on PCA-Stacking

    CN109781706A

  • Method for rapidly identifying positive and negative Gram bacteria

    CN110702664A

  • Method for rapidly identifying bacteria and fungi by utilizing Raman spectra

    CN111624190A

  • Novel coronavirus detection method and system based on enhanced Raman spectrum and neural network

    CN112798529A