Method for rapidly identifying acinetobacter complex
By combining surface-enhanced Raman spectroscopy with machine learning algorithms, the complexity and low accuracy issues of Acinetobacter complex identification were resolved, achieving fast and accurate identification results, with the support vector machine algorithm performing excellently.
Patent Information
- Application Number
- CN202410331790.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-22
- Publication Date
- 2025-09-23
AI Technical Summary
Existing technologies for the identification of Acinetobacter complex are complex, time-consuming, and have low accuracy, making it difficult to achieve rapid and accurate identification.
Combining surface-enhanced Raman spectroscopy (SERS) with machine learning algorithms, a rapid identification method for Acinetobacter complex was established through the preparation of Raman-enhanced substrates, acquisition of spectral signals, data quality control, characteristic peak analysis and clustering, and machine learning model training.
Rapid and accurate identification of Acinetobacter complex was achieved, and identification efficiency and accuracy were improved. In particular, the support vector machine algorithm achieved effective classification with an accuracy of 98.33% and a stability of 96.73%.
Smart Images

Figure CN120685401A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of spectral analysis technology, and specifically relates to using Raman spectroscopy to quickly analyze different Acinetobacter complex groups, qualitatively analyzing the distribution of spectral characteristic peaks, and comparing the accuracy of different machine learning models in identifying Acinetobacter complex groups. Background Art
[0002] The Acinetobacter complex is a group of four Acinetobacter species, including A. baumannii, A. pituitii, A. nosocomialis, and A. calciticus. The first three are common pathogens of community-acquired and nosocomial infections. A. baumannii, in particular, is a ubiquitous opportunistic pathogen resistant to multiple antibiotics, causing severe and difficult-to-treat infections worldwide, with infection rates increasing annually. Furthermore, A. baumannii is the most common source of respiratory-associated pneumonia, accounting for 15% of nosocomial infections in critical care hospital wards and having the highest mortality rate. Therefore, rapid and accurate identification of the Acinetobacter complex is of clinical value. Common methods for identifying Acinetobacter include, but are not limited to, 16S rRNA sequencing, specific gene sequencing, whole genome sequencing (WGS), polymerase chain reaction (PCR), pulsed-field gel electrophoresis (PFGE), and DNA arrays. However, these methods are subject to various limitations, including complex procedures, long time-to-accuracy (TAT), low accuracy, and high cost, necessitating the development of new diagnostic approaches. In recent years, surface-enhanced Raman spectroscopy (SERS) combined with machine learning (ML) algorithms has enabled rapid and accurate identification of a variety of clinical pathogens, garnering extensive research.
[0003] Existing techniques for identifying Baumanii typing are mostly based on biological techniques. For example, patent application publication number CN108842006A discloses a method for rapid identification of Acinetobacter baumannii in serum. This method uses qPCR to detect changes in the number of nucleic acid copies of progeny phage produced after phage infection of pathogenic bacteria, enabling rapid identification of the pathogenic baumannii. This method, which eliminates the need for complex sequencing techniques and biological background, simply combines collected Raman spectra with artificial intelligence analysis methods to achieve rapid identification of the Acinetobacter complex. Summary of the Invention
[0004] The purpose of the present invention is to address the deficiencies of the above-mentioned prior art and provide a method for rapid identification of Acinetobacter complex. By combining surface light-enhanced Raman spectroscopy with machine learning algorithms, the problems of cumbersome detection process and long time consumption are solved.
[0005] The technical solutions of the present invention are as follows:
[0006] A method for rapid identification of Acinetobacter complex comprising the following steps:
[0007] (1) Identification of bacterial species: All strains were confirmed by MALDI-TOF mass spectrometry;
[0008] (2) Bacterial collection: isolation and culture of Acinetobacter complex strains;
[0009] (3) Substrate preparation: sodium citrate solution is added to a silver nitrate solution that has been heated to boiling, and the mixture is stirred and reacted. The resulting reaction solution is centrifuged, and the resulting precipitate is resuspended in deionized water to obtain a Raman enhancement substrate with negatively charged nanosilver particles;
[0010] (4) Sample preparation: The three different complex group strains isolated and cultured in step (2) were inoculated into phosphate buffer solution, and then mixed with the Raman enhancement substrate prepared in step (3) to obtain three surface enhanced Raman spectroscopy (SERS) test samples of the Acinetobacter complex group;
[0011] (5) SERS detection: The SERS test sample of the Acinetobacter complex obtained in step (4) is subjected to multiple, multi-site SERS spectral signal acquisition to obtain spectral fingerprints of the corresponding Acinetobacter complex samples, thereby constructing a SERS fingerprint database of the Acinetobacter complex;
[0012] (6) Data quality control: The cosmic peak generated during sampling was removed from the spectral fingerprint of the Acinetobacter complex sample obtained in step (5), and the average Raman spectrum and standard deviation of each Acinetobacter complex sample were calculated to check the data distribution and repeatability;
[0013] (7) Characteristic peak analysis: Deconstruct the spectral characteristic peaks of the Acinetobacter complex sample through spectral deconvolution analysis, and find the biological significance of each characteristic peak;
[0014] (8) Data clustering: The Raman spectral data after quality control in step (6) were clustered using the unsupervised learning algorithms principal component analysis (PCA) and orthogonal partial least squares discriminant analysis (OPLS-DA). The differences between different Acinetobacter strains were clustered into different clusters in the classification quadrant diagram of the Raman vector. The sample was determined to belong to the cluster based on the distribution of different clusters, and the results of the cluster analysis were evaluated using the three indicators of R2X, R2Y and Q2;
[0015] (9) Data identification: The SERS spectral data after quality control in step (6) are analyzed using different artificial intelligence algorithm methods, divided into a training set, a validation set, and a test set by uniform random sampling, and a classifier is trained. K-fold cross validation is used for verification, where K is any integer from 1 to 10. The sample data are classified and labeled by the trained classifier and stored in a database;
[0016] (10) Data evaluation: Select test set samples to test the predictive ability, use precision, recall rate, ROC curve and confusion matrix to modify the final judgment model, evaluate the performance of different machine learning algorithms, and select the best judgment model, including:
[0017]
[0018]
[0019]
[0020] Among them: Precision represents accuracy, Recall represents recall rate, AUC represents area under the ROC curve, which is regarded as a performance indicator; TP, FP, TN and FN represent the number of true positives, false positives, true negatives and false negatives respectively, and the confusion matrix is constructed by these four cases.
[0021] In the present invention, in step (1), the process of bacterial species identification is as follows:
[0022] (a) Three strains of Acinetobacter were collected: 10–30 strains of Acinetobacter baumannii, 10–30 strains of Acinetobacter pitei, and 10–30 strains of Acinetobacter hospitalis, and confirmed by MALDI-TOF mass spectrometry;
[0023] (b) The mass spectrometry parameters were set as follows: laser frequency of 75 Hz, number of shots of 100, and mass-to-charge ratio of 2000–20000 Da in RUO mode.
[0024] In a preferred embodiment, the process of bacterial species identification is as follows:
[0025] (a) Three strains of Acinetobacter were collected: 20 strains of Acinetobacter baumannii, 20 strains of Acinetobacter pitei, and 20 strains of Acinetobacter hospitalis, and confirmed by MALDI-TOF mass spectrometry;
[0026] (b) The mass spectrometry parameters were set as follows: laser frequency of 75 Hz, number of shots of 100, and mass-to-charge ratio of 2000–20000 Da in RUO mode.
[0027] For the present invention, in step (2), the isolation and cultivation process is as follows: three Acinetobacter complex strains are cultured in Luria-Bertani liquid medium to the exponential growth phase, and after centrifugation, the supernatant and strain pellets are obtained. The obtained strain pellets are resuspended in deionized water, and the concentration of the strains is determined by plate count method on a blood agar plate cultured at 37°C for 24 hours. One colony is randomly selected for each strain and mixed evenly with 50-150 μL of sterile ddH2O to prepare a bacterial solution for use.
[0028] In a preferred embodiment, the isolation and culture process is as follows: three Acinetobacter complex strains are cultured in Luria-Bertani liquid medium to the exponential growth phase. After centrifugation, the supernatant and strain pellets are obtained. The obtained strain pellets are resuspended in deionized water. The concentration of the strains is determined by plate count method on blood agar plates cultured at 37°C for 24 hours. One colony is randomly selected for each strain and mixed evenly with 100 μL of sterile ddH2O to prepare a bacterial solution for use.
[0029] When preparing a Raman-enhanced substrate, silver nitrate is dissolved in water and stirred and heated to boiling, wherein the concentration of the silver nitrate in the silver nitrate solution is 0.5 to 1.5 mol / L, preferably 1 mol / L. The reducing agent is sodium citrate, and the sodium citrate solution is prepared with a concentration of 0.5 to 1.5 wt%, preferably 1 wt%.
[0030] In a preferred embodiment, the preparation method of the Raman enhanced substrate is as follows:
[0031] (a) A 1 mmol / L silver nitrate solution was heated to boiling, and 8 mL of a 1 wt% sodium citrate solution was added while stirring. The mixture was stirred at 450-850 rpm for 20-60 min to obtain negatively charged silver nanoparticles.
[0032] (b) The reaction solution obtained in step (a) was centrifuged at a speed of 6000-9000 r / min for 4-9 minutes, the supernatant was discarded, and the resulting precipitate was resuspended in deionized water to obtain a Raman enhancement substrate with negatively charged nanosilver particles.
[0033] In a more preferred embodiment, the preparation method of the Raman enhanced substrate is as follows:
[0034] (a) A 1 mmol / L silver nitrate solution was heated to boiling, and 8 mL of a 1 wt% sodium citrate solution was added while stirring. The mixture was stirred at 650 rpm for 40 min to obtain negatively charged silver nanoparticles.
[0035] (b) 1 mL of the reaction solution obtained in step (1) was centrifuged at 7000 rpm for 7 min. The supernatant was discarded and the resulting precipitate was resuspended in 100 μL of deionized water to obtain a negatively charged silver nanoparticle solution, which was the Raman enhancement substrate.
[0036] Regarding the negatively charged silver nanoparticles mentioned above, the diameter of the silver nanoparticles is less than 11 nm.
[0037] For the present invention, in step (4), the three Acinetobacter strains isolated and cultured in step (2) are respectively inoculated into a phosphate buffer solution, and sequentially mixed with the surface enhanced Raman substrate solution prepared in step (3) in equal amounts according to a volume of 3 to 6 μL, preferably 5 μL. A volume of 2 to 5 μL of the mixed solution, preferably 2.5 μL, is taken and dropped on a clean silicon wafer in appropriate amounts multiple times to form circular spots. After natural drying in a safety cabinet, the solution is detected to obtain a surface enhanced Raman spectrum (SERS) of the Acinetobacter complex strain sample, thereby constructing a SERS fingerprint database of the Acinetobacter complex strain.
[0038] In the present invention, in step (5), SERS detection is suitable for collecting surface enhanced Raman signals of samples, and signals are collected from the SERS sample of the Acinetobacter complex strain obtained in step (4).
[0039] Preferably, during SERS detection, 600 to 900 sites are randomly selected for Raman spectrum detection for each sample to be tested. To improve data balance, 800 sites are further preferably selected, that is, a total of 800 SERS spectra are obtained for each sample.
[0040] Preferably, SERS detection uses inVia™ micro-Raman spectroscopy to collect data, and the sampling parameters are: Raman spectroscopy excitation wavelength is 785 nm, detector type: high-sensitivity CCD array, exposure time is 10 s, and wavelength scanning range is 400-2800 cm-1.
[0041] Preferably, before spectrum acquisition, the SERS detection is performed using the Raman peak of the silicon wafer at 520 cm-1 as a reference peak for wave number calibration, and the dark current is deducted within the same integration time.
[0042] For the present invention, in step (6), the specific steps of data quality control are as follows:
[0043] (a) removing cosmic spikes from the SERS signal collected in step (5) by visual inspection;
[0044] (b) performing average spectrum calculation on the SERS spectra of the different bacillus complex groups after removing the sharp peak in step (a), calculating the average Raman shift of all SERS spectra of the same type at a certain point, selecting the number of Raman shifts to be 600 to 2000, preferably 1161;
[0045] (c) Calculate the standard deviation of the Raman shift at a certain point for all SERS spectra of the same type after removing the cosmic spike in step (a). Then, complete the standard deviation of the 1161 Raman shifts screened in step (b). Observe the width of the standard error bands of the three Acinetobacter species to measure the quality of the collected SERS spectra.
[0046] In the present invention, in step (7), the spectral characteristic peak analysis deconstructs a more detailed characteristic peak composition by the spectral deconvolution method, and the biological significance of each characteristic peak is found by searching literature related to microorganisms.
[0047] In the present invention, in step (8), data clustering is used to determine whether the sample is in a specified interval and belongs to a specified category. It can be selected as needed, specifically unsupervised learning principal component analysis, density cluster analysis, K-means analysis, and supervised learning algorithms linear discriminant analysis, partial least squares, orthogonal partial least squares discriminant analysis.
[0048] In the present invention, in step (8), cluster analysis is performed on the Raman spectrum data after quality control in step (6). The specific process is as follows:
[0049] (a) The SERS data of three types of Acinetobacter after quality control were input into the PCA algorithm. The spectral data were subjected to dimensionality reduction clustering visualization. The principal components PCA1 to PCA7 after dimensionality reduction were selected as the principal components. The preferred number of principal components was two, namely PCA1 and PCA2. The clustering results of the Raman spectral principal component sample points in the classification quadrant were observed.
[0050] (b) The preprocessed SERS data were input into the OPLS-DA algorithm, and the clustering results were evaluated using the three indicators R2X, R2Y, and Q2. The larger the indicator, the better the model performance, and the difference between R2X and Q2 did not exceed 0.3.
[0051] For the present invention, in step (8), data clustering is only used as a preliminary judgment of the data to determine the quality of the data and the degree of distinguishability. In the present invention, data clustering can distinguish the three types of Acinetobacter complex to a certain extent, but the results overlap, indicating that they cannot be completely distinguished. Moreover, data clustering cannot provide an identification model. For Raman spectral data that are not recorded in the database, it is necessary to rewrite the training and divide them again. The purposes of subsequent steps (9) and (10) include the following two points: one is to be able to better distinguish the Acinetobacter complex; the other is to be able to train an optimal model for identifying the Acinetobacter complex. When Raman spectral data that are not recorded in the database enter, the model does not need to be retrained and can directly identify the Raman spectral data. Data clustering can be understood as a preliminary experiment of steps (9) and (10) and is a transition.
[0052] For the present invention, in step (9), the Acinetobacter complex is identified as a 3-classification model, and the artificial intelligence algorithms used in data identification are eight classic machine learning algorithms: support vector machine, extreme gradient boosting machine, random forest, guided clustering, linear discriminant analysis, decision tree, adaptive boosting and quadratic discriminant analysis.
[0053] For the present invention, in step (9), the machine learning process includes computer storage, a computer processor, and a computer program stored in the computer storage and executable on the computer processor, the computer storage storing the model optimization parameter range, and the computer processor performing the following steps when executing the computer program:
[0054] (a) The SERS fingerprints of three types of Acinetobacter after quality control were divided into data sets;
[0055] (b) pre-set different machine learning training parameter ranges and select the optimal parameter combination;
[0056] (c) Using the machine learning model in data identification to classify the SERS fingerprints of Acinetobacter strains, the classification results of different spectral data were obtained;
[0057] (d) Use data evaluation to evaluate the performance of different machine learning algorithms and select the best judgment model.
[0058] Step (a) involves using uniform random sampling to group the pre-built sample database into a training set, a validation set, and a test set in a 6:2:2 ratio, where the validation set data is derived from the training set. The training and validation set data are used to train and validate the model, while the test set data is completely independent of these two sets and is used only to test the model's performance on unknown data.
[0059] Step (b) includes: presetting parameter ranges for different machine learning algorithms, using a grid search method to traverse and calculate the model scores of each machine learning algorithm in different parameter combinations, and selecting the best model parameter combination for final model performance comparison;
[0060] Step (c) includes: data classification using a data analysis tool to automatically analyze the Raman signal data of all samples, training multiple classifiers, using 5-fold cross-validation for detection, and using the trained classifiers to classify and label the three Acinetobacter species and store them in a database;
[0061] Step (d) includes: selecting some of the remaining samples for predictive ability testing, using accuracy, precision, recall, ROC curve and confusion matrix to modify the final judgment model, evaluating the performance of different machine learning algorithms, and selecting the best judgment model.
[0062] The technical solution of the present invention has the following advantages:
[0063] A method for identifying and classifying Acinetobacter complex strains based on surface-enhanced Raman spectroscopy combined with machine learning algorithms is provided, which provides a fast and effective analysis approach for the application of Raman spectrometers and reflects the application advantages of Raman spectrometers.
[0064] A SERS fingerprint library of Acinetobacter complex was established through machine learning algorithms, and the newly collected Raman spectra were directly imported into the pre-trained model file for judgment. The comparison and combination of different machine learning methods and evaluation indicators can better ensure the accuracy of the identification results.
[0065] High accuracy: When distinguishing three different Acinetobacter species, the support vector machine algorithm achieved the best performance among all evaluation indicators (accuracy = 98.33%) and good stability (5-fold cross validation = 96.73%), indicating that this method is stable and reliable and can more effectively extract and analyze spectral data. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 The average Raman spectra and Raman spectrum deconvolution images of the three Acinetobacter complexes in the present invention;
[0067] Figure 2 This is a clustering result diagram of the three Acinetobacter complexes in the present invention using PCA and OPLS-DA algorithms;
[0068] Figure 3 The ROC curves and confusion matrix results of different machine learning algorithms in the identification task of three Acinetobacter complex groups in the present invention are shown;
[0069] Figure 4 This is a diagram showing the biological significance of the unique Raman characteristic peaks of the three Acinetobacter species in the present invention;
[0070] Figure 5 This is a performance score diagram of different machine learning algorithms in the present invention. DETAILED DESCRIPTION
[0071] Raman spectroscopy is based on the principle of inelastic scattering. That is, when incident light from a laser light source illuminates a substance, it is scattered by the molecules of the substance. A very small portion of the scattered light has a frequency different from the incident light. The change in the scattered light frequency depends on the structural characteristics of the illuminated substance. Different substances produce scattered light of specific frequencies under the same laser irradiation. Therefore, Raman spectroscopy can be used to achieve fast, simple, repeatable and non-destructive detection of material composition.
[0072] Artificial intelligence technology provides an efficient and accurate implementation for Raman spectroscopy-based material composition detection. Existing Raman spectroscopy machine learning algorithms are oriented toward specific substances to be tested. They transform the material identification problem of Raman spectroscopy into a machine learning classification problem. Machine learning models are trained based on standard Raman spectra of known substances, and the trained models are used to accurately identify test samples.
[0073] This specification and claims do not use differences in names as a way to distinguish components, but use differences in components' functions as the criteria for distinction. As mentioned in the description and claims throughout the text, "including" is an open-ended term and should be interpreted as "including but not limited to". "Approximately" means that within an acceptable error range, those skilled in the art can solve the technical problem within a certain error range and basically achieve the technical effect. In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the embodiments so that those skilled in the art can implement it with reference to the text of the specification. The equipment or raw materials used in the embodiments can all be obtained from the market.
[0074] instrument
[0075] Raman spectrometer (inVia™), MALDI-TOF mass spectrometer (Vitek MS), balance (ME104), water purifier (Arium mini), centrifuge (Centrifuge 5424R), refrigerator (BCD-406WDPD), magnetic stirrer (RCTbasic), clean bench (SW-CJ-2F), ultra-low temperature refrigerator (994), low-temperature incubator (PR505750R-CN), pH meter (FE28). The models and manufacturers of the instruments used in the implementation case are detailed in Table 1 below.
[0076] Table 1 Experimental instrument information
[0077]
[0078] Drugs and reagents
[0079] Silver nitrate (Sinopharm Chemical Reagent), sodium citrate (Sinopharm Chemical Reagent), sodium chloride (Sinopharm Chemical Reagent), and Luria-Bertani liquid medium (JingBio, Beijing).
[0080] Specific examples Rapid identification of Acinetobacter complex strains using the method of the present invention
[0081] 1. Method
[0082] (1) Identification of bacterial species:
[0083] The specific process is as follows:
[0084] (a) Three strains of Acinetobacter were collected: 20 strains of Acinetobacter baumannii, 20 strains of Acinetobacter pitei, and 20 strains of Acinetobacter hospitalis, and confirmed by MALDI-TOF mass spectrometry;
[0085] (b) The mass spectrometry parameters were set as follows: laser frequency of 75 Hz, number of shots of 100, and mass-to-charge ratio of 2000–20000 Da in RUO mode.
[0086] (2) Bacterial collection: All Acinetobacter strains were isolated and cultured from clinical samples. The bacterial collection process was as follows: three Acinetobacter complex strains were cultured in Luria-Bertani liquid medium until the exponential growth phase. After centrifugation, the supernatant and strain pellets were obtained. The obtained strain pellets were resuspended in deionized water, and the strain concentrations were determined by plate count method on blood agar plates cultured at 37°C for 24 h.
[0087] (3) Substrate preparation:
[0088] The preparation process of the Raman enhanced substrate is as follows:
[0089] (a) 33.72 mg of silver nitrate was weighed and dissolved in 200 mL of deionized water to obtain a 1 mol / L silver nitrate solution. The solution was heated to boiling using a magnetic stirrer. 8 mL of a 1 wt % sodium citrate solution was then added at once while stirring. The mixture was stirred at 650 rpm for 40 min to obtain a negatively charged silver nanoparticle reaction solution having a diameter of <11 nm.
[0090] (b) 1 mL of the silver nanoparticle reaction solution obtained in step (a) was centrifuged at 7000 rpm for 7 min. The supernatant was discarded and the resulting precipitate was resuspended in 100 μL of deionized water to obtain a Raman enhancement substrate solution of negatively charged silver nanoparticles. The resulting solution was uniformly milky white and stored in the dark at room temperature until use.
[0091] (4) Sample preparation: The three different complex group strains isolated and cultured in step (2) were inoculated into phosphate buffer solution, and then mixed with the surface-enhanced Raman substrate solution prepared in step (3) in equal amounts of 5 μL. 2.5 μL of the mixed solution was dropped onto a clean silicon wafer several times to form circular spots, and then naturally dried in a safety cabinet before testing.
[0092] (5) SERS detection: The air-dried SERS sample in step (4) was taken for measurement. During the measurement, multiple and multi-site Raman spectroscopy sampling was adopted. 800 SERS fingerprint spectra were obtained for each sample, thereby constructing a SERS fingerprint database of the Acinetobacter complex.
[0093] The parameters for Raman spectrum sampling during the measurement process are as follows: the excitation wavelength of the Raman spectrum is 785 nm, the detector type is a high-sensitivity CCD array, the exposure time is 10 s, the wavelength scanning range is 400-2800 cm-1, the output data type is txt, the laser intensity is 2, the spectrum acquisition time is 10 s, and before spectrum acquisition, the Raman peak of the silicon wafer at 520 cm-1 is used as the reference peak for wavenumber calibration, and the dark current is subtracted at the same integration time.
[0094] (6) Data quality control:
[0095] The preparation process is as follows:
[0096] (a) removing cosmic spikes from the SERS signal collected in step (5) by visual inspection;
[0097] (b) performing spectrum average calculation on the SERS spectra of the different bacillus complexes after removing the sharp peak in step (a), calculating the average Raman shift of all SERS spectra of the same type at a certain point, and selecting the number of Raman shifts as 1161;
[0098] (c) Calculate the standard deviation of the Raman shift at a certain point for all SERS spectra of the same type after removing the cosmic peak in step (a), and then complete the standard deviation value of the 1161 Raman shifts screened in step (b).
[0099] (7) Characteristic peak analysis: Spectral characteristic peak analysis uses the method of spectral deconvolution to deconstruct more detailed characteristic peak composition, and searches for the biological significance of each characteristic peak by searching literature related to microorganisms.
[0100] (8) Data clustering:
[0101] The SERS spectrum clustering analysis process of three Acinetobacter samples is as follows:
[0102] (a) Inputting the SERS spectrum of Acinetobacter after data quality control in step (5) into the PCA unsupervised learning algorithm model, performing dimensionality reduction clustering visualization operations on the spectral data, selecting the principal components PCA1 and PCA2 to determine the classification quadrants, and observing the clustering of spectral sample points in the classification quadrants;
[0103] (b) The quality-controlled SERS spectra of Acinetobacter were input into the OPLS-DA supervised learning algorithm model. By comparing the distribution differences of different Acinetobacter SERS spectral sample points, a classification map of different Acinetobacter Raman vectors was formed. The clustering results were evaluated using the three indicators R2X, R2Y, and Q2.
[0104] (9) Data Identification: The Raman spectral data after quality control in step (6) were automatically analyzed using different machine learning algorithms. The SERS spectral datasets of different Acinetobacter complexes were divided into training, validation, and test sets using uniform random sampling. A classifier was trained and tested using 5-fold cross-validation. The trained classifiers were used to assign and label the sample data and store them in a database. All parameters were optimized before training. The machine learning parameter settings are shown in Table 2 below.
[0105] Table 2 Machine learning parameter settings for the three-classification of Acinetobacter complex
[0106]
[0107] (10) Data evaluation: Select test set samples to test the predictive ability, use precision, recall rate, ROC curve and confusion matrix to modify the final judgment model, evaluate the performance of different machine learning algorithms, and select the best judgment model, including:
[0108]
[0109]
[0110]
[0111] Among them: Precision represents accuracy, Recall represents recall rate, AUC represents area under the ROC curve, which is regarded as a performance indicator; TP, FP, TN and FN represent the number of true positives, false positives, true negatives and false negatives respectively, and the confusion matrix is constructed by these four cases.
[0112] 2. Results
[0113] Figure 1 The average Raman spectra and deconvolution of the Raman spectra of the three Acinetobacter species in the present invention are shown. The shaded area in the figure represents the standard error band of the SERS spectrum of each type of Acinetobacter. The width of the error band indicates the quality of the data. The deconvolution spectrum magnifies the difference between the three average spectra. It can be seen that the difference between 100-1800 cm -1 Within the range, multiple differences are fitted out.
[0114] Figure 2Figure (A) PCA and Figure (B) OPLS-DA algorithms are used in the present invention to cluster the three types of Acinetobacter spectral data in the form of scatter plots. It was found that the PCA algorithm was unable to cluster Acinetobacter, and a large amount of overlap occurred between the sample points; while for the OPLS-DA algorithm, different Acinetobacter were clustered into different clusters with partial overlap, indicating that the OPLS-DA algorithm can distinguish different Acinetobacter to a certain extent. The OPLS-DA clustering results of the three types of Acinetobacter are circled with dotted circles, and the results of the cluster analysis are evaluated using the three indicators R2X, R2Y and Q2. R2X = 0.994, R2Y = 0.702, Q2 = 0.686, indicating that the different Acinetobacter data are separable to a certain extent using the OPLS-DA algorithm.
[0115] Figure 3 Figure 2 is the ROC curve and confusion matrix diagram of eight machine learning algorithms in the present invention. Figure (A) shows the ROC curve results and area under the curve AUC values of different algorithms. According to the size of the AUC value, we found that the performance of SVM is the best. To this end, we used the confusion matrix to show in detail the detailed classification of each Acinetobacter by SVM. The result shows that the SVM model can accurately identify Acinetobacter baumannii. For the recognition ability of Acinetobacter nosocomial, 2% of the spectral data are mistakenly identified as Acinetobacter baumannii and Acinetobacter pituitarius, and for the recognition ability of Acinetobacter pituitarius, 1% of the spectral data are mistakenly identified as Acinetobacter baumannii and Acinetobacter nosocomial. The average identification accuracy of the whole model on the test set is 98%.
[0116] Figure 4 The table below shows the molecular vibrations and biological significance of the three Acinetobacter species at their respective unique characteristic peaks. These differences can be preliminarily considered as the basis for identifying different Acinetobacter species.
[0117] Figure 5 The performance evaluation results of the eight machine learning algorithms used in this study using five evaluation metrics are shown. The results show that SVM performs best, with the highest accuracy (Accuracy = 98.33%) and the best stability (5-Fold = 96.73%), demonstrating that SVM is an effective method for distinguishing different Acinetobacter species from SERS spectra. The remaining algorithms, except QDA, also achieved accuracy exceeding 80%.
[0118] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that it is still possible to modify the technical solutions described in the aforementioned embodiments, or to make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for rapid identification of Acinetobacter complex, characterized in that: It includes the following steps: (1) Identification of bacterial species: All strains were confirmed by MALDI-TOF mass spectrometry; (2) Bacterial collection: isolation and culture of Acinetobacter complex strains; (3) Substrate preparation: sodium citrate solution is added to a silver nitrate solution that has been heated to boiling, and the mixture is stirred and reacted. The resulting reaction solution is centrifuged, and the resulting precipitate is resuspended in deionized water to obtain a Raman enhancement substrate with negatively charged nanosilver particles; (4) Sample preparation: The three different complex group strains isolated and cultured in step (2) were inoculated into phosphate buffer solution, and then mixed with the Raman enhancement substrate prepared in step (3) to obtain three surface enhanced Raman spectroscopy (SERS) test samples of the Acinetobacter complex group; (5) SERS detection: The SERS test sample of the Acinetobacter complex obtained in step (4) is subjected to multiple, multi-site SERS spectral signal acquisition to obtain spectral fingerprints of the corresponding Acinetobacter complex samples, thereby constructing a SERS fingerprint database of the Acinetobacter complex; (6) Data quality control: The cosmic peak generated during sampling was removed from the spectral fingerprint of the Acinetobacter complex sample obtained in step (5), and the average Raman spectrum and standard deviation of each Acinetobacter complex sample were calculated to check the data distribution and repeatability; (7) Characteristic peak analysis: Deconstruct the spectral characteristic peaks of the Acinetobacter complex sample through spectral deconvolution analysis, and find the biological significance of each characteristic peak; (8) Data clustering: The Raman spectral data after quality control in step (6) were clustered using the unsupervised learning algorithms principal component analysis (PCA) and orthogonal partial least squares discriminant analysis (OPLS-DA). The differences between different Acinetobacter strains were clustered into different clusters in the classification quadrant diagram of the Raman vector. The sample was determined to belong to the cluster based on the distribution of different clusters, and the results of the cluster analysis were evaluated using the three indicators of R2X, R2Y and Q2; (9) Data identification: The SERS spectral data after quality control in step (6) are analyzed using different artificial intelligence algorithm methods, divided into a training set, a validation set, and a test set by uniform random sampling, and a classifier is trained. K-fold cross validation is used for verification, where K is any integer from 1 to 10. The sample data are classified and labeled by the trained classifier and stored in a database; (10) Data evaluation: Select test set samples to test the predictive ability, use precision, recall rate, ROC curve and confusion matrix to modify the final judgment model, evaluate the performance of different machine learning algorithms, and select the best judgment model, including: Among them: Precision represents accuracy, Recall represents recall rate, AUC represents area under the ROC curve, which is regarded as a performance indicator; TP, FP, TN and FN represent the number of true positives, false positives, true negatives and false negatives respectively, and the confusion matrix is constructed by these four cases.
2. The method for rapid identification of Acinetobacter complex according to claim 1, characterized in that In step (1), the process of bacterial species identification is as follows: (a) Three strains of Acinetobacter were collected: 10–30 strains of Acinetobacter baumannii, 10–30 strains of Acinetobacter pitei, and 10–30 strains of Acinetobacter hospitalis, and confirmed by MALDI-TOF mass spectrometry; (b) The mass spectrometry parameters were set as follows: laser frequency of 75 Hz, number of shots of 100, and mass-to-charge ratio of 2000–20000 Da in RUO mode.
3. The method for rapid identification of Acinetobacter complex according to claim 1, wherein In step (2), the isolation and culture process is as follows: three strains of Acinetobacter complex are cultured in Luria-Bertani liquid medium until the exponential growth phase, and after centrifugation, the supernatant and strain particles are obtained, and the obtained strain particles are resuspended in deionized water. The concentration of the strains is determined by plate count method on a blood agar plate cultured at 37°C for 24 hours. One colony is randomly selected for each strain and mixed evenly with 50-150 μL of sterile ddH2O to prepare a bacterial solution for use.
4. The method for rapid identification of Acinetobacter complex according to claim 1, wherein In step (3), the preparation method of the Raman enhanced substrate is as follows: (a) A 1 mmol / L silver nitrate solution was heated to boiling, and 8 mL of a 1 wt% sodium citrate solution was added while stirring. The mixture was stirred at 450-850 rpm for 20-60 min to obtain negatively charged silver nanoparticles. (b) The reaction solution obtained in step (a) was centrifuged at a speed of 6000-9000 r / min for 4-9 minutes, the supernatant was discarded, and the resulting precipitate was resuspended in deionized water to obtain a Raman enhancement substrate with negatively charged nanosilver particles.
5. The method for rapid identification of Acinetobacter complex according to claim 1, characterized in that In step (4), 3 to 6 μL of the Raman enhancement substrate solution with negatively charged silver nanoparticles prepared in step (3) is mixed with an equal amount of the Acinetobacter strain sample cultured in step (2). 2 to 5 μL of the mixed sample is titrated on a clean silicon wafer to form circular spots, and the sample is naturally dried in a safety cabinet to obtain a SERS sample to be tested.
6. The method for rapid identification of Acinetobacter complex according to claim 1, wherein In step (5), 600 to 900 sites are selected from each SERS sample to be tested obtained in step (4) by multi-site scanning; the parameters of Raman spectrum sampling are as follows: the excitation wavelength of Raman spectrum is 785 nm, the detector type is a high-sensitivity CCD array, the exposure time is 10 s, and the wavelength scanning range is 400 to 2800 cm -1 .
7. The method for rapid identification of Acinetobacter complex according to claim 1, wherein In step (6), the specific steps of data quality control are as follows: (a) removing cosmic spikes from the SERS signal collected in step (5) by visual inspection; (b) performing spectrum average calculation on the SERS spectra of the different bacillus complexes after removing the sharp peak in step (a), and calculating the average Raman shift of all SERS spectra of the same type at a certain point, with the number of Raman shifts selected being 600 to 2000; (c) Calculate the standard deviation of the Raman shift at a certain point for all SERS spectra of the same type after removing the cosmic spike in step (a).
8. The method for rapid identification of Acinetobacter complex according to claim 1, wherein In step (7), the spectral deconvolution method is used to deconstruct the characteristic peak composition in more detail, and the biological significance of each characteristic peak is found by searching the literature related to microorganisms.
9. The method for rapid identification of Acinetobacter complex according to claim 1, wherein In step (8), cluster analysis is performed on the Raman spectral data after quality control in step (6). The specific process is as follows: (a) The SERS data of three types of Acinetobacter after quality control were input into the PCA algorithm. The spectral data were subjected to dimensionality reduction and clustering visualization. The principal components PCA1 to PCA7 after dimensionality reduction were selected as the principal components. The clustering results of the Raman spectral principal component sample points in the classification quadrant were observed. (b) The preprocessed SERS data were input into the OPLS-DA algorithm, and the clustering results were evaluated using the three indicators R2X, R2Y, and Q2. The larger the indicator, the better the model performance, and the difference between R2X and Q2 did not exceed 0.
3.
10. The method for rapid identification of Acinetobacter complex according to claim 1, characterized in that In step (9), the artificial intelligence algorithms used in the data identification are eight classic machine learning algorithms: support vector machine, extreme gradient boosting machine, random forest, guided clustering, linear discriminant analysis, decision tree, adaptive boosting and quadratic discriminant analysis; the SER data set is divided into training set, validation set and test set in a ratio of 6:2:2; 5-fold cross validation is used for testing, and K is 5; all test set data are only used to test model performance, and grid search is used to search for the best parameter combination.
Citation Information
Patent Citations
Rapid identification method for Acinetobacter baumannii in serum and application of rapid identification method
CN108842006A