Method for rapidly identifying ST typing of acinetobacter
By combining surface-enhanced Raman spectroscopy and machine learning algorithms, the problem of tedious and time-consuming typing and identification of Acinetobacter was solved, and fast and accurate ST typing and identification of Acinetobacter were achieved. In particular, the use of convolutional neural networks improved the accuracy of identification.
Patent Information
- Application Number
- CN202410300198.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-15
- Publication Date
- 2025-09-16
AI Technical Summary
The existing methods for typing and identifying Acinetobacter are cumbersome and time-consuming, making it difficult to achieve rapid and accurate typing and identification.
Combining surface-enhanced Raman spectroscopy (SERS) with machine learning algorithms, a rapid identification method for ST typing of Acinetobacter was established through the preparation of Raman-enhanced substrates, sample preparation, data quality control, characteristic peak analysis and machine learning model training.
The rapid and accurate identification of ST typing of Acinetobacter was achieved, and the identification efficiency and accuracy were improved. In particular, through the application of the convolutional neural network model, an accuracy rate of 97.78% was achieved.
Smart Images

Figure CN120656550A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of spectral analysis technology, and specifically relates to using Raman spectroscopy to quickly analyze different ST typing of Acinetobacter, qualitatively analyzing the distribution of spectral characteristic peaks, and comparing different machine learning models to examine the accuracy of Acinetobacter ST typing identification. Background Art
[0002] Traditionally, the Acinetobacter complex consists of four species: A. calcoaceticus, A. baumannii, A. nosocomialis, and A. pituitii. With the exception of A. calcoaceticus, the remaining three species are the most commonly implicated in hospital- and community-acquired infections, causing respiratory-associated pneumonia, wound infections, urinary tract infections, and bloodstream infections. Therefore, to control the spread of Acinetobacter within hospitals and the community, analysis of Acinetobacter isolates is essential. However, despite being genetically closely related, all species of the Acinetobacter complex are phenotypically indistinguishable, necessitating molecular methods for accurate identification. Traditional methods for typing Acinetobacter include biotyping, serotyping, phage typing, immunotyping, drug susceptibility typing, and multilocus sequence typing (MLST). However, these techniques are cumbersome and time-consuming, making rapid and accurate identification of Acinetobacter species a challenge for clinical laboratories. In recent years, surface-enhanced Raman spectroscopy (SERS) combined with machine learning (ML) algorithms has been extensively studied for its rapid and accurate identification of bacterial pathogen species and subtypes.
[0003] Existing technologies for identifying Baumanii types mostly rely on genomics. For example, patent application publication number CN110910960B discloses a rapid molecular serotyping method for Acinetobacter baumannii. This method involves sequencing Baumanii strains and uploading the genome sequence for identification. The present invention eliminates the need for complex sequencing technology and simply combines the collected Raman spectra with artificial intelligence analysis methods to rapidly identify Acinetobacter complex groups and their types. Summary of the Invention
[0004] The purpose of the present invention is to provide a method for rapid identification of ST typing of Acinetobacter based on the existing technology, which solves the problems of cumbersome detection process and long time consumption by combining surface enhanced Raman spectroscopy with machine learning algorithm.
[0005] The technical solutions of the present invention are as follows:
[0006] A method for rapidly identifying ST typing of Acinetobacter spp. comprises the following steps:
[0007] (1) Bacterial collection: isolation and culture of Acinetobacter strains;
[0008] (2) Bacterial identification: After obtaining the nucleic acid sequence of the Acinetobacter strain obtained in step (1) by genome sequencing, the ST typing corresponding to different Acinetobacter strains is confirmed by multi-locus sequence analysis based on the housekeeping gene;
[0009] (3) Substrate preparation: sodium citrate solution is added to a silver nitrate solution that has been heated to boiling, and the mixture is stirred and reacted. The resulting reaction solution is centrifuged, and the resulting precipitate is resuspended in deionized water to obtain a Raman enhancement substrate with negatively charged nanosilver particles;
[0010] (4) Sample preparation: The different ST typing strains identified in step (2) were inoculated into phosphate buffer solution, and then mixed with 3-6 μL of the Raman enhancement substrate with negatively charged nanosilver particles prepared in step (3). Then, 2-5 μL of the mixed sample was titrated on a clean silicon wafer to form circular spots, and the sample was naturally dried in a safety cabinet to obtain the surface enhanced Raman spectroscopy (SERS) test sample of the Acinetobacter complex and ST typing;
[0011] (5) SERS detection: SERS spectral signal sampling of the SERS test samples of different ST types of Acinetobacter obtained in step (4) is performed multiple times and at multiple locations to obtain spectral fingerprints of the corresponding ST type samples, thereby constructing a SERS fingerprint database of different ST type strains of Acinetobacter;
[0012] (6) Data quality control: remove the cosmic spikes generated during the acquisition in step (5), and calculate the average Raman spectrum and standard deviation of each Acinetobacter ST type to check the data distribution and repeatability;
[0013] (7) Characteristic peak analysis: The spectral characteristic peaks of the ST-typed strains of Acinetobacter were deconstructed by spectral deconvolution analysis, and the overall distribution of the spectral characteristic peaks was viewed through a dot matrix plot;
[0014] (8) Data clustering: The Raman spectral data after quality control in step (6) were clustered using orthogonal partial least squares discriminant analysis to compare the differences between strains with different ST typing. The strains were clustered into different clusters in the classification quadrant of the Raman vector. The sample was determined to belong to the cluster based on the distribution of the different clusters. The results of the cluster analysis were evaluated using the three indicators of R2X, R2Y and Q2.
[0015] (9) Data identification: The SERS spectral data after quality control in step (6) are analyzed using different machine learning algorithms, divided into a training set, a validation set, and a test set using uniform random sampling, and a classifier is trained. K-fold cross validation is used for verification, where K is any integer from 1 to 10. The sample data are classified and labeled using the trained classifier and stored in a database;
[0016] (10) Data evaluation: Select test set samples to test the predictive ability, use accuracy, precision, recall, ROC curve and confusion matrix to modify the final judgment model, evaluate the performance of different machine learning algorithms, and select the best judgment model, including:
[0017]
[0018]
[0019]
[0020] Where: Precision represents precision, Recall represents recall, and AUC represents the area under the ROC curve, which is considered a performance indicator; TP, FP, TN, and FN represent the number of true positives, false positives, true negatives, and false negatives, respectively. These four cases constitute the confusion matrix.
[0021] (11) Performance verification: In the ST typing multi-classification task, in order to verify the reliability of the deep learning algorithm, we used the unsupervised learning algorithm t-distributed stochastic neighbor embedding (TSNE) to compare and analyze the spatial distribution of the original SERS data and the data after model processing.
[0022] In the present invention, in step (1), the Acinetobacter strain is a bacterial sample obtained by clinical isolation and culture. In a preferred embodiment, the process of bacterial collection is as follows: the Acinetobacter strain is cultured in Luria-Bertani liquid medium until the exponential growth phase, and after centrifugation, a supernatant and a strain pellet are obtained. The obtained strain pellet is resuspended in deionized water, and the concentration of the strain is determined by plate count method on a blood agar plate cultured at 37°C for 24 hours.
[0023] In the present invention, in step (2), when the ST typing of Acinetobacter in step (1) is determined using the MLST typing system, the number of housekeeping gene loci can be selected according to specific needs, and can be 5, 6, 7, 8 or 9.
[0024] Preferably, when determining the ST typing of the Acinetobacter strain, the number of housekeeping genes can be selected according to demand. For example, the number of housekeeping genes is 7, and the specific housekeeping genes are gapA, infB, mdh, pgi, phoE, rpoB or tonB. The ST typing is as follows: ST2, ST10, ST25, ST33, ST40, ST46, ST52, ST63, ST64, ST68, ST71, ST77, ST106, ST119, ST132, ST203, ST204, ST217, ST220, ST221, ST321, ST336, ST338, ST357, ST396, ST410, ST433 , ST457, ST516, ST629, ST768, ST795, ST821, ST33, ST1159, ST1264, ST1276, ST1333, ST1433, ST1828.
[0025] When preparing a Raman-enhanced substrate, silver nitrate is dissolved in water and stirred and heated to boiling, wherein the concentration of the silver nitrate in the silver nitrate solution is 0.5 to 1.5 mol / L, preferably 1 mol / L. The reducing agent is sodium citrate, and the sodium citrate solution is prepared with a concentration of 0.5 to 1.5 wt%, preferably 1 wt%.
[0026] In a preferred embodiment, the preparation method of the Raman enhanced substrate is as follows:
[0027] (a) A 1 mmol / L silver nitrate solution was heated to boiling, and 8 mL of a 1 wt% sodium citrate solution was added while stirring. The mixture was stirred at 500-700 rpm for 40-60 min to obtain negatively charged silver nanoparticles.
[0028] (b) The reaction solution obtained in step (a) was centrifuged at a speed of 5000-8000 r / min for 6-10 min, the supernatant was discarded, and the resulting precipitate was resuspended in deionized water to obtain a Raman enhancement substrate with negatively charged nanosilver particles.
[0029] In a more preferred embodiment, the preparation method of the Raman enhanced substrate is as follows:
[0030] (a) A 1 mmol / L silver nitrate solution was heated to boiling, and 8 mL of a 1 wt% sodium citrate solution was added while stirring. The mixture was stirred at 650 rpm for 40 min to obtain negatively charged silver nanoparticles.
[0031] (b) 1 mL of the reaction solution obtained in step (1) was centrifuged at 7000 rpm for 7 min. The supernatant was discarded and the resulting precipitate was resuspended in 100 μL of deionized water to obtain a negatively charged silver nanoparticle solution, which was the Raman enhancement substrate.
[0032] Regarding the negatively charged silver nanoparticles mentioned above, the diameter of the silver nanoparticles is less than 11 nm.
[0033] For the present invention, in step (4), the sample preparation is as follows: the different ST typing strains identified in steps (1) and (2) are inoculated into phosphate buffer solution respectively, and then mixed with the surface enhanced Raman substrate solution prepared in step (3) in equal amounts according to a volume of 3 to 6 μL, preferably 5 μL, and a volume of 2 to 5 μL of the mixed solution, preferably 2.5 μL, is dropped on a clean silicon wafer several times to form circular spots, and then naturally dried in a safety cabinet and tested to obtain surface enhanced Raman spectra (SERS) of the ST typing strain samples of Acinetobacter, thereby constructing a SERS fingerprint database of ST typing strains of Acinetobacter.
[0034] In the present invention, in step (5), SERS detection is suitable for collecting surface enhanced Raman signals of samples, and signals are collected from the SERS sample of the ST typing strain of Acinetobacter obtained in step (4).
[0035] Preferably, during SERS detection, 40 to 70 sites are randomly selected for each sample to be tested for Raman spectrum detection. In order to improve data balance, 50 sites are further preferably selected.
[0036] Preferably, SERS detection uses inVia™ micro-Raman spectroscopy to collect data, and the sampling parameters are: Raman spectroscopy excitation wavelength is 785 nm, detector type: high-sensitivity CCD array, exposure time is 10 s, and wavelength scanning range is 400-2800 cm-1.
[0037] Preferably, before spectrum acquisition, the SERS detection is performed using the Raman peak of the silicon wafer at 520 cm-1 as a reference peak for wave number calibration, and the dark current is deducted within the same integration time.
[0038] For the present invention, in step (6), the specific steps of data quality control are as follows:
[0039] (a) removing cosmic spikes from the SERS signal collected in step (5) by visual inspection;
[0040] (b) performing average spectrum calculation on the SERS spectra of the different ST types of bacilli after removing the sharp peak in step (a), calculating the average Raman shift of all SERS spectra of the same type at a certain point, selecting the number of Raman shifts to be 600 to 2000, preferably 1161;
[0041] (c) Calculate the standard deviation of the Raman shift at a certain point for all SERS spectra of the same type after removing the cosmic peak in step (a), and then complete the standard deviation value of the 1161 Raman shifts screened in step (b).
[0042] In the present invention, in step (7), a more detailed characteristic peak composition is deconstructed by the spectral deconvolution method. Since there are many types of ST typing, we use a dot matrix to display the characteristic peak composition of each type.
[0043] In a preferred embodiment, the characteristic peaks of the spectrum are specifically analyzed as follows:
[0044] (a) performing spectral deconvolution analysis on the average spectrum obtained in step (6), parsing the detailed characteristic peak composition of the average spectrum of each Acinetobacter ST typing and mining the differences between the spectra;
[0045] (b) Due to the large number of ST types, in order to fully demonstrate the differences in characteristic peaks between different ST types, the characteristic peak composition of each type is displayed using a dot matrix diagram, and the differences in characteristic peaks are viewed based on the node size and distribution.
[0046] In the present invention, in step (8), data clustering is used to determine whether the sample is in a specified interval and belongs to a specified category. It can be selected as needed, and the specific analysis method can be principal component analysis, linear discriminant analysis, partial least squares or orthogonal partial least squares discriminant analysis.
[0047] For the present invention, orthogonal partial least squares discriminant analysis (OPLS-DA) is used to qualitatively analyze the SERS fingerprints of the ST typing strains of Acinetobacter, discriminant analysis of the clustering of different strains, and quantitative evaluation indicators are used to evaluate the results of the cluster analysis. Preferably, in the present invention, the results of the cluster analysis are evaluated using the three indicators R2X, R2Y and Q2. The closer the values of the three indicators R2X, R2Y and Q2 are to 1, and the difference between the values of R2X and Q2 does not exceed 0.3, the better the result of data clustering.
[0048] For the present invention, in step (8), data clustering is only used as a preliminary judgment of the data to determine the quality of the data and the degree of distinguishability. In the present invention, data clustering can distinguish different ST types of Acinetobacter to a certain extent, but the results overlap, indicating that it cannot be completely distinguished. Moreover, data clustering cannot provide an identification model. For Raman spectral data that are not recorded in the database, it is necessary to rewrite the training and divide it again. The purposes of subsequent steps (9) and (10) include the following two points: one is to be able to better distinguish the ST types of Acinetobacter; the other is to be able to train an optimal model for identifying the ST types of Acinetobacter. When Raman spectral data that are not recorded in the database enter, the model does not need to be retrained and can directly identify the Raman spectral data. Data clustering can be understood as a preliminary experiment of steps (9) and (10) and is a transition.
[0049] For the present invention, in step (9), the identification of ST typing strains of Acinetobacter is 40 classifications, and the machine learning algorithms used are support vector machine, extreme gradient boosting machine, random forest, guided clustering, linear discriminant analysis, decision tree, adaptive boosting, quadratic discriminant analysis and deep learning algorithm convolutional neural network to construct the optimal identification model.
[0050] For the present invention, in step (9), the machine learning process includes computer storage, a computer processor, and a computer program stored in the computer storage and executable on the computer processor, wherein the computer storage stores a model optimization parameter range, and when the computer processor executes the computer program, the following steps are performed:
[0051] (a) Divide the Raman spectroscopy data signal after quality control into data sets;
[0052] (b) pre-set different machine learning training parameter ranges and select the optimal parameter combination;
[0053] (c) Using the machine learning model in data identification to classify the SERS fingerprints of ST-type strains of Acinetobacter, the classification results of different spectral data were obtained;
[0054] (d) Use data evaluation to evaluate the performance of different machine learning algorithms and select the best decision model;
[0055] (e) Verify the feature extraction performance of the convolutional layer of the deep learning algorithm convolutional neural network.
[0056] Step (a) involves grouping the pre-built sample database into a training set and a validation set using uniform random sampling, and constructing a test set by uniformly random sampling from the validation set. The training and validation set data are used for model training and validation, while the test set data is completely independent of these two sets and is used only to test the performance of the model.
[0057] Step (b) includes: setting parameter ranges for different machine learning algorithms, using a grid search approach to enumerate the model scores for each parameter combination, and selecting the best model parameter combination for final model performance comparison;
[0058] Step (c) includes: data classification using a data analysis tool to automatically analyze the Raman signal data of all samples, train multiple classifiers, use 5-fold cross-validation for detection, and classify and label the ST typing of Acinetobacter using the trained classifiers and store them in a database;
[0059] Step (d) includes: selecting some of the remaining samples for predictive ability testing, using accuracy, precision, recall, ROC curve and confusion matrix to modify the final judgment model, evaluating the performance of different machine learning algorithms, and selecting the best judgment model;
[0060] Step (e) includes: using the TSNE algorithm to visualize the original input matrix, checking the distribution of the spectral sample points in the Raman feature coordinate system, and then visualizing the features extracted by the convolutional neural network 4 to 8 layers, preferably 8 layers, and measuring the feature extraction performance of the convolutional neural network by comparing the distribution differences of the spectral sample points in the feature coordinate system.
[0061] Adopt the technical scheme of the present invention, the advantages are as follows:
[0062] A method for identification and classification of Acinetobacter ST typing based on surface-enhanced Raman spectroscopy combined with machine learning algorithm is provided, which provides a fast and effective analysis approach for the application of Raman spectrometers and reflects the application advantages of Raman spectrometers.
[0063] A SERS fingerprint library for ST typing of Acinetobacter was established through machine learning algorithms, and the newly collected Raman spectra were directly imported into the library for judgment. The comparison and combination of different machine learning methods and evaluation indicators can better ensure the accuracy of the identification results.
[0064] High accuracy: In the identification task of distinguishing ST types, the convolutional neural network achieved ideal identification results (accuracy = 97.78%, 5-fold cross-validation = 98.00%), indicating that this method is stable and reliable and can more effectively extract and analyze spectral data of different ST types of Acinetobacter. BRIEF DESCRIPTION OF THE DRAWINGS
[0065] Figure 1 : is the evolutionary tree diagram of ST typing of Acinetobacter in the present invention;
[0066] Figure 2 is the average Raman spectrum of different ST types of Acinetobacter in the present invention;
[0067] Figure 3 This is the deconvolution spectrum of different ST types of Acinetobacter in the present invention;
[0068] Figure 4 This is a dot matrix diagram showing the characteristic peak distribution of different ST types of Acinetobacter in the present invention;
[0069] Figure 5 This is the OPLS-DA clustering result diagram of different ST types of Acinetobacter in the present invention;
[0070] Figure 6 The score graph of different machine learning algorithms in the identification task of different ST types of Acinetobacter;
[0071] Figure 7 This is the confusion matrix diagram of the convolutional neural network algorithm in the identification task of different ST types of Acinetobacter;
[0072] Figure 8 Analyze the performance verification of convolutional neural network feature extraction for TSNE algorithm. DETAILED DESCRIPTION
[0073] Raman spectroscopy is based on the principle of inelastic scattering. That is, when incident light from a laser light source illuminates a substance, it is scattered by the molecules of the substance. A very small portion of the scattered light has a frequency different from the incident light. The change in the scattered light frequency depends on the structural characteristics of the illuminated substance. Different substances produce scattered light of specific frequencies under the same laser irradiation. Therefore, Raman spectroscopy can be used to achieve fast, simple, repeatable and non-destructive detection of material composition.
[0074] Artificial intelligence technology provides an efficient and accurate implementation for Raman spectroscopy-based material composition detection. Existing Raman spectroscopy machine learning algorithms are oriented toward specific substances to be tested. They transform the material identification problem of Raman spectroscopy into a machine learning classification problem. Machine learning models are trained based on standard Raman spectra of known substances, and the trained models are used to accurately identify test samples.
[0075] This specification and claims do not use differences in names as a way to distinguish components, but use differences in components' functions as the criteria for distinction. As mentioned in the description and claims throughout the text, "including" is an open-ended term and should be interpreted as "including but not limited to". "Approximately" means that within an acceptable error range, those skilled in the art can solve the technical problem within a certain error range and basically achieve the technical effect. In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the embodiments so that those skilled in the art can implement it with reference to the text of the specification. The equipment or raw materials used in the embodiments can all be obtained from the market.
[0076] instrument
[0077] Raman spectrometer (inViaTM), balance (ME104), water purifier (Arium mini), centrifuge (Centrifuge5424R), refrigerator (BCD-406WDPD), magnetic stirrer (RCT basic), clean bench (SW-CJ-2F), ultra-low temperature refrigerator (994), low-temperature incubator (PR505750R-CN), pH meter (FE28). The models and manufacturers of the instruments used in the implementation case are detailed in Table 1 below.
[0078] Table 1 Experimental instrument information
[0079]
[0080] Drugs and reagents
[0081] Silver nitrate (Sinopharm Chemical Reagent), sodium citrate (Sinopharm Chemical Reagent), sodium chloride (Sinopharm Chemical Reagent), and Luria-Bertani liquid medium (JingBio, Beijing).
[0082] Specific examples Rapid identification of ST typing of Acinetobacter using the method of the present invention
[0083] 1. Method
[0084] (1) Bacterial collection: All Acinetobacter strains were obtained from clinically isolated bacterial samples. In a preferred embodiment, the bacterial collection process is as follows: the Acinetobacter strains were cultured in Luria-Bertani liquid medium until the exponential growth phase. After centrifugation, the supernatant and strain pellets were obtained. The obtained strain pellets were resuspended in deionized water, and the strain concentration was determined by plate count method on blood agar plates cultured at 37°C for 24 hours.
[0085] (2) Typing identification: After the nucleic acid sequence of the strain obtained in step (1) is obtained by genome sequencing, the ST typing of the strain is confirmed according to the housekeeping gene by the multi-locus sequence analysis (MLST) method; wherein the number of housekeeping genes is 7, specifically gapA, infB, mdh, pgi, phoE, rpoB or tonB, and the specific typing is as follows: ST2, ST10, ST25, ST33, ST40, ST46, ST52, ST63, ST64, ST68 , ST71, ST77, ST106, ST119, ST132, ST203, ST204, ST217, ST220, ST221, ST321, ST336, ST338, ST357, ST396, ST4 10. ST433, ST457, ST516, ST629, ST768, ST795, ST821, ST33, ST1159, ST1264, ST1276, ST1333, ST1433, ST1828.
[0086] (3) Substrate preparation:
[0087] The preparation process of the Raman enhanced substrate is as follows:
[0088] (a) 33.72 mg of silver nitrate was weighed and dissolved in 200 mL of deionized water to obtain a 1 mol / L silver nitrate solution. The solution was heated to boiling using a magnetic stirrer. 8 mL of a 1 wt % sodium citrate solution was then added at once while stirring. The mixture was stirred at 650 rpm for 40 min to obtain a negatively charged silver nanoparticle reaction solution having a diameter of <11 nm.
[0089] (b) 1 mL of the silver nanoparticle reaction solution obtained in step (a) was centrifuged at 7000 rpm for 7 min. The supernatant was discarded and the resulting precipitate was resuspended in 100 μL of deionized water to obtain a Raman enhancement substrate solution of negatively charged silver nanoparticles. The resulting solution was uniformly milky white and stored in the dark at room temperature until use.
[0090] (4) Sample preparation: The different ST typing strains identified in step (2) were inoculated into a phosphate buffer solution, and then sequentially mixed with the surface enhanced Raman substrate solution prepared in step (3) in equal amounts of 5 μL. 2.5 μL of the mixed solution was dropped onto a clean silicon wafer several times to form circular spots. After natural drying in a safety cabinet, the surface enhanced Raman spectroscopy (SERS) samples of Acinetobacter complex and ST typing were obtained.
[0091] (5) SERS detection: The air-dried SERS sample in step (4) was taken for measurement. During the measurement, multiple and multi-site Raman spectroscopy sampling was adopted. 50 SERS fingerprint spectra were obtained for each sample, thereby constructing a SERS fingerprint database of ST-type strains of Acinetobacter;
[0092] The Raman spectrum acquisition parameters during the measurement process were: an excitation wavelength of 785 nm, a high-sensitivity CCD array detector, a 10-second exposure time, a wavelength scan range of 400–2800 cm⁻¹, output data type: txt, a laser intensity of 2, and a spectral acquisition time of 10 seconds. Wavenumber calibration was performed using the silicon wafer's Raman peak at 520 cm⁻¹ as a reference peak before spectral acquisition, and dark current was subtracted using the same integration time. Raman spectroscopy was performed at 45 sites per sample.
[0093] (6) Data quality control:
[0094] The preparation process is as follows:
[0095] (a) removing cosmic spikes from the SERS signal collected in step (5) by visual inspection;
[0096] (b) performing spectrum average calculation on the SERS spectra of the different ST-typed strains after removing the sharp peak in step (a), and calculating the average Raman shift of all SERS spectra of the same type at a certain point, with the number of Raman shifts being 1161;
[0097] (c) Calculate the standard deviation of the Raman shift at a certain point for all SERS spectra of the same type after removing the cosmic peak in step (a), and then complete the standard deviation value of the 1161 Raman shifts screened in step (b).
[0098] (7) Characteristic peak analysis: To comprehensively demonstrate the differences in characteristic peaks between different ST types, we used a dot matrix to display the characteristic peak composition of each type.
[0099] (8) Data clustering: The Raman spectral data after quality control in step (5) were clustered using orthogonal partial least squares discriminant analysis (OPLS-DA). The Raman spectrum category was determined by comparing the differences between the SERS spectra of different ST types of Acinetobacter strains, and a classification quadrant diagram of the ST type Raman vector was formed. The different ST type strains were distinguished by observing the clustering of Raman spectral sample points in the classification quadrant diagram. The results of the cluster analysis were evaluated using the three indicators of R2X, R2Y and Q2.
[0100] (9) Data Identification: The Raman spectral data after quality control in step (5) were automatically analyzed using various machine learning algorithms. The SERS spectral datasets of the different ST-typed strains were divided into training, validation, and test sets using uniform random sampling. A classifier was trained and tested using 5-fold cross-validation. The trained classifiers were used to assign and label the sample data and store them in a database. All parameters were optimized before training. The machine learning parameter settings are shown in Table 2 below.
[0101] Table 2 Machine learning parameter settings for the Acinetobacter ST typing task
[0102]
[0103] (10) Data evaluation: Select test set samples to test the predictive ability, use precision, recall rate, ROC curve and confusion matrix to modify the final judgment model, evaluate the performance of different machine learning algorithms, and select the best judgment model, including:
[0104]
[0105]
[0106]
[0107] Among them: Precision represents accuracy, Recall represents recall rate, AUC represents area under the ROC curve, which is regarded as a performance indicator; TP, FP, TN and FN represent the number of true positives, false positives, true negatives and false negatives respectively, and the confusion matrix is constructed by these four cases.
[0108] (11) Performance verification: The TSNE algorithm is used to visualize the original input matrix and check the distribution of the spectral sample points in the Raman feature coordinate system. Then, the features extracted by the last layer of the convolutional neural network, i.e., the 8th layer, are visualized. The feature extraction performance of the convolutional neural network is measured by comparing the distribution differences of the spectral sample points in the feature coordinate system.
[0109] 2. Results
[0110] Figure 1 This is the evolutionary tree clustering result of different ST typing strains in the present invention. It can be seen in the figure that different strains with the same typing are clustered to the same position;
[0111] Figure 2 The average Raman spectra of different ST-typed strains in the present invention are shown in Figure 1. The shaded area in the figure represents the standard error band of the SERS spectrum of each ST-type. The width of the error band indicates the quality of the data.
[0112] Figure 3 This is the deconvolution Raman spectra of different ST typing strains in the present invention. The differences between different ST typing are magnified in the figure. The differences between different ST typing can be viewed by observing the convolution peak distribution of different typing;
[0113] Figure 4 This is a dot matrix diagram of the spectral characteristic peaks fitted by the average spectrum lines of different ST typing in the present invention. Each node indicates that the strain has the characteristic peak, and the size of the node indicates the signal intensity of the characteristic peak at that location.
[0114] Figure 5 This is the clustering result obtained by clustering the spectral data of different ST types using the OPLS-DA algorithm in the present invention in the form of a scatter plot. Different ST types are clustered into different clusters, but there is a lot of overlap. Only ST25, ST68, ST220, ST336, ST396, and ST410 can be distinguished, indicating that the OPLS-DA algorithm cannot distinguish different ST-type strains. The three indicators R2X, R2Y, and Q2 were used to evaluate the cluster analysis results. The results showed that R2X = 0.994, R2Y = 0.272, and Q2 = 0.271. The R2X and Q2 results differed significantly, indicating that the SERS data of different ST types cannot be distinguished using OPLS-DA, so more advanced computational analysis methods are needed.
[0115] Figure 6 The following is a score chart of different machine learning algorithms on SERS data of different ST typing strains. The results show that the convolutional neural network has the best performance, with all scores exceeding 95%. The identification performance is better than other algorithms, and the support vector machine also achieves strong classification performance. This also shows that the support vector machine is an efficient recognition method. However, the convolutional neural network is more suitable for multi-class classification tasks due to its stronger feature extraction ability.
[0116] Figure 7 The confusion matrix of the convolutional neural network with the best performance shows that the convolutional neural network model is able to completely and accurately identify most of the typing strains. Some misidentifications mostly occur in ST63. For example, 20% of the SERS data of ST119 were misidentified as ST63, and 9% of the SERS data of ST45 were misidentified as ST63.
[0117] Figure 8This is the verification process of the convolutional neural network model performance. Figure (A) shows the SERS sample point clustering result after the original input matrix is analyzed by the TSNE algorithm. It can be seen that most sample points overlap and there is no pattern between different sample clusters; Figure (B) is a visualization of the abstract feature data output by the last convolutional layer after feature extraction through multiple convolutional layers. It can be seen that the SERS data of different ST types processed by the convolutional neural network are effectively clustered into different clusters, which verifies that convolutional neural networks are an excellent feature extraction method and an efficient tool for identifying different types of strains.
[0118] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that it is still possible to modify the technical solutions described in the aforementioned embodiments, or to make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for rapid identification of ST typing of Acinetobacter, characterized in that: It includes the following steps: (1) Bacterial collection: isolation and culture of Acinetobacter strains; (2) Bacterial identification: After obtaining the nucleic acid sequence of the Acinetobacter strain obtained in step (1) by genome sequencing, the ST typing corresponding to different Acinetobacter strains is confirmed by multi-locus sequence analysis based on the housekeeping gene; (3) Substrate preparation: sodium citrate solution is added to a silver nitrate solution that has been heated to boiling, and the mixture is stirred and reacted. The resulting reaction solution is centrifuged, and the resulting precipitate is resuspended in deionized water to obtain a Raman enhancement substrate with negatively charged nanosilver particles; (4) Sample preparation: The different ST typing strains identified in step (2) were inoculated into phosphate buffer solution, and then mixed with 3-6 μL of the Raman enhancement substrate with negatively charged nanosilver particles prepared in step (3). Then, 2-5 μL of the mixed sample was titrated on a clean silicon wafer to form circular spots, and the sample was naturally dried in a safety cabinet to obtain the surface enhanced Raman spectroscopy (SERS) test sample of the Acinetobacter complex and ST typing; (5) SERS detection: SERS spectral signal sampling of the SERS test samples of different ST types of Acinetobacter obtained in step (4) is performed multiple times and at multiple locations to obtain spectral fingerprints of the corresponding ST type samples, thereby constructing a SERS fingerprint database of different ST type strains of Acinetobacter; (6) Data quality control: remove the cosmic spikes generated during the acquisition in step (5), and calculate the average Raman spectrum and standard deviation of each Acinetobacter ST type to check the data distribution and repeatability; (7) Characteristic peak analysis: The spectral characteristic peaks of the ST-typed strains of Acinetobacter were deconstructed by spectral deconvolution analysis, and the overall distribution of the spectral characteristic peaks was viewed through a dot matrix plot; (8) Data clustering: The Raman spectral data after quality control in step (6) were clustered using orthogonal partial least squares discriminant analysis to compare the differences between strains with different ST typing. The strains were clustered into different clusters in the classification quadrant of the Raman vector. The sample was determined to belong to the cluster based on the distribution of the different clusters. The results of the cluster analysis were evaluated using the three indicators of R2X, R2Y and Q2. (9) Data identification: The SERS spectral data after quality control in step (6) are analyzed using different machine learning algorithms, divided into a training set, a validation set, and a test set using uniform random sampling, and a classifier is trained. K-fold cross validation is used for verification, where K is any integer from 1 to 10. The sample data are classified and labeled using the trained classifier and stored in a database; (10) Data evaluation: Select test set samples to test the predictive ability, use accuracy, precision, recall, ROC curve and confusion matrix to modify the final judgment model, evaluate the performance of different machine learning algorithms, and select the best judgment model, including: Where: Precision represents precision, Recall represents recall, and AUC represents the area under the ROC curve, which is considered a performance indicator; TP, FP, TN, and FN represent the number of true positives, false positives, true negatives, and false negatives, respectively. These four cases constitute the confusion matrix. (11) Performance verification: In the ST typing multi-classification task, in order to verify the reliability of the deep learning algorithm, we used the unsupervised learning algorithm t-distributed stochastic neighbor embedding (TSNE) to compare and analyze the spatial distribution of the original SERS data and the data after model processing.
2. The method for rapid identification of ST typing of Acinetobacter according to claim 1, wherein In step (1), the bacterial collection process is as follows: the Acinetobacter strain is cultured in Luria-Bertani liquid medium to the exponential growth phase, and after centrifugation, a supernatant and strain pellets are obtained. The obtained strain pellets are resuspended in deionized water, and the concentration of the strain is determined by the plate count method on a blood agar plate cultured at 37°C for 24 hours.
3. The method for rapid identification of ST typing of Acinetobacter according to claim 1, wherein In step (2), the housekeeping gene is gapA, infB, mdh, pgi, phoE, rpoB or tonB.
4. The method for rapid identification of ST typing of Acinetobacter according to claim 1, wherein In step (3), the preparation method of the Raman enhanced substrate is as follows: (a) A 1 mmol / L silver nitrate solution was heated to boiling, and 8 mL of a 1 wt% sodium citrate solution was added while stirring. The mixture was stirred at 500-700 rpm for 40-60 min to obtain negatively charged silver nanoparticles. (b) The reaction solution obtained in step (a) was centrifuged at a speed of 5000-8000 r / min for 6-10 min, the supernatant was discarded, and the resulting precipitate was resuspended in deionized water to obtain a Raman enhancement substrate with negatively charged nanosilver particles.
5. The method for rapid identification of ST typing of Acinetobacter according to claim 1, wherein In step (5), the multi-site SERS spectrum signal sampling is performed at 40 to 70 sites; the Raman spectrum sampling parameters are as follows: the excitation wavelength of the Raman spectrum is 785 nm, the detector type is a high-sensitivity CCD array, the exposure time is 10 s, and the wavelength scanning range is 400-2800 cm -1 .
6. The method for rapid identification of ST typing of Acinetobacter according to claim 1, wherein In step (6), the specific steps of data quality control are as follows: (a) removing cosmic spikes from the SERS signal collected in step (5) by visual inspection; (b) performing average spectrum calculation on the SERS spectra of the different ST types of bacilli after removing the sharp peak in step (a), and calculating the average Raman shift of all SERS spectra of the same type at a certain point, with the number of Raman shifts selected to be 600 to 2000; (c) Calculate the standard deviation of the Raman shift at a certain point for all SERS spectra of the same type after removing the cosmic spike in step (a).
7. The method for rapid identification of ST typing of Acinetobacter according to claim 1, wherein In step (7), the spectral deconvolution method is used to deconstruct a more detailed characteristic peak composition. Since there are many types of ST types, we use a dot matrix to display the characteristic peak composition of each type.
8. The method for rapid identification of ST typing of Acinetobacter according to claim 1, wherein In step (8), the SERS data of ST typing of Acinetobacter were input into the OPLS-DA algorithm, and the clustering results were evaluated using the three evaluation indicators R2X, R2Y and Q2. The closer the values of the three indicators R2X, R2Y and Q2 were to 1, and the difference between the values of R2X and Q2 did not exceed 0.3, the better the data clustering results were.
9. The method for rapid identification of ST typing of Acinetobacter according to claim 1, wherein In step (9), the machine learning algorithms are support vector machine, extreme gradient boosting machine, random forest, guided clustering, linear discriminant analysis, decision tree, adaptive boosting, quadratic discriminant analysis and deep learning algorithm convolutional neural network to construct the optimal identification model; the SER dataset is divided into training set, validation set and test set in a ratio of 6:2:2; 5-fold cross validation is used for testing, and the K is 5.
10. The method for rapid identification of ST typing of Acinetobacter according to claim 1, characterized in that In step (11), in order to verify the reliability of the deep learning algorithm in step (9) in the ST typing classification task, the distribution of the original typing SERS data in the feature space and the distribution of the data output after convolution and abstraction by the convolutional layer of the convolutional neural network in the space are compared by the TSNE algorithm, and the number of the convolutional layers is 4 to 8.
Citation Information
Patent Citations
A rapid molecular serotype analysis method for Acinetobacter baumannii
CN110910960B