Method for rapidly identifying salmonella serotype

By combining surface-enhanced Raman spectroscopy with machine learning, the time-consuming and tedious problem of Salmonella serotype identification was solved, and rapid and accurate serotype identification was achieved. The accuracy of the machine learning model was as high as 99%.

CN120685613APending Publication Date: 2025-09-23GUANGDONG GENERAL HOSPITAL
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410316607.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-03-20
Publication Date
2025-09-23

AI Technical Summary

Technical Problem

The existing methods for identifying Salmonella serotypes are time-consuming and cumbersome. Traditional methods require multiple reagent materials and complex sample pretreatment, and are unable to quickly and accurately identify the spectral information differences between different bacteria.

Method used

Combining surface-enhanced Raman spectroscopy with machine learning algorithms, by preparing nanosilver particle substrates, collecting Raman spectral data, performing data quality control and characteristic peak analysis, and using orthogonal partial least squares discriminant analysis and multiple machine learning algorithms for clustering and identification, a rapid identification model was established.

Benefits of technology

It has achieved rapid and accurate identification of Salmonella serotypes, improved detection efficiency, and the accuracy of the machine learning model has reached over 99%, with good stability and the ability to effectively distinguish different serotypes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120685613A_ABST
    Figure CN120685613A_ABST
Patent Text Reader

Abstract

The invention provides a method for rapidly identifying salmonella serotypes, which solves the problems of tedious detection process, long consumed time and the like in a mode of combining surface enhanced Raman spectroscopy preparation and a machine learning algorithm. On the other hand, Raman spectrum libraries of different salmonella serotypes are established through machine learning, newly collected Raman spectrums are directly imported into the libraries for judgment, and the accuracy of identification results can be better guaranteed through comparison and combination of different machine learning methods and evaluation indexes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of spectral analysis technology, and specifically relates to using Raman spectroscopy to quickly analyze different Salmonella serotypes, qualitatively analyzing the distribution of spectral characteristic peaks, and comparing the accuracy of different machine learning models in identifying Salmonella serotypes. Background Art

[0002] Salmonella is a Gram-negative, facultative Enterobacteriaceae. Poultry and poultry products are a major source of Salmonella contamination. Improper cooking, heating, and food reprocessing can contribute to Salmonella outbreaks and infections worldwide. Low nutrient concentrations and temperatures are a major challenge in controlling Salmonella in agriculture and food processing. Traditional microbial identification methods are time-consuming and labor-intensive. Plate culture, the gold standard for detecting Salmonella, requires bacterial pre-enrichment, isolation, biochemical testing, and serum identification—a complex, time-consuming process that prevents rapid identification of Salmonella. Traditional detection methods, such as polymerase chain reaction (PCR) and enzyme-linked immunosorbent assay (ELISA), require multiple reagents and materials, and sample pre-treatment is very demanding for the operator. Therefore, rapid detection methods for Salmonella in food continue to be developed and are a pressing need. Furthermore, rapid pathogen detection would be extremely useful for quality control in large-scale food processing and production processes. In recent years, Raman spectroscopy has been increasingly used in the field of pathogenic microorganisms due to its rapid, non-destructive detection, high sensitivity, and high efficiency. However, due to the ubiquity of biomacromolecules such as nucleic acids, proteins, lipids, and sugars in pathogenic bacteria, the spectral information differences between bacteria are minimal, not to mention the differences between different serotypes within the same genus. To accurately detect Salmonella serotypes from highly similar Raman spectroscopy data, advanced machine learning algorithms have been introduced. Numerous studies have confirmed the effectiveness of combining machine learning with Raman spectroscopy in bacterial classification and identification.

[0003] Existing technologies for microbial typing and identification mostly rely on polymerase chain reaction methods, which extract DNA and design primers to identify serotypes. For example, patent application publication number CN104059977B discloses a Salmonella serotype identification method and kit. This method extracts DNA from a test sample, uses PCR to detect the resulting DNA template, and combines CRISPR technology with primer detection. This method is often time-consuming. With the continuous development of machine learning and artificial intelligence technologies, methods based on machine learning and intelligent Raman spectroscopy data processing and identification will be the development trend of Raman spectroscopy instrument processing and identification. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for rapid identification of Salmonella serotypes based on the existing technology. By combining surface-enhanced Raman spectroscopy with machine learning algorithms, the problems of complicated and time-consuming detection process can be solved.

[0005] The technical solutions of the present invention are as follows:

[0006] A method for rapidly identifying Salmonella serotypes comprises the following steps:

[0007] (1) Bacterial collection: Isolation and culture of different serotypes of Salmonella;

[0008] (2) Substrate preparation: sodium citrate solution is added to a boiling silver nitrate solution and stirred for reaction. The resulting reaction solution is centrifuged and the resulting precipitate is resuspended in deionized water to obtain a Raman enhancement substrate with negatively charged nanosilver particles.

[0009] (3) Sample preparation: Each type of Salmonella cultured in step (1) is inoculated into a phosphate buffer solution, and then mixed with 3 to 6 μL of the Raman enhancement substrate prepared in step (2). Then, 2 to 5 μL of the mixed sample is titrated on a clean silicon wafer to form circular spots, and the sample is naturally dried in a safety cabinet to obtain high-sensitivity Raman signal detection samples for different Salmonella species to be tested;

[0010] (4) SERS detection: The high-sensitivity Raman signal of the Salmonella sample obtained in step (3) is subjected to multiple, multi-site Raman spectrum sampling to obtain corresponding SERS fingerprints of different Salmonella species, thereby constructing a SERS fingerprint database of different Salmonella species;

[0011] (5) Data quality control: The SERS spectra of different Salmonella species obtained in step (4) were sequentially de-peaked, and the average spectrum and standard deviation of the spectra were calculated;

[0012] (6) Difference evaluation: The data after data quality control in step (5) are subjected to quality evaluation, and the differences between the original SERS spectra of different Salmonella are examined by deconvolution spectral analysis;

[0013] (7) Characteristic peak analysis: Analyze the characteristic peaks of the spectrum obtained by deconvolution fitting in step (6) difference assessment;

[0014] (8) Spectral preprocessing: performing data normalization on the data after quality control in step (5);

[0015] (9) Data clustering: The Raman spectral data obtained through quality control in step (5) and pre-processed in step (7) were clustered using orthogonal partial least squares discriminant analysis (OPLS-DA). By comparing the distribution differences of Salmonella spectral samples in the feature space, a classification quadrant diagram of the Raman vectors of different Salmonella was formed. The clustering performance was determined based on the degree of dispersion between different clusters, and the results of the clustering analysis were evaluated using the three indicators of R2X, R2Y, and Q2.

[0016] (10) Data identification: Using different machine learning algorithms to automatically analyze the Raman spectral data preprocessed in step (7), the spectral data sets of different Salmonella are divided into training sets, validation sets, and test sets by uniform random sampling, and a classifier is trained and tested using K-fold cross validation, where K is any integer from 1 to 10. The sample data are classified and labeled by the trained classifier and stored in a database;

[0017] (11) Data evaluation: Select test set samples to test the predictive ability, use precision, recall rate, ROC curve and confusion matrix to modify the final judgment model, evaluate the performance of different machine learning algorithms, and select the best judgment model, including:

[0018]

[0019]

[0020]

[0021] Among them: Precision represents accuracy, Recall represents recall rate, AUC represents area under the ROC curve, which is regarded as a performance indicator; TP, FP, TN and FN represent the number of true positives, false positives, true negatives and false negatives respectively, and the confusion matrix is ​​constructed by these four cases.

[0022] In the present invention, in step (1), the different serotypes of Salmonella are bacterial samples isolated from clinical samples. In a preferred embodiment, the process of bacterial isolation and culture is as follows: different Salmonella are cultured in Luria-Bertani (LB) liquid medium until the exponential growth phase, and after centrifugation, a supernatant and strain pellets are obtained. The obtained strain pellets are resuspended in deionized water, and the concentration of the strains is determined by plate count method on blood agar plates cultured at 37°C for 24 hours.

[0023] In the present invention, when preparing the Raman enhancement substrate in step (2), silver nitrate is dissolved in water and stirred and heated to boiling, wherein the concentration of the silver nitrate in the silver nitrate solution is 0.5 to 1.5 mol / L, preferably 1 mol / L. The reducing agent is sodium citrate, and the sodium citrate is prepared into a sodium citrate solution with a concentration of 0.5 to 1.5 wt%, preferably 1 wt%.

[0024] In a preferred embodiment, the preparation method of the Raman enhanced substrate is as follows:

[0025] (a) A 1 mmol / L silver nitrate solution was heated to boiling, and 8 mL of a 1 wt% sodium citrate solution was added while stirring. The mixture was stirred at 500-800 rpm for 30-50 min to obtain negatively charged silver nanoparticles.

[0026] (b) The reaction solution obtained in step (a) was centrifuged at 6500-8500 r / min for 6-9 min, the supernatant was discarded, and the precipitate was resuspended in deionized water to obtain a negatively charged silver nanoparticle solution.

[0027] In a more preferred embodiment, the preparation method of the Raman enhanced substrate is as follows:

[0028] (a) A 1 mmol / L silver nitrate solution was heated to boiling, and 8 mL of a 1 wt% sodium citrate solution was added while stirring. The mixture was stirred at 650 rpm for 40 min to obtain negatively charged silver nanoparticles.

[0029] (b) 1 mL of the reaction solution obtained in step (a) was centrifuged at 7000 rpm for 7 min. The supernatant was discarded and the resulting precipitate was resuspended in 100 μL of deionized water to obtain a negatively charged silver nanoparticle solution, which served as the Raman enhancement substrate.

[0030] Regarding the negatively charged silver nanoparticles mentioned above, the diameter of the silver nanoparticles is less than 11 nm.

[0031] In the present invention, in steps (3) and (4), the preparation of a high-sensitivity SERS sample of Salmonella includes mixing the high-sensitivity Raman substrate obtained in step (3) with the bacteria cultured in step (1), dropping appropriate amounts onto a silicon wafer multiple times, and performing detection after natural drying to obtain corresponding SERS spectra of Salmonella of different serotypes, thereby constructing a SERS spectrum database of Salmonella;

[0032] Preferably, 5 μL of Raman spectroscopy enhancement substrate solution with negatively charged nanosilver particles is mixed with an equal amount of Salmonella on a vortex mixer for 30 seconds, and then 3 μL of the mixed sample is dropped on a clean silicon wafer to form a circular spot, and then naturally dried in a safety cabinet for detection;

[0033] Preferably, SERS detection uses inVia™ micro-Raman spectroscopy to collect data. The sampling parameters are as follows: the excitation wavelength of the Raman spectrum is 785 nm, the detector type is a high-sensitivity CCD array, the exposure time is 10 s, and the wavelength scanning range is 500-1800 cm -1 ;

[0034] Preferably, SERS detection is performed using a silicon wafer at 520 cm before spectrum acquisition. -1 The Raman peak at is used as the reference peak for wavenumber calibration, and the dark current is deducted within the same integration time.

[0035] Preferably, during SERS detection, 700 to 1000 sites are randomly selected for Raman spectroscopy detection for each sample to be tested. In order to improve the balance of spectral data between different Salmonella species, 800 sites are further preferably selected.

[0036] In the present invention, in step (5), the Raman spectra of the different Salmonella samples obtained in step (4) are de-peaked, the average Raman spectrum is plotted and the standard deviation is calculated, and the smoothness of the spectrum and the size of the error band are observed to determine the repeatability of the data.

[0037] For the present invention, in a preferred embodiment, the difference analysis in step (6) is as follows: the SERS fingerprint spectrum after data quality control in step (5) is deconstructed using spectral deconvolution spectrum to obtain the full spectral characteristic peaks of each Salmonella. The deconvolution function can be selected as needed, such as Gaussian function, Lorentz function or central frequency symmetric function (vogit) of Gaussian combined with Lorentz function, preferably vogit function.

[0038] For the present invention, in step (7), the characteristic peak analysis process is as follows:

[0039] (a) comparing the spectral characteristic peaks generated by the differential analysis in step (6) to identify the unique characteristic peaks of each Salmonella species;

[0040] (b) selecting 6 to 10, preferably 8, spectral characteristic peaks common to each type of Salmonella from the spectral characteristic peaks generated by the differential analysis in step (6), and using a box plot to analyze the signal differences of different bacteria at the same characteristic peaks;

[0041] (c) For the characteristic peaks of step (a) and step (b), search the literature to find the molecular vibration and biological significance corresponding to each characteristic peak.

[0042] For the present invention, in step (8), the data clustering algorithm includes: maximum and minimum normalization, vector normalization, z-score normalization, decimal scaling normalization and logarithmic normalization; preferably, the present invention normalizes each spectral data to the range of [0,1] through maximum and minimum normalization.

[0043] For the present invention, in step (9), data clustering is used to determine whether the sample is in a specified interval and category. The qualitative analysis method is specifically selected according to actual needs, and can be unsupervised learning principal component analysis, density cluster analysis, K-means analysis, and supervised learning algorithm linear discriminant analysis, partial least squares, orthogonal partial least squares discriminant analysis. Preferably, the present invention uses OPLS-DA algorithm to cluster different Salmonella spectra, and uses R2X, R2Y and Q2 to evaluate the results of cluster analysis. The closer the values ​​of R2X, R2Y and Q2 are to 1, and the difference between the values ​​of R2X and Q2 does not exceed 0.3, the better the result of data clustering.

[0044] For the present invention, in step (9), data clustering is only used as a preliminary judgment of the data to determine the quality of the data and the degree of distinguishability. In the present invention, data clustering can distinguish different serotypes of Salmonella to a certain extent, but the results are overlapping, indicating that they cannot be completely distinguished. Moreover, data clustering cannot provide an identification model. For Raman spectral data that are not recorded in the database, it is necessary to rewrite the training and divide it again. The purposes of the subsequent steps (10) and (11) include the following two points: one is to be able to better distinguish the serotypes of different strains; the other is to be able to train an optimal model for identifying different Salmonella. When Raman spectral data that are not recorded in the database enter, the model does not need to be retrained and can directly identify the Raman spectral data. Data clustering can be understood as a preliminary experiment of steps (10) and (11) and is a transition.

[0045] In the present invention, in step (10), the different machine learning algorithms can specifically be support vector machine (SVM), random forest (RF), extreme gradient boosting (XGBoost), adaptive boosting (AdaBoost), decision tree (DT) and quadratic discriminant analysis (QDA). Among them, the support vector machine (SVM) algorithm has the highest accuracy and the best effect.

[0046] For the present invention, in step (10), the data is identified as a machine learning 4-classification model, the machine learning process includes computer storage, a computer processor, and a computer program stored in the computer memory and executable on the computer processor, the computer memory stores a model optimization parameter range, and the computer processor performs the following steps when executing the computer program:

[0047] a) dividing the quality control and pre-processed Raman spectroscopy data signals into data sets;

[0048] b) Pre-set different machine learning training parameter ranges and select the optimal parameter combination;

[0049] c) using the machine learning model in data identification to classify the Raman spectral data of different Salmonella species to obtain classification results of the different spectral data;

[0050] d) Use data evaluation to evaluate the performance of different machine learning algorithms and select the best judgment model.

[0051] Step a) involves grouping the pre-built sample database into a training set and a validation set using uniform random sampling, and constructing a test set from a uniform random sampling of the validation set. The training and validation set data are used for model training and validation, while the test set data is completely independent of these two sets and is used only to test the model's performance.

[0052] Step b) includes: setting parameter ranges for different machine learning algorithms, using a grid search approach to enumerate the model scores for each parameter combination, and selecting the best model parameter combination for final model performance comparison.

[0053] Step c) includes data classification, which involves automatically analyzing the Raman signal data of all samples using a data analysis tool, training multiple classifiers, and performing detection using 5-fold cross-validation. The spectral data of different strains are assigned and labeled as ST categories using the trained classifiers and stored in a database.

[0054] Step d) includes: selecting some of the remaining samples for predictive ability testing, using accuracy, precision, recall, ROC curve and confusion matrix to modify the final judgment model, evaluating the performance of different machine learning algorithms, and selecting the best judgment model.

[0055] Adopt the technical scheme of the present invention, the advantages are as follows:

[0056] The use of machine learning combined with the SERS method to classify and identify Salmonella samples provides a rapid detection method. The application of Raman spectrometer provides a fast and effective analysis approach, reflecting the application advantages of Raman spectrometer.

[0057] A Raman spectrum library of different glycogen samples is established through machine learning and deep learning algorithms, and the newly collected fingerprint spectra are directly imported into the library for judgment. The comparison and combination of different machine learning and deep learning methods and evaluation indicators can better ensure the accuracy of the identification results.

[0058] High accuracy: In the task of distinguishing different Salmonella samples, the optimal model support vector machine of the present invention scored more than 99% for each indicator; good stability: The 5-fold cross-validation score of the convolutional neural network was 99.98%, which was close to the scores of each scoring indicator, indicating that the model did not overfit and had strong universality. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] Figure 1 The average Raman spectra and Raman spectrum deconvolution images of different Salmonella serotypes in the present invention;

[0060] Figure 2 This is a graph showing the difference in the characteristic peaks of Salmonella in the present invention;

[0061] Figure 3 This is the clustering result diagram of Salmonella before and after pretreatment by OPLS-DA analysis in the present invention;

[0062] Figure 4 This is a performance gradient graph for training different machine learning algorithm models in the present invention;

[0063] Figure 5 Confusion matrix of different machine learning algorithm models in the present invention;

[0064] Figure 6 The biological significance diagram of the unique characteristic peaks of each Salmonella in the present invention;

[0065] Figure 7 This is a performance score diagram of different machine learning algorithms in the present invention. DETAILED DESCRIPTION

[0066] Raman spectroscopy is based on the principle of inelastic scattering. That is, when incident light from a laser light source illuminates a substance, it is scattered by the molecules of the substance. A very small portion of the scattered light has a frequency different from the incident light. The change in the scattered light frequency depends on the structural characteristics of the illuminated substance. Different substances produce scattered light of specific frequencies under the same laser irradiation. Therefore, Raman spectroscopy can be used to achieve fast, simple, repeatable and non-destructive detection of material composition.

[0067] Artificial intelligence technology provides an efficient and accurate implementation for Raman spectroscopy-based material composition detection. Existing Raman spectroscopy machine learning algorithms are oriented toward specific substances to be tested. They transform the material identification problem of Raman spectroscopy into a machine learning classification problem. Machine learning models are trained based on standard Raman spectra of known substances, and the trained models are used to accurately identify test samples.

[0068] This specification and claims do not use differences in names as a way to distinguish components, but use differences in components' functions as the criteria for distinction. As mentioned in the description and claims throughout the text, "including" is an open-ended term and should be interpreted as "including but not limited to". "Approximately" means that within an acceptable error range, those skilled in the art can solve the technical problem within a certain error range and basically achieve the technical effect. In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the embodiments so that those skilled in the art can implement it with reference to the text of the specification. The equipment or raw materials used in the embodiments can all be obtained from the market.

[0069] instrument

[0070] Raman spectrometer (inVia™), tissue homogenizer (Bio-Gen Series), balance (ME104), water purifier (Arium mini), centrifuge (Centrifuge 5424R), refrigerator (BCD-406WDPD), magnetic stirrer (RCT basic), clean bench (SW-CJ-2F), ultra-low temperature refrigerator (994), low-temperature incubator (PR505750R-CN), pH meter (FE28). The models and manufacturers of the instruments used in the implementation case are detailed in Table 1 below.

[0071] Table 1 Experimental instrument information

[0072]

[0073] Drugs and reagents

[0074] Silver nitrate (Sinopharm Chemical Reagent), sodium citrate (Sinopharm Chemical Reagent), sodium chloride (Sinopharm Chemical Reagent), and Luria-Bertani liquid medium (JingBio, Beijing).

[0075] Specific embodiment: Rapid identification of Salmonella serotypes using the method of the present invention

[0076] 1. Method

[0077] (1) Bacterial collection: All Salmonella strains were obtained from a clinical microbiology laboratory. The obtained Salmonella strains were cultured in Luria-Bertani (LB) liquid medium until the exponential growth phase. After centrifugation, the supernatant and strain pellets were obtained. The obtained pellets were resuspended in 2 mL of deionized water. The concentration of the strains was determined by the plate count method on blood agar plates cultured at 37°C for 24 h.

[0078] (2) Substrate preparation:

[0079] The preparation process of the Raman enhanced substrate is as follows:

[0080] a) 33.72 mg of silver nitrate was weighed and dissolved in 200 ml of deionized water to obtain a 1 mmol / L silver nitrate solution. The solution was heated to boiling using a magnetic stirrer. 8 mL of a 1 wt % sodium citrate solution was then added at once while stirring. The solution was stirred at 650 r / min for 40 min to obtain negatively charged silver nanoparticles having a diameter of <11 nm.

[0081] b) 1 mL of the reaction solution obtained in step a) was centrifuged at 7000 rpm for 7 min. The supernatant was discarded and the resulting precipitate was resuspended in 100 μL of deionized water to obtain a Raman enhancement substrate containing negatively charged silver nanoparticles. The resulting solution was uniformly milky white and stored in the dark at room temperature until use.

[0082] (3) Sample preparation: 5 μL of Raman spectroscopy enhancement substrate solution with negatively charged silver nanoparticles was mixed with an equal amount of Salmonella on a vortex mixer for 30 s. Then, 3 μL of the mixed sample was dropped onto a clean silicon wafer to form a circular spot, and the sample was dried naturally in a safety cabinet for detection.

[0083] (4) SERS detection: The air-dried SERS sample from step (3) was taken for measurement. During the measurement, multiple, multi-site Raman spectral sampling was performed to obtain the Raman spectra of the corresponding ST-typed strain samples, thereby constructing a Raman spectral database of Klebsiella pneumoniae with different ST types.

[0084] The Raman spectrum sampling parameters are as follows: the excitation wavelength of the Raman spectrum is 785 nm, the detector type is a high-sensitivity CCD array, the exposure time is 10 s, and the wavelength scanning range is 500-1800 cm -1 , output data type: txt, laser intensity: 2, spectrum acquisition time: 10s, before spectrum acquisition, use silicon wafer at 520cm -1 The Raman peak at is used as the reference peak for wavenumber calibration, and the dark current is subtracted at the same integration time. 800 sites are selected for Raman spectroscopy detection on each sample.

[0085] (5) Data quality control: The Raman spectra of the different Salmonella strain samples obtained in step (4) were de-peaked, the average Raman spectrum was plotted, and the standard deviation was calculated. The smoothness of the spectrum and the size of the standard error were observed to determine the repeatability of the data.

[0086] (6) Difference evaluation: The spectral data were processed by deconvolution vogit function to improve the distinguishability of Raman spectra between different Salmonella strain samples.

[0087] (7) Characteristic peak analysis:

[0088] The specific process is as follows:

[0089] (a) comparing the spectral characteristic peaks generated by the differential analysis in step (6) to identify the unique characteristic peaks of each Salmonella species;

[0090] (b) For the spectral characteristic peaks generated by the differential analysis in step (6), eight spectral characteristic peaks common to each type of Salmonella were selected, and the signal differences of different bacteria at the same characteristic peaks were analyzed using a box plot;

[0091] (c) For the characteristic peaks of step (a) and step (b), search the literature to find the molecular vibration and biological significance corresponding to each characteristic peak.

[0092] (8) Spectral preprocessing: The spectral data after data quality control in step (5) are normalized using maximum and minimum values ​​to control the spectral signal intensity range at each Raman shift to be within the range of [0, 1].

[0093] (9) Data clustering: The Raman spectral data obtained through quality control in step (5) and preprocessing in step (8) were clustered using orthogonal partial least squares discriminant analysis (OPLS-DA). The Raman spectrum category was determined by comparing the distribution differences of Salmonella of different data types in the characteristic peak coordinate system, and a classification quadrant diagram of the Raman vectors of different Salmonella was formed. Different Salmonella strains were distinguished by observing the clustering of Raman spectral sample points in the classification quadrant diagram. The results of the clustering analysis were evaluated using the three indicators of R2X, R2Y and Q2.

[0094] (10) Data Identification: The Raman spectral data preprocessed in step (8) were automatically analyzed using different machine learning algorithms. The SERS spectral datasets of different Salmonella species were divided into training, validation, and test sets using uniform random sampling. A classifier was trained and tested using 5-fold cross-validation. The trained classifiers were used to classify and label the sample data and store them in a database. All parameters were optimized before training. The machine learning parameter settings are shown in Table 2 below.

[0095] Table 2 Machine learning parameter settings for Salmonella serotypes

[0096]

[0097] (11) Data evaluation: Select test set samples to test the predictive ability, use precision, recall rate, ROC curve and confusion matrix to modify the final judgment model, evaluate the performance of different machine learning algorithms, and select the best judgment model, including:

[0098]

[0099]

[0100]

[0101] Among them: Precision represents accuracy, Recall represents recall rate, AUC represents area under the ROC curve, which is regarded as a performance indicator; TP, FP, TN and FN represent the number of true positives, false positives, true negatives and false negatives respectively, and the confusion matrix is ​​constructed by these four cases.

[0102] 2. Results

[0103] Figure 1 are the average Raman spectra and Raman spectrum deconvolution diagrams of different Salmonella serotypes in the present invention, wherein, Figure 1 The shaded area of ​​each Salmonella serotype spectrum in (A) is the standard error band of each spectrum. The width of the error band indicates the repeatability of the spectral data. From the results, we can see that the error band size of each bacteria is within the acceptable range, indicating that these spectra are highly repeatable. Due to the high similarity between the average spectra, spectral deconvolution can amplify the differences between the spectra. It can be seen that Figure 1 In (B), the deconvolution algorithm deconstructs the most characteristic peaks of Salmonella Enteritidis.

[0104] Figure 2 These are the eight representative spectral characteristic peaks shared by the four different Salmonella species in the present invention. A box plot was used to calculate whether each spectral characteristic peak differed between different bacteria. It can be seen that the P values ​​of the eight characteristic peaks were all less than 0.0001, indicating that there were statistical differences in the peaks between different samples.

[0105] Figure 3 The clustering result of different Salmonella serotypes obtained by clustering the spectral data of 4 Salmonella before and after pretreatment in the form of a scatter plot using the OPLS-DA algorithm in the present invention is shown. Different Salmonella are clustered into different clusters, indicating that the OPLS-DA algorithm can be used to distinguish the spectral data between different Salmonella. Figure 3 Different bacterial species are circled with different dashed circles, and the results of cluster analysis are evaluated using the three indicators R2X, R2Y and Q2. Figure 3 As can be seen, when the spectral data was not normalized, R2X = 0.998, R2Y = 0.774, and Q2 = 0.702. However, after spectral normalization, R2X = 0.942, R2Y = 0.953, and Q2 = 0.951. Each metric score exceeded 0.90, and the difference between R2X and Q2 was within 0.3, indicating that the OPLS-DA algorithm can effectively identify the four different Salmonella species using the normalized spectral data.

[0106] Figure 4 This is the performance gradient diagram of different machine learning algorithm models in the present invention. Figure 5 As can be seen from the figure, the support vector machine scores 100% for every parameter combination. Compared to other learning algorithms, such as AdaBoost, DT, QDA, RF, and XGB, performance degrades in certain parameter combinations. This shows that the SVM algorithm has the highest stability.

[0107] Figure 5 The confusion matrix diagram of different machine learning algorithms in the present invention. The rows of the confusion matrix represent the true values ​​of the SERS signal, the columns represent the predicted values ​​of the model for the SERS signal, and the diagonal represents the score of the model in each ST classification. Figure 5 It can be seen that the SVM model has the best classification performance, with an average recognition accuracy of 99.50%. Only 2% of Salmonella typhimurium were misidentified as Salmonella typhi. Among the other algorithms, the performance of RF was second only to SVM, with 7% of Salmonella typhimurium being misidentified as Salmonella typhi.

[0108] Figure 6 The table below shows the molecular vibrations and biological significance of the unique characteristic peaks of the four different Salmonella species. These differences can be preliminarily considered as the basis for identifying different Salmonella species.

[0109] Figure 7 Performance evaluation results of eight machine learning algorithms in this paper using five evaluation metrics are shown. The SVM model achieved the best performance, with the highest accuracy (Accuracy = 99.38%) and the best stability (5-Fold = 99.98%), demonstrating that SVM is an effective method for distinguishing the SERS spectra of four different Salmonella species. The remaining seven algorithms also achieved accuracy rates exceeding 85%.

[0110] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that it is still possible to modify the technical solutions described in the aforementioned embodiments, or to make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for rapid identification of Salmonella serotypes, characterized in that: It includes the following steps: (1) Bacterial collection: Isolation and culture of different serotypes of Salmonella; (2) Substrate preparation: sodium citrate solution is added to a boiling silver nitrate solution and stirred for reaction. The resulting reaction solution is centrifuged and the resulting precipitate is resuspended in deionized water to obtain a Raman enhancement substrate with negatively charged nanosilver particles. (3) Sample preparation: Each type of Salmonella cultured in step (1) is inoculated into a phosphate buffer solution, and then mixed with 3 to 6 μL of the Raman enhancement substrate prepared in step (2). Then, 2 to 5 μL of the mixed sample is titrated on a clean silicon wafer to form circular spots, and the sample is naturally dried in a safety cabinet to obtain high-sensitivity Raman signal detection samples for different Salmonella species to be tested; (4) SERS detection: The high-sensitivity Raman signal of the Salmonella sample obtained in step (3) is subjected to multiple, multi-site Raman spectrum sampling to obtain corresponding SERS fingerprints of different Salmonella species, thereby constructing a SERS fingerprint database of different Salmonella species; (5) Data quality control: The SERS spectra of different Salmonella species obtained in step (4) were sequentially de-peaked, and the average spectrum and standard deviation of the spectrum were calculated; (6) Difference evaluation: The data after data quality control in step (5) are subjected to quality evaluation, and the differences between the original SERS spectra of different Salmonella are examined by deconvolution spectral analysis; (7) Characteristic peak analysis: Analyze the characteristic peaks of the spectrum obtained by deconvolution fitting in step (6) difference assessment; (8) Spectral preprocessing: performing data normalization on the data after quality control in step (5); (9) Data clustering: The Raman spectral data obtained through quality control in step (5) and pre-processed in step (7) were clustered using orthogonal partial least squares discriminant analysis (OPLS-DA). By comparing the distribution differences of Salmonella spectral samples in the feature space, a classification quadrant diagram of the Raman vectors of different Salmonella was formed. The clustering performance was determined based on the degree of dispersion between different clusters, and the results of the clustering analysis were evaluated using the three indicators of R2X, R2Y, and Q2. (10) Data identification: Using different machine learning algorithms to automatically analyze the Raman spectral data preprocessed in step (7), the spectral data sets of different Salmonella are divided into training sets, validation sets, and test sets by uniform random sampling, and a classifier is trained and tested using K-fold cross validation, where K is any integer from 1 to 10. The sample data are classified and labeled by the trained classifier and stored in a database; (11) Data evaluation: Select test set samples to test the predictive ability, use precision, recall rate, ROC curve and confusion matrix to modify the final judgment model, evaluate the performance of different machine learning algorithms, and select the best judgment model, including: Among them: Precision represents accuracy, Recall represents recall rate, AUC represents area under the ROC curve, which is regarded as a performance indicator; TP, FP, TN and FN represent the number of true positives, false positives, true negatives and false negatives respectively, and the confusion matrix is ​​constructed by these four cases.

2. The method for rapid identification of Salmonella serotypes according to claim 1, wherein In step (1), the process of bacterial isolation and cultivation is as follows: different Salmonella strains are cultured in Luria-Bertani liquid medium to the exponential growth phase, and after centrifugation, the supernatant and strain pellets are obtained. The obtained strain pellets are resuspended in deionized water, and the concentration of the strains is determined by the plate count method on a blood agar plate cultured at 37°C for 24 hours.

3. The method for rapid identification of Salmonella serotypes according to claim 1, wherein In step (2), the preparation method of the Raman enhanced substrate is as follows: (a) A 1 mmol / L silver nitrate solution was heated to boiling, and 8 mL of a 1 wt% sodium citrate solution was added while stirring. The mixture was stirred at 500-800 rpm for 30-50 min to obtain negatively charged silver nanoparticles. (b) The reaction solution obtained in step (a) was centrifuged at 6500-8500 r / min for 6-9 min, the supernatant was discarded, and the resulting precipitate was resuspended in deionized water to obtain a Raman enhancement substrate with negatively charged nanosilver particles.

4. The method for rapid identification of Salmonella serotypes according to claim 1, wherein In step (4), the multi-site Raman spectrum sampling is 700 to 1000 sites; the parameter conditions of the Raman spectrum sampling are: the excitation wavelength of the Raman spectrum is 785nm, the detector type is a high-sensitivity CCD array, the exposure time is 10s, and the wavelength scanning range is 500-1800cm -1 .

5. The method for rapid identification of Salmonella serotypes according to claim 1, wherein In step (5), the Raman spectra of the different Salmonella samples obtained in step (4) are de-peaked, the average Raman spectrum is plotted and the standard deviation is calculated, and the smoothness of the spectrum and the width of the standard error band are observed to determine the repeatability of the data.

6. The method for rapid identification of Salmonella serotypes according to claim 1, characterized in that: In step (6), the SERS spectra of different Salmonella serotype samples are evaluated by using spectral deconvolution to amplify the subtle differences between the characteristic peaks, deconstructing all the characteristic peaks contained in each Salmonella to view the differences between the characteristic peak distributions.

7. The method for rapid identification of Salmonella serotypes according to claim 1, wherein In step (7), the unique spectral characteristic peaks of each Salmonella obtained by deconvolution spectrum fitting in the differential analysis of step (6) are analyzed for comparative analysis, and the differences between the common characteristic peaks of 6 to 10 different Salmonella are selected and analyzed using a box plot, and the molecular vibration and biological significance corresponding to each characteristic peak are queried.

8. The method for rapid identification of Salmonella serotypes according to claim 1, wherein In step (8), a normalization operation is performed on the SERS data after data quality control in step (5). The normalization method is selected according to actual needs and can be maximum and minimum normalization, vector normalization, z-score normalization, decimal calibration normalization, and logarithmic normalization.

9. The method for rapid identification of Salmonella serotypes according to claim 1, characterized in that: In step (9), the OPLS-DA algorithm is used to perform data clustering on the data after data quality control in step (5) and spectral normalization in step (8). The closer the values ​​of the three indicators R2X, R2Y and Q2 are to 1, and the difference between the values ​​of R2X and Q2 does not exceed 0.3, the better the data clustering result is.

10. The method for rapid identification of Salmonella serotypes according to claim 1, characterized in that: In step (10), the machine learning algorithms are support vector machine (SVM), random forest (RF), extreme gradient boosting (XGBoost), adaptive boosting (AdaBoost), decision tree (DT) and quadratic discriminant analysis (QDA); the spectral dataset is divided into training set, validation set and test set in a ratio of 6:2:2; 5-fold cross validation is used for testing, and the K is 5.

Citation Information

Patent Citations

  • A method and kit for identifying Salmonella serotypes

    CN104059977B