Method for rapidly identifying trace antibiotics in milk
By combining surface-enhanced Raman spectroscopy and machine learning algorithms, the problem of tedious and time-consuming antibiotic detection in milk was solved, and rapid and accurate identification of antibiotics in milk was achieved. In particular, an efficient method for identifying antibiotics in milk was established through sample pretreatment, standard solution preparation, Raman signal detection and data preprocessing.
Patent Information
- Application Number
- CN202410303462.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-18
- Publication Date
- 2025-09-19
AI Technical Summary
The existing technology for detecting antibiotics in milk is cumbersome and time-consuming. In addition, methods such as high-performance liquid chromatography and liquid chromatography-tandem mass spectrometry are expensive and require professional operation, making it difficult to achieve rapid and accurate identification of antibiotics in milk.
Surface-enhanced Raman spectroscopy (SERS) was combined with a machine learning algorithm to establish a method for rapid identification of trace antibiotics in milk through sample pretreatment, standard solution preparation, Raman signal detection, data preprocessing, cluster analysis and machine learning model training.
It achieves rapid and accurate identification of antibiotics in milk, provides a highly sensitive and accurate detection method, and the application of machine learning algorithms improves the efficiency and accuracy of detection.
Smart Images

Figure CN120668627A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of Raman spectroscopy detection, and specifically to a method for rapidly detecting different antibiotics in human milk using Raman spectroscopy, analyzing the distribution of characteristic peaks in the spectrum, and comparing the performance of different machine learning models in detecting antibiotics in human milk. Background Art
[0002] Breast milk, the fluid produced by postpartum women's breasts for feeding infants, is rich in essential nutrients for infant growth, such as protein, lactose, fat, microorganisms, and various immunoglobulins. It is a crucial source of nutrients for child health and survival. However, during breastfeeding, many drugs enter the infant's body through breast milk, including chloramphenicol, tetracycline, doxycycline, and sulfonamides. Long-term, low-level intake of antibiotics by pregnant women can threaten the health and safety of infants and can lead to developmental defects, malformations, and congenital cataracts in the fetus or infant. Therefore, the exposure of breastfeeding women to various drugs warrants attention. Currently, commonly used methods for detecting antibiotics in breast milk include high-performance liquid chromatography (HPLC), liquid chromatography-tandem mass spectrometry (HPLC-MS), capillary electrophoresis (CE), and enzyme-linked immunosorbent assay (ELASA). However, these techniques are cumbersome and time-consuming, require expensive equipment, and require specialized personnel, making the rapid and accurate identification of antibiotics in breast milk challenging. In recent years, surface-enhanced Raman spectroscopy (SERS) combined with machine learning (ML) algorithms has been widely studied for its rapid and accurate detection of antibiotics in milk. However, there are currently few studies combining SERS and ML to detect antibiotics in human breast milk.
[0003] Existing methods for identifying antibiotics in milk mostly rely on observing the distribution of characteristic Raman spectral peaks and simple linear model analysis. Chinese Patent Publication No. CN111551537A discloses a SERS-based method for detecting tetracycline in milk. This method uses nanomaterials combined with linear models to modify the detection limit, but such analysis often falls far short of practical application. However, with the advancement of artificial intelligence (AI) technology, intelligent Raman spectroscopy analysis methods based on machine learning will be the future of rapid Raman spectroscopy identification of antibiotics in human milk. Summary of the Invention
[0004] The purpose of the present invention is to address the deficiencies of the above-mentioned existing technologies and provide a method for rapidly identifying trace antibiotics in milk. The method solves the problems of cumbersome and time-consuming detection process by combining surface-enhanced Raman spectroscopy with machine learning algorithms.
[0005] The technical solution adopted in the present invention is as follows:
[0006] A method for rapidly identifying trace amounts of antibiotics in milk comprises the following steps:
[0007] (1) Sample pretreatment: Take the human milk sample to be tested, place it in a centrifuge and centrifuge to separate the upper fat layer; then add sulfosalicylic acid solution and centrifuge to obtain the supernatant to obtain the pretreated human milk sample;
[0008] (2) preparing standard solutions: weighing doxycycline and tetracycline standards, adding them to the human milk sample obtained in step (1), mixing them to obtain doxycycline and tetracycline standard solutions, and preparing doxycycline and tetracycline human milk standard solutions of different concentrations for later use;
[0009] (3) Sample preparation: adding sodium citrate solution to a silver nitrate solution heated to boiling to react with stirring, centrifuging the resulting reaction solution, and resuspending the resulting precipitate in deionized water to obtain a Raman enhancement substrate with negatively charged nanosilver particles. The human milk antibiotic standard solution samples of different concentrations obtained in step (2) are mixed with the Raman substrate to obtain high-sensitivity Raman signal detection samples of different concentrations;
[0010] (4) SERS detection: The high-sensitivity Raman signal detection samples of different concentrations obtained in step (3) are titrated onto a silicon wafer and air-dried, and Raman spectrum sampling is performed multiple times and at multiple locations to obtain the corresponding Raman spectra of human milk antibiotic samples of different concentrations, thereby constructing a Raman spectrum database of human milk antibiotics of different concentrations;
[0011] (5) Detection limit selection: Calculate the Raman spectral signal intensity of the human milk antibiotics at different concentrations obtained in step (4) to determine the minimum concentration signal detection limit;
[0012] (6) Mixed sample preparation: the two antibiotic samples with the lowest concentration detection limit determined in step (5) are mixed, ultrasonically treated, and mixed with the high-sensitivity Raman enhanced substrate material prepared in step (3) to obtain a SERS detection sample of the human milk mixed antibiotic sample;
[0013] (7) Data preprocessing: performing curve smoothing operation on the Raman spectra of the lowest concentration SERS sample in steps (5) and (6);
[0014] (8) Data clustering: performing cluster analysis on the Raman spectral data obtained after preprocessing in step (7), determining the category of the sample by comparing the differences between different antibiotics, forming a Raman vector classification quadrant, and distinguishing the Raman spectra of different human milk antibiotics by observing the clustering results of the spectral sample points in the classification quadrant;
[0015] (9) Data identification: Use different machine learning algorithms to fit the Raman spectra preprocessed in step (7). The Raman spectrum dataset of all human milk antibiotic samples is divided into training set, validation set, and test set by random sampling. Different model parameter ranges are set, and grid search is used to find the optimal model parameters.
[0016] (10) Model evaluation: The test set samples divided in step (9) are used to test the model capability, and the performance of different machine learning algorithms is evaluated using different evaluation indicators to select the best model, including:
[0017]
[0018]
[0019]
[0020]
[0021] Among them: Accuracy represents accuracy, Precision represents precision, Recall represents recall, and F1 score represents F1 score, which is regarded as a performance indicator; TP, FP, TN, and FN represent the number of true positives, false positives, true negatives, and false negatives, respectively, and the data is judged based on these four situations.
[0022] In the present invention, in step (1), the sample pretreatment process is as follows: all collected fresh milk samples are stored in a refrigerator at -20°C, taken out and returned to room temperature before the experiment begins, 10 ml of milk is centrifuged in a centrifuge at 6000-8000 r / min for 10 minutes to remove the upper layer of human milk fat, then 5% sulfosalicylic acid is added, mixed and allowed to stand for 10 minutes to remove protein and polymorphism, and then centrifuged in a centrifuge at 6000-8000 r / min for 10 minutes, and the supernatant is collected for later use. The centrifuge speed is preferably 7500 r / min.
[0023] In the present invention, in step (2), each antibiotic is mixed with pre-treated human milk supernatant to prepare antibiotic single substance stock solutions, and antibiotic samples with different concentrations are obtained by a stepwise dilution method.
[0024] Preferably, the process for preparing the two antibiotic standard solutions is as follows: 0.9618 mg of tetracycline and 0.8888 mg of doxycycline powder are respectively added to 2 ml of pre-treated human milk supernatant to prepare the two antibiotic single substance stock solutions with a final concentration of 1×10-6 M; then, a stepwise dilution method is used to prepare antibiotic standard solutions with concentrations ranging from 1×10-6 M to 1×10-10 M for use.
[0025] Preferably, the two antibiotic standard solutions are prepared into 1×10 -6 M, 1×10 -7 M, 1×10 -8 M, 1×10 -9 M, 1×10 -10 M concentration solution.
[0026] In the present invention, in step (3), when preparing a high-sensitivity Raman signal detection sample, silver nitrate is dissolved in water and stirred and heated to boiling, wherein the concentration of silver nitrate in the silver nitrate solution is 0.5 to 1.5 mol / L, preferably 1 mol / L. The reducing agent is sodium citrate, and the sodium citrate is prepared into a sodium citrate solution with a concentration of 0.5 to 1.5 wt%, preferably 1 wt%. The antibiotic solutions of different concentrations prepared in step (2) are mixed with the silver nitrate solution in equal amounts.
[0027] In a preferred embodiment, the method for preparing a sample for high-sensitivity Raman signal detection is as follows:
[0028] (a) heating a 1 mol / L silver nitrate solution to boiling, adding 8 mL of a 1 wt % sodium citrate solution while stirring, and stirring at 500-900 rpm for 20-50 min to obtain a negatively charged silver nanoparticle solution; then centrifuging at 6000-8000 rpm for 6-8 min, discarding the supernatant, and resuspending the resulting precipitate in deionized water to obtain a negatively charged silver nanoparticle solution;
[0029] (b) 5 to 30 μL of the negatively charged silver nanoparticle solution obtained in step (a) was mixed with an equal amount of standard antibiotic solutions of different concentrations, titrated onto a clean silicon wafer to form circular spots, and dried naturally in a safety cabinet for detection.
[0030] In a more preferred embodiment, the method for preparing a sample for high-sensitivity Raman signal detection is as follows:
[0031] (a) A 1 mol / L silver nitrate solution was heated to boiling. 8 mL of a 1 wt % sodium citrate solution was added while stirring. The mixture was stirred at 650 rpm for 40 min to obtain a negatively charged silver nanoparticle reaction solution. 1 mL of the negatively charged silver nanoparticle reaction solution was then centrifuged at 7000 rpm for 7 min. The supernatant was discarded and the resulting precipitate was resuspended in 100 μL of deionized water to obtain a negatively charged silver nanoparticle solution, which served as the Raman enhancement substrate.
[0032] (b) 15 μL of the negatively charged silver nanoparticle solution obtained in step (a) was mixed with an equal amount of standard antibiotic solutions of different concentrations, titrated onto a clean silicon wafer to form circular spots, and dried naturally in a safety cabinet for detection.
[0033] In the present invention, in step (4), the multi-site Raman spectrum sampling is 50 to 150 sites; SERS detection uses Anton Paar TM Cora100 handheld Raman spectrometer, Raman spectrum sampling parameters are: Raman spectrum excitation wavelength is 785nm, excitation power is 25mW, spectral resolution is 1nm, spectral wave number resolution is 10cm-1, detection spectrum range is 400-2300cm-1. Among them, the detection spectrum range is preferably 530-1800cm-1. -1 .
[0034] In the present invention, in step (5), the average Raman spectra of antibiotic standard solutions of different concentrations are calculated, the differences in Raman spectral signal intensities at different concentrations are compared, and the minimum concentration detection limit is determined to be 1×10-7M based on the distribution of characteristic peak differences.
[0035] In the present invention, in step (6), the mixed sample preparation process is specifically as follows:
[0036] (a) After determining the minimum detection limit, the two antibiotic solutions were mixed at the minimum detection limit and sonicated for 10 min to obtain a homogenous sample.
[0037] (b) 10-20 μL of the mixed antibiotic solution sample prepared in step (a) was mixed with an equal amount of antibiotic standard solutions of different concentrations, titrated onto a clean silicon wafer to form circular spots, and dried naturally in a safety cabinet for testing.
[0038] In a preferred embodiment, the mixed sample preparation process is as follows:
[0039] (a) After determining the minimum concentration detection limit, 0.5×10 -9 M doxycycline and 0.5 × 10 -9M tetracycline solution was mixed at the minimum concentration detection limit and sonicated for 10 min to obtain a homogenous sample;
[0040] (b) 15 μL of the mixed antibiotic solution sample prepared in step (a) was mixed with an equal amount of antibiotic standard solutions of different concentrations, titrated onto a clean silicon wafer to form circular spots, and dried naturally in a safety cabinet for testing.
[0041] In the present invention, in step (7), spectral data preprocessing methods are used to improve the quality of spectral data collected from different human milk antibiotic samples, such as peak removal, curve smoothing, baseline correction, and normalization.
[0042] Preferably, the collected Raman spectra are subjected to a curve smoothing preprocessing operation using a Savitzky Golay (SG) polynomial filtering algorithm, the fitting order is set to 3, the influence of external noise on the spectral data is removed, the average Raman spectrum and the standard deviation are plotted, and the differences between the antibiotic Raman spectra and the repeatability of the spectral data are checked by observing the distribution of the spectral characteristic peaks and the size of the standard error.
[0043] In the present invention, in step (8), principal component analysis (PCA) and orthogonal partial least squares discriminant analysis (OPLS-DA) are performed on the pre-processed Raman spectrum. The specific operation steps are as follows:
[0044] (a) All three antibiotic spectral data after preprocessing were input into the PCA unsupervised learning algorithm model. Dimensionality reduction and clustering visualization were performed on the spectral data. The principal components PCA1 to PCA7 after dimension reduction were selected as principal components. The clustering results of the Raman spectral principal component sample points in the classification quadrant were observed.
[0045] (b) All three preprocessed antibiotic spectral data were input into the OPLS-DA supervised learning algorithm model, and the clustering results were evaluated using the three indicators R2X, R2Y and Q2. The larger the indicator, the better the model performance, and the difference between R2X and Q2 did not exceed 0.3.
[0046] Preferably, the spectral qualitative analysis uses a PCA algorithm with two principal components, namely PCA1 and PCA2, and determines whether the sample is in a specified interval and the category it belongs to based on the principal components.
[0047] For the present invention, in step (9), the machine learning algorithm is specifically a decision tree (DT), a random forest (RF) and a support vector machine (SVM); the Raman spectroscopy data set is divided into a training set, a validation set and a test set in a ratio of 6:2:2, the training set and the validation set are used for fitting and verification, and the test set data is only used to test the model performance; the grid parameters are used to find the optimal model parameter combination.
[0048] Preferably, in step (9), the data is identified as a machine learning 3-classification model, the machine learning process includes computer storage, a computer processor, and a computer program stored in the computer storage and executable on the computer processor, the computer storage stores a model optimization parameter range, and the computer processor performs the following steps when executing the computer program:
[0049] a) dividing all SERS Raman signals after data preprocessing into different data subsets;
[0050] b) Pre-set different machine learning training parameter ranges, fit the model, and select the optimal parameter combination;
[0051] c) using the trained machine learning model to classify and predict the SERS spectra of three antibiotics to obtain different model files;
[0052] d) Use multiple data evaluations to evaluate the performance of different machine learning algorithms and select the best identification model.
[0053] Step a) involves randomly partitioning a pre-built human milk antibiotic SERS spectral database into a training set, a validation set, and a test set in a ratio of 6:2:2 using a random sampling method. The training set and validation set are used to train the fitting model and verify model performance, respectively. The test data does not participate in model training or verification, but is input into the training model file as unfamiliar data to test the model's ability to identify unknown data types.
[0054] Step b) includes: pre-setting the model parameter ranges that need to be fitted for different machine learning algorithm models, using grid search plus cross-validation to enumerate the model scores of each parameter combination, and selecting the best parameter combination for each model for final model performance comparison.
[0055] Step c) includes: data classification and prediction using the best model parameter combination fitted in step a) to perform model training, comparing the performance of different models on the data set, selecting the best model, and using the trained best classifier to classify and label the human milk antibiotic spectral data and store them in a database.
[0056] Step d) includes: selecting test set samples to test the model's predictive ability, using accuracy, precision, recall, F1 score, 5-fold cross-validation and training time to modify the identification model, evaluating the performance of different machine learning algorithms, and selecting the best identification model.
[0057] The beneficial effects of the present invention are:
[0058] Machine learning combined with the SERS method was used to identify two antibiotic elements and a mixture of two antibiotic elements in human milk, providing a rapid identification and classification method for antibiotics in human milk, providing a fast and effective analysis approach for the application of Raman spectrometers, and reflecting the application advantages of Raman spectrometers.
[0059] Through machine learning, a Raman spectrum library of different human milk antibiotics is established, and the newly collected fingerprint spectra are directly imported into the library for judgment. The comparison and combination of different machine learning methods and evaluation indicators can better ensure the accuracy of the identification results.
[0060] High accuracy: In distinguishing different samples, the optimal model SVM of the present invention has a score of more than 98% for each indicator; good stability: the 5-fold cross-validation score of the SVM model is 98.14%, which is close to the scores of each scoring indicator, indicating that the model is not overfitting and has strong universality. BRIEF DESCRIPTION OF THE DRAWINGS
[0061] Figure 1 Establishing an analysis process combining machine learning algorithms for the Raman spectroscopy fingerprint of human milk antibiotics in the present invention;
[0062] Figure 2 This is a linear relationship diagram of the lowest concentration detection limit and concentration gradient of human milk antibiotics in the present invention;
[0063] Figure 3 This is a graph showing the RSD calculation scores of human milk antibiotics at designated characteristic peaks in the present invention;
[0064] Figure 4 The average Raman spectrum of human milk antibiotics in the present invention and the biological significance diagram of the spectral characteristic peaks;
[0065] Figure 5 This is the clustering result diagram of SERS sample points of different antibiotics obtained by PCA and OPLS-DA data cluster analysis in the present invention;
[0066] Figure 6 This is a performance gradient graph for training different machine learning algorithm models in the present invention;
[0067] Figure 7 This is a radar chart of the scores of different machine learning algorithm models in the present invention. DETAILED DESCRIPTION
[0068] Raman spectroscopy is based on the principle of inelastic scattering. That is, when incident light from a laser light source illuminates a substance, it is scattered by the molecules of the substance. A very small portion of the scattered light has a frequency different from the incident light. The change in the scattered light frequency depends on the structural characteristics of the illuminated substance. Different substances produce scattered light of specific frequencies under the same laser irradiation. Therefore, Raman spectroscopy can be used to achieve fast, simple, repeatable and non-destructive detection of material composition.
[0069] Artificial intelligence technology provides an efficient and accurate implementation for Raman spectroscopy-based material composition detection. Existing Raman spectroscopy machine learning algorithms are oriented toward specific substances to be tested. They transform the material identification problem of Raman spectroscopy into a machine learning classification problem. Machine learning models are trained based on standard Raman spectra of known substances, and the trained models are used to accurately identify test samples.
[0070] This specification and claims do not use differences in names as a way to distinguish components, but use differences in components' functions as the criteria for distinction. As mentioned in the description and claims throughout the text, "including" is an open-ended term and should be interpreted as "including but not limited to". "Approximately" means that within an acceptable error range, those skilled in the art can solve the technical problem within a certain error range and basically achieve the technical effect. In order to make the purpose, technical solutions and advantages of the present invention clearer, the present invention is further described in detail below in conjunction with the embodiments so that those skilled in the art can implement it with reference to the text of the specification. The equipment or raw materials used in the embodiments can all be obtained from the market.
[0071] instrument
[0072] Raman spectrometer (Anton Paar TM The models and manufacturers of the instruments used in the implementation case are detailed in Table 1.
[0073] Table 1 Experimental instrument information
[0074]
[0075] Drugs and reagents
[0076] Silver nitrate (Sinopharm Chemical Reagent), sodium citrate (Sinopharm Chemical Reagent), sodium chloride (Sinopharm Chemical Reagent), doxycycline (Shanghai McLean), tetracycline (Shanghai McLean).
[0077] Specific examples: Rapid identification of trace antibiotics in milk using the method of the present invention
[0078] 1. Method
[0079] (1) Sample pretreatment: The collected fresh milk samples were stored in a refrigerator at -20°C. Before the experiment, they were taken out and returned to room temperature. 10 ml of milk was centrifuged at 6000-8000 r / min for 10 min to remove the upper fat layer of human milk. Then, 5% sulfosalicylic acid was added, mixed, and allowed to stand for 10 min to remove proteins and polymorphisms. The samples were then centrifuged at 6000-8000 r / min for 10 min. The supernatant was collected and set aside.
[0080] (2) Prepare standard solution: add 0.9618 mg of tetracycline and 0.8888 mg of doxycycline powder to 2 ml of pre-treated human milk supernatant to prepare a final concentration of 1×10 -6 M of the two antibiotics; then, the concentration range from 1×10 -6 M~1×10 -10 M antibiotic standard solution, set aside;
[0081] (3) Sample preparation:
[0082] The sample preparation method for high-sensitivity Raman signal detection is as follows:
[0083] (a) 33.72 mg of silver nitrate was weighed and dissolved in 200 mL of deionized water to obtain a 1 mol / L silver nitrate solution. The solution was heated to boiling using a magnetic stirrer. 8 mL of a 1 wt% sodium citrate solution was then added at once while stirring. The reaction was stirred at 650 r / min for 40 min to obtain a negatively charged silver nanoparticle reaction solution with a diameter of <11 nm. Subsequently, 1 mL of the negatively charged silver nanoparticle reaction solution was centrifuged at 7000 r / min for 7 min. The supernatant was discarded and the resulting precipitate was resuspended in 100 μL of deionized water to obtain a negatively charged silver nanoparticle solution, which served as the Raman enhancement substrate.
[0084] (b) 15 μL of the negatively charged silver nanoparticle solution obtained in step (a) was mixed with an equal amount of human milk antibiotic standard solution of different concentrations, titrated onto a clean silicon wafer to form circular spots, and dried naturally in a safety cabinet for detection.
[0085] (4) SERS detection: The high-sensitivity Raman signal detection samples of different concentrations obtained in step (3) were titrated onto a silicon wafer and dried, and Raman spectrum sampling was performed multiple times and at multiple locations to obtain the corresponding Raman spectra of human milk antibiotic samples of different concentrations, thereby constructing a Raman spectrum database of human milk antibiotics of different concentrations; the parameter conditions for Raman spectrum sampling were: the excitation wavelength of the Raman spectrum was 785 nm, the excitation power was 25 mW, the spectral resolution was 1 nm, and the spectral wavenumber resolution was 10 cm -1 The detection spectrum range is 400-2300cm -1 The actual spectral range is 530-1800cm -1 .
[0086] (5) Detection limit selection: Calculate the average Raman spectra of the antibiotic standard solutions with different concentrations obtained in step (4), compare the differences in Raman spectral signal intensities at different concentrations, and determine the minimum concentration detection limit as 1×10 -9 M.
[0087] (6) Preparation of mixed samples: After determining the minimum concentration detection limit in step (5), the two antibiotic solutions were mixed at the minimum concentration detection limit and sonicated for 10 min to obtain a uniform sample; then 15 μL of the prepared mixed solution sample was taken and mixed evenly with 15 μL of negatively charged nanosilver particle substrate solution, and titrated onto a clean silicon wafer to form circular spots, which were then naturally dried in a safety cabinet for SERS signal detection.
[0088] (7) Data preprocessing: The Raman spectra of the lowest concentration SERS samples in steps (5) and (6) are smoothed to remove the influence of external noise on the spectral data. The average Raman spectrum and standard deviation are plotted. The differences between the antibiotic Raman spectra and the repeatability of the spectral data are checked by observing the distribution of the spectral characteristic peaks and the size of the standard error.
[0089] (8) Data clustering:
[0090] The specific steps are as follows:
[0091] a) Input all three pre-processed antibiotic spectral data into the PCA unsupervised learning algorithm model, perform dimensionality reduction clustering visualization on the spectral data, select the principal components PCA1 and PCA2 to determine the classification quadrants, and observe the clustering of spectral sample points in the classification quadrants;
[0092] b) All three preprocessed antibiotic spectral data were input into the OPLS-DA supervised learning algorithm model. By comparing the distribution differences of the spectral sample points of different antibiotics, classification box plots of the Raman vectors of different antibiotics were formed. The clustering results were evaluated using the three indicators R2X, R2Y, and Q2. Larger indicators indicate better model performance, and the difference between R2X and Q2 should not exceed 0.3.
[0093] (9) Data identification: Three machine learning algorithms, decision tree, random forest and support vector machine, were used to identify the concentration of 1×10 -9 The SERS signals of human milk antibiotics were automatically analyzed. The SERS dataset was divided into training, validation, and test sets, and a classifier was trained. The trained classifier assigned and labeled the sample data, which was then stored in a database. All parameters were optimized before training. The machine learning parameter settings are shown in Table 2 below.
[0094] Table 2 Machine learning parameter settings for human milk antibiotic discrimination
[0095]
[0096] (10) Model evaluation: The test set samples divided in step (9) are used to test the model capability. The final discriminant model is modified using accuracy, precision, recall, F1 score, 5-fold cross validation and training time. The performance of different machine learning algorithms is evaluated and the best discriminant model is selected. Specifically, the following are included:
[0097]
[0098]
[0099]
[0100]
[0101] The model's judgment of data is mainly divided into four situations: true positive (TP), false positive (FP), true negative (TN) and false negative (FN).
[0102] 2. Results
[0103] Figure 2 The concentration gradient diagram and linear relationship diagram of two human milk elemental antibiotics show that as the concentration decreases, the spectral intensity decreases linearly, and at a concentration of 1×10-10M, the characteristic peaks of the antibiotic spectrum cannot be well displayed in the full Raman spectrum, and the minimum concentration detection limit is determined to be 1×10-9M.
[0104] Figure 3is the relative standard deviation of the Raman intensity of two human milk elemental antibiotics at the specified characteristic peaks, which is used to measure the relative magnitude of the fluctuation range of antibiotic spectral data; Figure 3 The three bars corresponding to each Sample No. in (A) represent from left to right: 1218 cm-1RSD=3.15%, 1316 cm-1RSD=1.47% and 1628 cm-1RSD=4.98%; Figure 3 The three bars corresponding to each Sample No. in (B) represent from left to right: 1238 cm-1 RSD = 4.01%, 1314 cm-1 RSD = 4.75% and 1614 cm-1 RSD = 4.79%.
[0105] Figure 4 The PCA algorithm partitioned the SERS spectral data of human milk samples containing two single-agent antibiotics and a mixture of the two antibiotics, but some spectral data points overlapped. The OPLS-DA algorithm effectively partitioned the different SERS data sets into distinct clusters, evaluating the clustering results using the R2X = 0.971, R2Y = 0.903, and Q2 = 0.88 metrics. All scores exceeded 0.85, and the differences in R2X and Q2 were within 0.3, indicating that OPLS-DA performed well in clustering the SERS data of the three antibiotics.
[0106] Figure 5 Grid search was used to optimize parameters for different machine learning algorithms. The results showed that the support vector machine scored 98.85% for every parameter combination. Compared to other learning algorithms, both decision trees and random forests experienced performance degradation in certain parameter combinations. This shows that the SVM algorithm has the highest stability.
[0107] Figure 6 A radar chart shows the scores of different machine learning algorithms on SERS spectral data from three antibiotics. The results show that the support vector machine achieved scores exceeding 98% in all areas. Compared to other learning algorithms, the decision tree and random forest accuracy scores were 92.73% and 96.36%, respectively, both lower than the support vector machine. This indicates that the SVM algorithm has the highest accuracy and performs best.
[0108] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that it is still possible to modify the technical solutions described in the aforementioned embodiments, or to make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for rapid identification of trace antibiotics in milk, characterized in that: It includes the following steps: (1) Sample pretreatment: Take the human milk sample to be tested, place it in a centrifuge and centrifuge to separate the upper fat layer; then add sulfosalicylic acid solution and centrifuge to obtain the supernatant to obtain the pretreated human milk sample; (2) preparing standard solutions: weighing doxycycline and tetracycline standards, adding them to the human milk sample obtained in step (1), mixing them to obtain doxycycline and tetracycline standard solutions, and preparing doxycycline and tetracycline human milk standard solutions of different concentrations for later use; (3) Sample preparation: adding sodium citrate solution to a silver nitrate solution heated to boiling to react with stirring, centrifuging the resulting reaction solution, and resuspending the resulting precipitate in deionized water to obtain a Raman enhancement substrate with negatively charged nanosilver particles. The human milk antibiotic standard solution samples of different concentrations obtained in step (2) are mixed with the Raman substrate to obtain high-sensitivity Raman signal detection samples of different concentrations; (4) SERS detection: The high-sensitivity Raman signal detection samples of different concentrations obtained in step (3) are titrated onto a silicon wafer and air-dried, and Raman spectrum sampling is performed multiple times and at multiple locations to obtain the corresponding Raman spectra of human milk antibiotic samples of different concentrations, thereby constructing a Raman spectrum database of human milk antibiotics of different concentrations; (5) Detection limit selection: Calculate the Raman spectral signal intensity of the human milk antibiotics at different concentrations obtained in step (4) to determine the minimum concentration signal detection limit; (6) Mixed sample preparation: the two antibiotic samples with the lowest concentration detection limit determined in step (5) are mixed, ultrasonically treated, and mixed with the high-sensitivity Raman enhanced substrate material prepared in step (3) to obtain a SERS detection sample of the human milk mixed antibiotic sample; (7) Data preprocessing: performing curve smoothing operation on the Raman spectra of the lowest concentration SERS sample in steps (5) and (6); (8) Data clustering: performing cluster analysis on the Raman spectral data obtained after preprocessing in step (7), determining the category of the sample by comparing the differences between different antibiotics, forming a Raman vector classification quadrant, and distinguishing the Raman spectra of different human milk antibiotics by observing the clustering results of the spectral sample points in the classification quadrant; (9) Data identification: Use different machine learning algorithms to fit the Raman spectra preprocessed in step (7). The Raman spectrum dataset of all human milk antibiotic samples is divided into training set, validation set, and test set by random sampling. Different model parameter ranges are set, and grid search is used to find the optimal model parameters. (10) Model evaluation: The test set samples divided in step (9) are used to test the model capability, and the performance of different machine learning algorithms is evaluated using different evaluation indicators to select the best model, including: Among them: Accuracy represents accuracy, Precision represents precision, Recall represents recall, and F1 score represents F1 score, which is regarded as a performance indicator; TP, FP, TN, and FN represent the number of true positives, false positives, true negatives, and false negatives, respectively, and the data is judged based on these four situations.
2. The method for rapid identification of trace antibiotics in milk according to claim 1, characterized in that: In step (1), the sample pretreatment process is as follows: all collected fresh milk samples are stored in a -20°C refrigerator, taken out and restored to room temperature before the start of the experiment, 10 ml of milk is centrifuged in a centrifuge at a speed of 6000-8000 r / min for 10 minutes to remove the upper fat of the human milk, then 5% concentration of sulfosalicylic acid is added, mixed and allowed to stand for 10 minutes to remove proteins and polymorphisms, and then centrifuged in a centrifuge at a speed of 6000-8000 r / min for 10 minutes, and the supernatant is taken for later use.
3. The method for rapid identification of trace antibiotics in milk according to claim 1, characterized in that: In step (2), the standard solutions of the two antibiotics were prepared as follows: 0.9618 mg of tetracycline and 0.8888 mg of doxycycline powder were added to 2 ml of pre-treated human milk supernatant to prepare a final concentration of 1×10 -6 Two antibiotic elemental solutions of M; Subsequently, a serial dilution method was used to prepare the -6 M~1×10 -10 M antibiotic standard solution for later use.
4. The method for rapid identification of trace antibiotics in milk according to claim 1, characterized in that: In step (3), the method for preparing the high-sensitivity Raman signal detection sample is as follows: (a) heating a 1 mol / L silver nitrate solution to boiling, adding 8 mL of a 1 wt % sodium citrate solution while stirring, and stirring at 500-900 rpm for 20-50 min to obtain a negatively charged silver nanoparticle solution; then centrifuging at 6000-8000 rpm for 6-8 min, discarding the supernatant, and resuspending the resulting precipitate in deionized water to obtain a negatively charged silver nanoparticle solution; (b) 5 to 30 μL of the negatively charged silver nanoparticle solution obtained in step (a) was mixed with an equal amount of standard antibiotic solutions of different concentrations, titrated onto a clean silicon wafer to form circular spots, and dried naturally in a safety cabinet for detection.
5. The method for rapid identification of trace antibiotics in milk according to claim 1, characterized in that: In step (4), the number of sites in the multi-site Raman spectrum sampling is 50 to 150 sites; the parameter conditions of the Raman spectrum sampling are: the excitation wavelength of the Raman spectrum is 785nm, the excitation power is 25mW, the spectral resolution is 1nm, and the spectral wavenumber resolution is 10cm -1 The detection spectrum range is 400-2300cm -1 .
6. The method for rapid identification of trace antibiotics in milk according to claim 1, characterized in that: In step (5), the average Raman spectra of antibiotic standard solutions at different concentrations are calculated, and the differences in Raman spectral signal intensities at different concentrations are compared. Based on the distribution of characteristic peak differences, the minimum concentration detection limit is determined to be 1×10-7M.
7. The method for rapid identification of trace antibiotics in milk according to claim 1, characterized in that: In step (6), the mixed sample preparation process is as follows: (a) After determining the minimum detection limit, the two antibiotic solutions were mixed at the minimum detection limit and sonicated for 10 min to obtain a homogenous sample. (b) 10-20 μL of the mixed antibiotic solution sample prepared in step (a) was mixed with an equal amount of antibiotic standard solutions of different concentrations, titrated onto a clean silicon wafer to form circular spots, and dried naturally in a safety cabinet for testing.
8. The method for rapid identification of trace antibiotics in milk according to claim 1, characterized in that: In step (7), the collected Raman spectra are subjected to a curve smoothing preprocessing operation to remove the influence of external noise on the spectral data, and the average Raman spectrum and standard deviation are plotted. By observing the distribution of the spectral characteristic peaks and the size of the standard error, the differences between the antibiotic Raman spectra and the repeatability of the spectral data can be checked.
9. The method for rapid identification of trace antibiotics in milk according to claim 1, characterized in that: In step (8), principal component analysis (PCA) and orthogonal partial least squares discriminant analysis (OPLS-DA) are performed on the pre-processed Raman spectrum. The specific operation steps are as follows: (a) All three antibiotic spectral data after preprocessing were input into the PCA unsupervised learning algorithm model. Dimensionality reduction and clustering visualization were performed on the spectral data. The principal components PCA1 to PCA7 after dimension reduction were selected as principal components. The clustering results of the Raman spectral principal component sample points in the classification quadrant were observed. (b) All three preprocessed antibiotic spectral data were input into the OPLS-DA supervised learning algorithm model, and the clustering results were evaluated using the three indicators R2X, R2Y and Q2. The larger the indicator, the better the model performance, and the difference between R2X and Q2 did not exceed 0.
3.
10. The method for rapid identification of trace antibiotics in milk according to claim 1, characterized in that: In step (9), the machine learning algorithm is specifically a decision tree (DT), a random forest (RF) and a support vector machine (SVM); the Raman spectroscopy data set is divided into a training set, a validation set and a test set in a ratio of 6:2:2, the training set and the validation set are used for fitting and verification, and the test set data is only used to test the model performance; the grid parameters are used to find the optimal model parameter combination.
Citation Information
Patent Citations
Method for detecting tetracycline in milk based on surface enhanced Raman technology
CN111551537A
Cited By
Rapid qualitative and quantitative analysis method for quinolone antibiotics in soil
CN122042636A