A multi-concentration mixed SERS data-based drug qualitative detection method
By constructing a multi-concentration mixed SERS dataset and optimizing the random forest model using the improved NSGA-II algorithm, the problem of decreased classification reliability of low-concentration drug samples was solved, and high-accuracy qualitative drug detection was achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- NANTONG UNIV
- Filing Date
- 2025-04-30
- Publication Date
- 2026-05-19
AI Technical Summary
Existing machine learning techniques suffer from a significant drop in reliability when classifying low-concentration drug samples, making it difficult to meet the requirements for high accuracy. Furthermore, traditional SERS spectral analysis has low accuracy for similar drugs.
A multi-concentration mixed SERS dataset was constructed, and the hyperparameters of the random forest model were optimized by combining the improved NSGA-II algorithm to build an optimized classification model and improve the classification accuracy of low-concentration drug samples.
By using a hybrid dataset and optimizing with the NSGA-II algorithm, the classification accuracy of low-concentration drug samples was significantly improved, enabling reliable classification of low-concentration drugs and enhancing the model's universality and adaptability.
Smart Images

Figure CN120577280B_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of analytical chemistry technology, specifically relating to a method for qualitative detection of drugs based on multi-concentration mixed SERS data. Background Technology
[0002] In recent years, drugs whose main pharmacological effects are central nervous system stimulation or inhibition have become a global social problem due to their addictive nature. According to statistics from the United Nations Office on Drugs and Crime, the number of drug users exceeded 290 million in 2022, with opioids and amphetamine derivatives accounting for the majority of abuse. Notably, with the development of new formulation technologies, these substances have moved beyond traditional solid forms, gradually evolving towards liquid and trace forms, exhibiting low concentrations in biological samples and environmental samples. This evolutionary trend has led to technical bottlenecks in existing rapid screening technologies based on physical morphology recognition, such as insufficient sensitivity and decreased specificity. Particularly in key application scenarios such as drug-impaired driving monitoring, relapse monitoring of drug addicts, and prevention of prescription drug abuse, the development of analytical methods with high sensitivity is urgently needed.
[0003] Currently, gas chromatography-mass spectrometry (GC-MS) and liquid chromatography-mass spectrometry (LC-MS) are the standard methods for drug detection in my country. These methods require significant time, large-scale instruments, and skilled technicians, and are still insufficient for rapid, on-site drug detection. While SERS detection using colloidal gold nanomaterials is fast and highly sensitive, interpreting SERS spectra requires specialized knowledge. When similar drugs have highly similar Raman spectra, the accuracy of manual interpretation drops significantly. Therefore, automated classification of SERS spectra of a large number of drugs based on machine learning methods has become one of the most important detection methods. However, due to the small Raman scattering cross-section of drug molecules, the accuracy of machine learning methods for classifying SERS spectral data drops drastically when dealing with drugs of unknown concentration, making it difficult to meet the industry's high accuracy requirements.
[0004] In classifying analytes using SERS spectroscopy, patent CN119510391A discloses a method for preparing a recessed SERS substrate for detecting exosomes in prostate cancer and its application. This method employs deep learning technology to achieve label-free SERS spectral classification of prostate cancer exosomes. However, this method requires high sample quality and quantity, making it difficult to achieve optimal classification accuracy when the sample size is limited. Summary of the Invention
[0005] This application provides a drug qualitative detection method based on multi-concentration mixed SERS data to solve the technical problem that the reliability of existing machine learning techniques is severely reduced when directly classifying low-concentration drug samples.
[0006] To solve the above-mentioned technical problems, one technical solution adopted in this application is: a method for qualitative detection of drugs based on multi-concentration mixed SERS data, comprising:
[0007] S1. Based on the SERS spectra of the same drug at high and low concentrations, construct a dataset of different drug concentrations;
[0008] S2. Mix the low-concentration and high-concentration SERS spectral data in the dataset to construct a mixed dataset;
[0009] S3. Based on the improved NSGA-II algorithm, the hyperparameters of the random forest model are optimized to construct an optimized classification model;
[0010] S4. Optimize the classification model based on the mixed dataset to classify drugs.
[0011] Furthermore, the method in step S1 includes:
[0012] S11. Sample Preparation: Stock solutions of morphine, methamphetamine, and methadone were diluted at three concentration gradients to obtain nine drug samples. Sodium citrate solution was then added to a preheated chloroauric acid solution and stirred. After cooling, the resulting solution was centrifuged, and the precipitate was redissolved in deionized water to obtain a colloidal SERS substrate containing gold nanoparticles. Next, the drug samples were added to a 96-well plate, and a fixed proportion of colloidal gold nanoparticles was added at a drug-to-gold nanoparticle solution ratio of 1:9. Finally, the mixed solution was incubated for 5–10 minutes to obtain a mixed solution.
[0013] S12. Data Acquisition: The mixed solution was transferred to a quartz tube using a pipette and placed under a Raman spectrometer. It was irradiated with a 785 nm laser with a laser energy of 30 mw and an integration time of 10 s. Ten sets of SERS spectral data were collected for each sample.
[0014] S13. Preprocess the acquired SERS spectral data, including baseline correction, denoising, and normalization.
[0015] S14. Based on nine drug samples, nine original datasets of the same size were constructed: A, B, C, D, E, F, G, H, and I. Dataset A consists of 0.25 ppm morphine SERS data; Dataset B consists of 10 ppm methadone SERS data; Dataset C consists of 12.5 ppm methadone SERS data; Dataset D consists of 0.5 ppm morphine SERS data; Dataset E consists of 20 ppm methadone SERS data; Dataset F consists of 25 ppm methadone SERS data; Dataset G consists of 1 ppm morphine SERS data; Dataset H consists of 40 ppm methadone SERS data; and Dataset I consists of 50 ppm methadone SERS data.
[0016] Among them, the low-concentration datasets refer to datasets A, B, and C; the high-concentration data refers to dataset DI; each of datasets A, B, C, D, E, F, G, H, and I contains N SERS spectral data.
[0017] S15. Datasets A, B, and C are low-concentration drug datasets, and datasets D, E, F, G, H, and I are high-concentration drug datasets. The sample size for each of the nine datasets is set to N. Further, the method in step S2 includes:
[0018] S21. Select any two datasets from datasets A, B, and C to obtain a total of 2N data points. Then, randomly select 0.7N SERS spectral data points from the remaining datasets to form a mixed dataset J. Data points are extracted from datasets A, B, and C.
[0019] S22. Based on the drug types in the remaining dataset, select a dataset of the same type but different concentration from datasets D, E, F, G, H, and I, and randomly select 0.3 N SERS spectral data from the dataset of the same type to add to the mixed dataset J.
[0020] Furthermore, step S3 includes:
[0021] S31. Obtain the F1-score objective function based on the random forest model;
[0022] S32. Set the range of values for hyperparameters in the random forest model;
[0023] S33. Configure the improved NSGA-II algorithm to initialize the population individuals;
[0024] S34. Based on the F1-score objective function, multi-objective optimization is performed in combination with the computational efficiency objective function to obtain the multi-objective function and calculate the fitness of the individual;
[0025] S35. Based on the multi-objective function and individual fitness, obtain the crowding degree of individuals and the hierarchical ranking of individuals in the population;
[0026] S36. Based on the hierarchical sorting of individuals in the population, prioritize the selection of individuals at lower levels with high crowding to obtain the optimal hyperparameters in the optimized classification model.
[0027] Furthermore, the method for obtaining the F1-score objective function in step S31 includes:
[0028] Based on formula (1), the F1-score model is obtained; where formula (1) is:
[0029] (1);
[0030] Wherein, Precision is the proportion of samples predicted as positive that are actually positive, and Recall is the proportion of samples that are actually positive that were correctly predicted.
[0031] Furthermore, the method for obtaining the multi-objective function in step S34 includes:
[0032] Based on formulas (2)-(4), the multi-objective function is obtained; where formulas (2)-(4) are:
[0033] (2);
[0034] (3);
[0035] (4);
[0036] in, and These represent the number of trees and the maximum depth of the trees, respectively.
[0037] Furthermore, the method for obtaining the crowding level of an individual in step S35 includes:
[0038] Based on formula (5), the crowding level of an individual is obtained; where formula (5) is:
[0039] (5);
[0040] in, and These are the target values of adjacent individuals in the target space; and For the goal The maximum and minimum values; Based on the individual The weight function is adaptively adjusted based on the position in the current generation.
[0041] The beneficial effects of this application are as follows: This application combines samples with different concentration ranges, exhibiting greater universality and adaptability compared to traditional classification methods that rely on single-concentration samples, thus meeting various needs. By using a hybrid SERS dataset, this application effectively improves the classification accuracy of machine learning models for low-concentration drug samples, solving the problem of decreased reliability of existing technologies on low-concentration samples. Regarding hyperparameter optimization of the machine learning model, an improved NSGA-II algorithm is used, enabling the simultaneous handling of multiple conflicting objectives to achieve a balance between model accuracy and efficiency. Attached Figure Description
[0042] Figure 1 This is a flowchart illustrating an embodiment of the drug qualitative detection method based on multi-concentration mixed SERS data of this application;
[0043] Figure 2 yes Figure 1 A flowchart illustrating step S3 in one embodiment;
[0044] Figure 3 In Embodiment 1 of this application, a random forest is used as the classifier, and the ROC curves of training set 1 and test set 1 are used.
[0045] Figure 4 In Embodiment 1 of this application, a random forest is used as the classifier, and the ROC curves of training set 2 and test set 2 are used.
[0046] Figure 5 The ROC curve for the first mixed dataset in Embodiment 1 of this application, using random forest as the classifier;
[0047] Figure 6 The ROC curve for the second mixed dataset in Embodiment 1 of this application, using random forest as the classifier;
[0048] Figure 7 The ROC curve for the third mixed dataset in Embodiment 1 of this application, using random forest as the classifier;
[0049] Figure 8 The ROC curve for the fourth mixed dataset in Embodiment 1 of this application, using random forest as the classifier;
[0050] Figure 9 The ROC curve for the fifth mixed dataset in Embodiment 1 of this application, using random forest as the classifier;
[0051] Figure 10The ROC curve for the sixth mixed dataset in Embodiment 1 of this application, using random forest as the classifier;
[0052] Figure 11 In Embodiment 2 of this application, a support vector machine is used as the classifier, and the ROC curves of training set 1 and test set 1 are used.
[0053] Figure 12 In Embodiment 2 of this application, a support vector machine is used as the classifier, and the ROC curves of training set 2 and test set 2 are used.
[0054] Figure 13 The ROC curves for the first mixed dataset in Embodiment 2 of this application, using a support vector machine as the classifier;
[0055] Figure 14 The ROC curves for the second mixed dataset in Embodiment 2 of this application, using a support vector machine as a classifier;
[0056] Figure 15 The ROC curves for the third mixed dataset in Embodiment 2 of this application, using a support vector machine as a classifier;
[0057] Figure 16 The ROC curves for the fourth mixed dataset in Embodiment 2 of this application, using a support vector machine as a classifier;
[0058] Figure 17 The ROC curves for the fifth mixed dataset in Embodiment 2 of this application, using a support vector machine as the classifier;
[0059] Figure 18 The ROC curves for the sixth mixed dataset in Embodiment 2 of this application, using a support vector machine as the classifier;
[0060] Figure 19 The figure shows a comparison of the ROC curves obtained after optimizing the hyperparameters of a random forest using the improved NSGA-II algorithm and the Bayesian optimization algorithm, respectively, in Embodiment 3 of this application. Detailed Implementation
[0061] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to specific embodiments.
[0062] Numerous specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may be practiced in other ways than those described herein, and therefore the invention is not limited to the specific embodiments disclosed in the following specification.
[0063] Example 1
[0064] See Figure 1 , Figure 1 This is a flowchart illustrating an embodiment of the drug qualitative detection method based on multi-concentration mixed SERS data of this application. The method includes:
[0065] S1. Based on the SERS spectra of the same drug at high and low concentrations, construct a dataset of different drug concentrations.
[0066] Specifically, step S1 includes:
[0067] S11. Sample Preparation: Stock solutions of morphine, methamphetamine, and methadone were diluted at three concentration gradients to obtain nine drug samples. Sodium citrate solution was then added to a preheated chloroauric acid solution and stirred. After cooling, the resulting solution was centrifuged, and the precipitate was redissolved in deionized water to obtain a colloidal SERS substrate containing gold nanoparticles. Next, the prepared drug samples (10 sets of each) were added to a 96-well plate, and a fixed proportion of colloidal gold nanoparticles was added at a drug-to-gold nanoparticle solution ratio of 1:9. Finally, the mixed solution was incubated for 5–10 minutes, awaiting the next step of SERS signal measurement.
[0068] Specifically, a total of 90 different samples were prepared, with each drug having three different concentrations and 10 samples for each concentration.
[0069] S12. Data Acquisition: The mixture of drug and gold nanoparticles was transferred to a quartz tube using a pipette and placed under a Raman spectrometer. It was irradiated with a 785 nm laser with a laser energy of 30 mw and an integration time of 10 s. Multiple measurements were performed on each sample, and the spectral data were saved to the computer.
[0070] Specifically, each sample was measured 10 times, resulting in a total of 900 sets of SERS spectral data.
[0071] S13. Preprocess the acquired spectral data, including baseline correction, denoising, and normalization.
[0072] S14. Construct nine original datasets of the same size: A, B, C, D, E, F, G, H, and I. Dataset A consists of 0.25 ppm morphine SERS data; dataset B consists of 10 ppm methadone SERS data; dataset C consists of 12.5 ppm methylamphetamine SERS data; dataset D consists of 0.5 ppm morphine SERS data; dataset E consists of 20 ppm methadone SERS data; dataset F consists of 25 ppm methylamphetamine SERS data; dataset G consists of 1 ppm morphine SERS data; dataset H consists of 40 ppm methadone SERS data; and dataset I consists of 50 ppm methylamphetamine SERS data. Each dataset contains 100 Raman spectra.
[0073] S15. Treat all data in datasets A, B, and C as low-concentration drug datasets, with a sample size of 300. Treat all data in datasets D, E, F, G, H, and I as high-concentration drug datasets, with a sample size of 600.
[0074] S2. Mix the low-concentration and high-concentration SERS spectral data in the dataset to construct a hybrid dataset.
[0075] Specifically, step S2 includes:
[0076] S21. Select any two datasets (e.g., A and B) from datasets A, B, and C, and take a total of 200 data points. Then, randomly select 70 SERS spectral data points from the remaining dataset (e.g., C) to form a mixed dataset J. It must be ensured that data are extracted from datasets A, B, and C.
[0077] S22. Based on the drug types in the remaining dataset (e.g., C) from the 70 data points extracted in S21, select a dataset of the same type as C but with different concentrations from datasets D, E, F, G, H, and I (e.g., F). Randomly select 30 SERS spectral data points from F and add them to the mixed dataset J to enhance the feature diversity of the dataset. Combine the SERS spectral data extracted from S21 and S22 to obtain the mixed dataset.
[0078] S3. Based on the improved NSGA-II algorithm, the hyperparameters of the random forest model are optimized to construct an optimized classification model.
[0079] For details, please refer to Figure 2 Step S3 includes:
[0080] S31. Obtain the F1-score objective function based on the random forest model.
[0081] Specifically, when using a probabilistic model with random forests, the F1-score is used as the primary optimization objective. The F1-score, which comprehensively considers precision and recall, effectively measures the performance of a classification model when handling imbalanced data. Its calculation formula is as follows:
[0082] (1);
[0083] Wherein, Precision is the proportion of samples predicted as positive that are actually positive, and Recall is the proportion of samples that are actually positive that were correctly predicted.
[0084] S32. Set the range of values for hyperparameters in the random forest model.
[0085] Specifically, set the range of values for hyperparameters in the random forest model, such as the maximum depth of the tree, the number of samples required for the minimum split of a node, and the minimum number of samples required for a leaf node.
[0086] S33. Configure the improved NSGA-II algorithm to initialize the population individuals.
[0087] Specifically, first, the population is initialized with a size of 100 to 200 individuals, and Latin hypercube sampling is used to ensure the diversity of the initial population. Then, crossover and mutation operations are configured in the genetic operations. Simulated binary crossover is used for crossover, with a crossover probability of 0.8 to 0.9, and polynomial mutation is used for mutation, with a mutation probability of 0.1 to 0.2. The top 20% of non-dominated solutions are retained in each generation to ensure the propagation of the optimal solution. Termination occurs when the maximum number of iterations is 50 to 100 generations, or when there is no significant improvement in the Pareto front for 10 consecutive generations.
[0088] S34. Based on the F1-score objective function, a multi-objective function is obtained by combining it with the computational efficiency objective function, and the fitness of individuals is calculated.
[0089] Specifically, first, cross-validation (such as 5-fold cross-validation) is used to calculate the F1-score, and training time is used as the computational efficiency objective function. Then, the F1-score objective function is transformed into:
[0090] (2);
[0091] (3);
[0092] This generates a multi-objective optimization function:
[0093] (4);
[0094] in, and These represent the number of trees and the maximum depth of the trees, respectively.
[0095] S35. Based on the multi-objective function and individual fitness, perform multi-objective optimization to obtain the crowding degree of individuals and the hierarchical ranking of individuals in the population.
[0096] First, a non-dominated sort is performed, stratifying the population and prioritizing non-dominated solutions. These individuals perform well in multi-objective optimization, meaning they are not dominated by other individuals across all objectives. Next, the crowding density of each individual is calculated to maintain solution set diversity. Individuals with higher crowding density are prioritized. The crowding density calculation formula is:
[0097] (5);
[0098] in, and These are the target values of adjacent individuals in the target space; and Let i be the maximum and minimum values of the target i. Based on the individual The weight function is adaptively adjusted based on the position in the current generation.
[0099] S36. Based on the hierarchical sorting of individuals in the population, prioritize the selection of individuals at lower levels with high crowding to obtain the optimal hyperparameters in the optimized classification model.
[0100] Specifically, based on the hierarchical ranking of individuals in the population, priority is given to selecting solution individuals from lower levels with high crowding, until the required number of parent individuals is reached. In this process, the selection of individuals follows a combination of non-dominated ranking and crowding, ensuring that the searched solutions are widely distributed in the target space and avoiding over-concentration in a certain region, thus avoiding local optima. Through multiple generations of crossover and mutation operations, the improved NSGA-II continuously optimizes the solution set, eventually converging to a set of Pareto optimal solutions, thus obtaining the optimal hyperparameter combination of the random forest model.
[0101] S4. Optimize the classification model based on the mixed dataset to classify drugs.
[0102] Specifically, step S4 includes:
[0103] S41. A random forest model combined with the NSGA-II algorithm was used to classify the SERS spectral data of drug samples with different concentrations in the experiment.
[0104] See Figure 3As shown, 70% of the data (210 SERS spectral data) was randomly selected from low-concentration drug datasets A, B, and C as training set 1, and the remaining 30% of the data (90 SERS spectral data) was used as test set 1. Under the random forest model combined with the NSGA-II algorithm, the classification accuracy of morphine, methamphetamine, and methadone can reach 61%, 74%, and 64%, respectively.
[0105] See Figure 4 As shown, 70% of the data (210 SERS spectral data) was randomly selected from the high-concentration drug datasets G, H, and I as training set 2, and the remaining 30% of the data (90 SERS spectral data) was used as test set 2. Under the random forest model, the classification accuracy of morphine, methamphetamine, and methadone can reach 88%, 86%, and 90%, respectively.
[0106] S42. Use the optimized machine learning model to classify or predict the SERS spectral data of drug samples in the mixed dataset in the experiment, and evaluate the effect of the model improvement.
[0107] See Figure 5-10 As shown, the ROC curve for mixed data under the random forest model is... Figure 5-10 The composition of the training and test sets in each graph is shown in Table 1:
[0108] Table 1 shows the composition of the dataset under different hybrid methods.
[0109] Dataset composition Figure 5 200 entries are selected from A and B, 70 entries are randomly selected from C, and 30 entries are randomly selected from F. Figure 6 200 entries are selected from A and B, 70 entries are randomly selected from C, and 30 entries are randomly selected from I. Figure 7 200 entries are selected from A and C, 70 entries are randomly selected from B, and 30 entries are randomly selected from E. Figure 8 200 entries are selected from A and C, 70 entries are randomly selected from B, and 30 entries are randomly selected from H. Figure 9 B and C each contain 200 entries. 70 entries are randomly selected from A, and 30 entries are randomly selected from D. Figure 10 200 entries are selected from B and C, 70 entries are randomly selected from A, and 30 entries are randomly selected from G.
[0110] from Figure 5-10 and Figure 2 Comparative analysis shows that by selecting one drug from the low-concentration dataset and incorporating high-concentration data of that drug into the training set, the classifier's classification accuracy for the remaining two low-concentration drugs significantly improves. For example, in... Figure 5 In the training set, we incorporated some data on higher concentrations (25 ppm) of methamphetamine, which improved the classification accuracy for morphine at 0.25 ppm and methadone at 10 ppm, increasing it from 61% to 77% and from 74% to 86%, respectively.
[0111] This application provides a qualitative drug detection method based on a hybrid SERS dataset. This embodiment innovatively employs a dynamic mixing mechanism (S1-S2) of high and low concentration spectral data, effectively solving the problem of low-concentration signals being easily interfered with by noise in traditional detection methods. The resulting classification accuracy far exceeds that of using only low-concentration SERS spectral data, achieving reliable classification of lower-concentration drugs. The NSGA-II algorithm (S3) is introduced to achieve an effective trade-off between classification performance and computational efficiency in the random forest model, improving the overall system optimization level. Actual sample validation shows that by introducing SERS spectral data of a certain drug under higher concentration conditions into the model training set, the classification accuracy of the random forest model for the other two drugs can be significantly improved on the newly constructed hybrid dataset.
[0112] Example 2
[0113] This embodiment provides a comparative experiment using the Support Vector Machine (SVM) algorithm, with the dataset from Embodiment 1 used. In step S3, the NSGA-II algorithm is used to optimize the hyperparameters of the SVM model, obtaining the optimal regularization parameters, kernel coefficients, and kernel type. The optimized SVM model is then used to classify the mixed dataset, and the results are compared with those from the previous classification using only a single-concentration dataset.
[0114] See Figure 11 As shown, if training set 1 and test set 1 are used, the classification accuracy of morphine, methamphetamine and methadone can reach 53%, 72% and 66% respectively under the optimized support vector machine model.
[0115] See Figure 12 As shown, if training set 2 and test set 2 are used, the classification accuracy of morphine, methamphetamine and methadone can reach 83%, 81% and 92% respectively under the optimized support vector machine model.
[0116] See Figure 13-18 As shown, Figures 12 to 18 The ROC curves obtained using the mixed data methods listed in Table 1, based on the support vector machine model, are presented respectively. Specifically, as shown... Figure 12 As shown, if we incorporate some data of higher concentrations (25 ppm) of methamphetamine into the training set, the classification accuracy for morphine at 0.25 ppm and methadone at 10 ppm is improved, from 53% to 65% and from 72% to 87%, respectively.
[0117] To evaluate the effectiveness of this data mixing method on other machine learning algorithms, Example 2 classified three drugs using the same dataset as in Example 1 and obtained their corresponding ROC curves. The comparison shows that regardless of the machine learning model used, the method of mixing multiple concentration data significantly improved the accuracy of classifying low-concentration drugs, demonstrating its good applicability and potential for widespread application in other machine learning algorithms.
[0118] Example 3
[0119] This embodiment conducts a comparative experiment based on a random forest model, using a dataset derived from Embodiment 1. Following step S3, the improved NSGA-II algorithm is used to optimize the hyperparameters of the random forest model, determining key parameters such as the optimal maximum tree depth, the minimum number of samples required for node splits, and the minimum number of samples required for leaf nodes. Subsequently, the optimized random forest model is used to classify the mixed dataset, and the classification results are compared with those of the random forest model after adjusting hyperparameters using Bayesian optimization, thus verifying the effectiveness of the proposed optimization method in improving model performance.
[0120] See Figure 19 Based on training set 1 and test set, after optimizing the hyperparameters of the random forest model using the improved NSGA-II algorithm, the classification accuracy for morphine, methamphetamine, and methadone reached 61%, 74%, and 64%, respectively. In contrast, the classification accuracy of the random forest model after adjusting the hyperparameters using the Bayesian optimization method was only 54%, 63%, and 61%. Through the improved NSGA-II optimization, the classification accuracy of the three types of drugs improved by 7%, 11%, and 3%, respectively, validating the effectiveness of the proposed method in improving model performance.
[0121] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.
Claims
1. A method for qualitative detection of drugs based on multi-concentration mixed SERS data, characterized in that, include: S1. Based on the SERS spectra of the same drug at high and low concentrations, construct a dataset of different drug concentrations; S2. Mix the low-concentration and high-concentration SERS spectral data in the dataset to construct a hybrid dataset; S3. Based on the improved NSGA-II algorithm, the hyperparameters of the random forest model are optimized to construct an optimized classification model; wherein, the method in step S3 includes: S31. Obtain the F1-score objective function based on the random forest model; S32. Set the range of values for hyperparameters in the random forest model; S33. Configure the improved NSGA-II algorithm and initialize the population individuals; S34. Based on the F1-score objective function, multi-objective optimization is performed in combination with the computational efficiency objective function to obtain the multi-objective function and calculate the fitness of the individual; Based on formulas (2)-(4), the multi-objective function is obtained; wherein, formulas (2)-(4) are: (2); (3); (4); in, and These represent the number of trees and the maximum depth of the trees, respectively. S35. Based on the multi-objective function and the individual fitness, obtain the crowding degree of the individual and the hierarchical ranking of the population individuals; Based on formula (5), the crowding degree of the individual is obtained; wherein, formula (5) is: (5); in, and These are the target values of adjacent individuals in the target space; and For the goal The maximum and minimum values; Based on the individual The weight function is adaptively adjusted based on the position in the current generation; S36. Based on the hierarchical sorting of the population individuals, individuals with low levels and high crowding are selected first to obtain the optimal hyperparameters in the optimized classification model; S4. Based on the optimized classification model of the mixed dataset, classify the drugs.
2. The method according to claim 1, characterized in that, The method of step S1 includes: S11. Sample Preparation: Stock solutions of morphine, methamphetamine, and methadone were diluted at three concentration gradients to obtain nine drug samples. Sodium citrate solution was then added to a preheated chloroauric acid solution and stirred. After cooling, the resulting solution was centrifuged, and the precipitate was redissolved in deionized water to obtain a colloidal SERS substrate containing gold nanoparticles. Next, the drug samples were added to a 96-well plate, and a fixed proportion of colloidal gold nanoparticles was added at a drug-to-gold nanoparticle solution ratio of 1:
9. Finally, the mixed solution was incubated for 5–10 minutes to obtain a mixed solution. S12. Data acquisition: The mixed solution was transferred to a quartz tube using a pipette and placed under a Raman spectrometer. It was irradiated with a 785 nm laser with a laser energy of 30 mw and an integration time of 10 s. Ten sets of SERS spectral data were acquired for each sample. S13. Preprocess the acquired SERS spectral data, including baseline correction, denoising, and normalization. S14. Based on the nine drug samples, construct nine original datasets of the same size: A, B, C, D, E, F, G, H, and I. Dataset A consists of 0.25 ppm morphine SERS data; Dataset B consists of 10 ppm methadone SERS data; Dataset C consists of 12.5 ppm methadone SERS data; Dataset D consists of 0.5 ppm morphine SERS data; Dataset E consists of 20 ppm methadone SERS data; Dataset F consists of 25 ppm methadone SERS data; Dataset G consists of 1 ppm morphine SERS data; Dataset H consists of 40 ppm methadone SERS data; and Dataset I consists of 50 ppm methadone SERS data. Among them, the low-concentration datasets refer to datasets A, B, and C; the high-concentration data refers to dataset DI; each of the datasets A, B, C, D, E, F, G, H, and I contains N SERS spectral data. S15. Datasets A, B, and C are low-concentration drug datasets, and datasets D, E, F, G, H, and I are high-concentration drug datasets. The sample size of each of the nine datasets is set to N.
3. The method according to claim 2, characterized in that, The method of step S2 includes: S21. Select any two datasets from datasets A, B, and C, and take a total of 2N data points. Then, randomly select 0.7N SERS spectral data points from the remaining datasets to form a mixed dataset J; wherein data are extracted from datasets A, B, and C. S22. Based on the drug types in the remaining dataset, select a dataset of the same type but different concentration from datasets D, E, F, G, H, and I, and randomly select 0.3N SERS spectral data from the dataset of the same type to add to the mixed dataset J.
4. The method according to claim 1, characterized in that, The method for obtaining the F1-score objective function in step S31 includes: Based on formula (1), the F1-score model is obtained; wherein, formula (1) is: (1); Wherein, Precision is the proportion of samples predicted as positive that are actually positive, and Recall is the proportion of samples that are actually positive that were correctly predicted.