Drug qualitative detection method based on multi-concentration mixed SERS (Surface Enhanced Raman Scattering) data
By constructing a multi-concentration mixed SERS dataset and optimizing the random forest model using the improved NSGA-II algorithm, the problem of insufficient classification of low-concentration drug samples was solved, and high accuracy classification of drugs was achieved, and the adaptability and accuracy of the model was improved.
Patent Information
- Application Number
- CN202510564058.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-30
- Publication Date
- 2025-09-02
- Estimated Expiration
- 2045-04-30
AI Technical Summary
The reliability of existing machine learning technologies has severely decreased when classifying low-concentration drug samples, making it difficult to meet the needs of high accuracy, and traditional SERS spectral judgments lack the accuracy of similar drugs.
A multi-concentration mixed SERS data set was constructed, and the random forest model was hyperparameter-optimized using the improved NSGA-II algorithm, and classified it with SERS spectral data of different concentrations.
The accuracy of the classification of low-concentration drug samples by machine learning models is improved, and the reliable classification of multiple drugs is achieved, the problem of reduced reliability of low-concentration samples is solved, and the universality and adaptability of the model is improved.
Smart Images

Figure CN120577280A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the technical field of analytical chemistry, and specifically relates to a qualitative drug detection method based on multi-concentration mixed SERS data. Background Art
[0002] In recent years, drugs whose main pharmacological effects are stimulation or inhibition of the central nervous system have become a global social problem due to their addictive characteristics. According to statistics from the United Nations Office on Drugs and Crime, the number of drug users in 2022 has exceeded 290 million, of which opioids and amphetamine derivatives account for the main types of abuse. It is worth noting that with the development of new formulation technologies, such substances have broken through the traditional solid formulation form and gradually evolved towards liquidization and trace amounts, showing low-concentration occurrence characteristics in biological samples and environmental samples. This evolutionary trend has caused the existing rapid screening technology based on physical morphology recognition to face technical bottlenecks such as insufficient sensitivity and decreased specificity. In particular, in key application scenarios such as drug-related driving supervision, relapse monitoring of drug addicts, and prevention and treatment of prescription drug abuse, there is an urgent and practical need to develop analytical methods with high-sensitivity detection capabilities.
[0003] Currently, gas chromatography-mass spectrometry and liquid chromatography-mass spectrometry are the standard methods for drug detection in my country. These methods require a significant amount of time, large instruments, and specialized technicians to operate, and are still insufficient for on-site, rapid drug detection. While SERS detection using colloidal gold nanomaterials is both fast and highly sensitive, the interpretation of SERS spectra requires specialized technical knowledge. When highly similar Raman spectra of similar drugs are present, the accuracy of manual judgment decreases significantly. Therefore, automated classification of SERS spectra of a large number of drugs based on machine learning has become one of the most important detection methods. However, due to the small Raman scattering cross-section of drug molecules, the accuracy of classifying SERS spectral data using machine learning methods decreases significantly when faced with unknown drug concentrations, making it difficult to meet the industry's high accuracy requirements.
[0004] Patent CN119510391A discloses a method for preparing and applying a recessed SERS substrate for prostate cancer exosome detection. This method utilizes deep learning technology to achieve label-free SERS spectral classification of prostate cancer exosomes. However, this method requires high sample quality and quantity, making it difficult to achieve high classification accuracy when the sample size is limited. Summary of the Invention
[0005] The present application provides a qualitative drug detection method based on multi-concentration mixed SERS data to solve the technical problem that the reliability of existing machine learning technology is seriously reduced when directly classifying lower concentration drug samples.
[0006] To solve the above technical problems, the present application adopts a technical solution: a qualitative drug detection method based on multi-concentration mixed SERS data, comprising:
[0007] S1. Construct a dataset of drug concentrations based on the SERS spectra of the same drug at high and low concentrations.
[0008] S2. Mixing the low-concentration and high-concentration SERS spectral data in the dataset to construct a mixed dataset;
[0009] S3. Based on the improved NSGA-II algorithm, hyperparameter optimization of the random forest model was performed to construct an optimized classification model.
[0010] S4. Optimize the classification model based on the mixed dataset to classify drugs.
[0011] Furthermore, the method of step S1 includes:
[0012] S11. Sample preparation: Morphine, methamphetamine, and methadone stock solutions were diluted according to three concentration gradients to obtain nine drug samples. A sodium citrate solution was then added to a boiling chloroauric acid solution and stirred. The resulting solution was cooled and centrifuged. The precipitate obtained by centrifugation was then redissolved in deionized water to obtain a colloidal SERS substrate containing gold nanoparticles. The drug samples were then added to a 96-well plate, and a fixed ratio of colloidal gold nanoparticles was added, with the drug and gold nanoparticle solution at a ratio of 1:9. Finally, the mixed solution was incubated for 5–10 minutes to obtain a mixed solution.
[0013] S12. Data Acquisition: The mixed solution was transferred to a quartz tube using a pipette and placed under a Raman spectrometer. Irradiation was performed using a 785 nm laser with a power of 30 mW and an integration time of 10 s. Ten sets of SERS spectra were collected for each sample.
[0014] S13. Preprocessing the collected SERS spectral data, including baseline correction, denoising, and normalization;
[0015] S14. Based on nine drug samples, nine original data sets A, B, C, D, E, F, G, H, and I of equal size were constructed; data set A consisted of SERS data of morphine at 0.25 ppm; data set B consisted of SERS data of methadone at 10 ppm; data set C consisted of SERS data of methadone at 12.5 ppm; data set D consisted of SERS data of morphine at 0.5 ppm; data set E consisted of SERS data of methadone at 20 ppm; data set F consisted of SERS data of methadone at 25 ppm; data set G consisted of SERS data of morphine at 1 ppm; data set H consisted of SERS data of methadone at 40 ppm; and data set I consisted of SERS data of methadone at 50 ppm.
[0016] The low-concentration data set refers to data sets A, B, and C; the high-concentration data set refers to data set DI; each of data sets A, B, C, D, E, F, G, H, and I contains N SERS spectra data;
[0017] S15. Datasets A, B, and C are low-concentration drug datasets, and data sets D, E, F, G, H, and I are high-concentration drug datasets. The sample capacity of each of the nine datasets is set to N.
[0018] Furthermore, the method of step S2 includes:
[0019] S21. Select 2N data points from each of the two datasets A, B, and C, and then randomly select 0.7N SERS spectral data points from the remaining dataset to form a mixed dataset J; wherein data points are all extracted from datasets A, B, and C;
[0020] S22. Based on the drug types in the remaining datasets, select a similar dataset from datasets D, E, F, G, H, and I that is of the same type but with a different concentration as the remaining datasets, and randomly select 0.3N SERS spectral data from the same dataset and add it to the mixed dataset J.
[0021] Further, step S3 includes:
[0022] S31. Based on the random forest model, obtain the F1-score objective function;
[0023] S32. Set the value range of hyperparameters in the random forest model;
[0024] S33. Configure the improved NSGA-II algorithm and initialize the population individuals;
[0025] S34. Based on the F1-score objective function, combined with the computational efficiency objective function, multi-objective optimization is performed to obtain the multi-objective function and calculate the fitness of the individual;
[0026] S35. Based on the multi-objective function and individual fitness, obtain the individual crowding degree and the hierarchical ranking of the population individuals;
[0027] S36. Based on the hierarchical sorting of population individuals, individuals at the lower levels with high crowding are prioritized to obtain the optimal hyperparameters in the optimized classification model.
[0028] Furthermore, the method for obtaining the F1-score objective function in step S31 includes:
[0029] Based on formula (1), the F1-score model is obtained; where formula (1) is:
[0030]
[0031] Among them, Precision is the proportion of samples predicted to be positive that are actually positive, and Recall is the proportion of samples that are correctly predicted to be positive.
[0032] Furthermore, the method for obtaining the multi-objective function in step S34 includes:
[0033] Based on formulas (2)-(4), a multi-objective function is obtained; wherein formulas (2)-(4) are:
[0034] P train =1-F1_score(2);
[0035] T train =n estimate ×max dpth (3);
[0036] multi object =[P train ,T train ] (4);
[0037] Among them, n estimate and max dpth Represents the number of trees and the maximum depth of the tree respectively.
[0038] Furthermore, the method for obtaining the individual congestion degree in step S35 includes:
[0039] Based on formula (5), the individual crowding degree is obtained; where formula (5) is:
[0040]
[0041] Among them, f i+1 and f i-1 are the target values of adjacent individuals in the target space; and is the maximum and minimum value of target i; w i (x) is a weight function that is adaptively adjusted based on the position of individual x in the current generation.
[0042] The beneficial effects of this application are as follows: Compared with traditional classification methods that rely on single-concentration samples, this application combines samples with different concentration ranges, and has greater universality and adaptability, and can meet various needs. By mixing SERS datasets, this application effectively improves the classification accuracy of machine learning models for lower-concentration drug samples, solving the problem of reduced reliability of existing technologies for low-concentration samples. In terms of hyperparameter optimization of machine learning models, the improved NSGA-II algorithm is used for model hyperparameter optimization, which can simultaneously handle multiple conflicting objectives to achieve a balance between model accuracy and efficiency. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 1 is a flow chart of an embodiment of a method for qualitative drug detection based on multi-concentration mixed SERS data of the present application;
[0044] Figure 2 yes Figure 1 A flow chart of an embodiment of step S3 in FIG.
[0045] Figure 3 This is an ROC curve diagram of Example 1 of the present application using random forest as the classifier, training set 1 and test set 1;
[0046] Figure 4 This is the ROC curve diagram of training set 2 and test set 2 using random forest as the classifier in Example 1 of the present application;
[0047] Figure 5 This is an ROC curve diagram for the first mixed data set using random forest as the classifier in Example 1 of the present application;
[0048] Figure 6 This is an ROC curve diagram for the second mixed data set using random forest as the classifier in Example 1 of the present application;
[0049] Figure 7 This is the ROC curve diagram for the third mixed data set using random forest as the classifier in Example 1 of the present application;
[0050] Figure 8 This is the ROC curve diagram for the fourth mixed data set using random forest as the classifier in Example 1 of the present application;
[0051] Figure 9 This is the ROC curve diagram for the fifth mixed data set using random forest as the classifier in Example 1 of the present application;
[0052] Figure 10 This is the ROC curve diagram for the sixth mixed data set using random forest as the classifier in Example 1 of the present application;
[0053] Figure 11 The ROC curves of training set 1 and test set 1 are shown in Example 2 of the present application using a support vector machine as a classifier.
[0054] Figure 12 In Example 2 of the present application, the support vector machine is used as the classifier, and the ROC curves of the training set 2 and the test set 2 are adopted;
[0055] Figure 13 2 are ROC curves of the first mixed data set using the support vector machine as the classifier in Example 2 of the present application;
[0056] Figure 14 2 are ROC curves of the second mixed data set using the support vector machine as the classifier in Example 2 of the present application;
[0057] Figure 15 2 are ROC curves of the third mixed data set using the support vector machine as the classifier in Example 2 of the present application;
[0058] Figure 16 2 are ROC curves of the fourth mixed data set using the support vector machine as the classifier in Example 2 of the present application;
[0059] Figure 17 2 are ROC curves of the fifth mixed data set using the support vector machine as the classifier in Example 2 of the present application;
[0060] Figure 18 2 are ROC curves of the sixth mixed data set using the support vector machine as the classifier in Example 2 of the present application;
[0061] Figure 19 This is a comparison of the ROC curves obtained after random forest hyperparameter optimization using the improved NSGA-II algorithm and the Bayesian optimization algorithm in Example 3 of the present application. DETAILED DESCRIPTION
[0062] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is further described in detail below with reference to specific embodiments.
[0063] In the following description, many specific details are set forth to facilitate a full understanding of the present invention. However, the present invention may also be implemented in other ways different from the description. Therefore, the present invention is not limited to the specific embodiments disclosed in the following specification.
[0064] Example 1
[0065] See Figure 1 , Figure 1 FIG. 1 is a flow chart of an embodiment of a method for qualitative drug detection based on multi-concentration mixed SERS data of the present application, the method comprising:
[0066] S1. Based on the SERS spectra of the same drug at high and low concentrations, a dataset of different drug concentrations was constructed.
[0067] Specifically, step S1 includes:
[0068] S11. Sample preparation: Take the stock solutions of three drugs, morphine, methamphetamine, and methadone, and dilute them according to three concentration gradients to obtain 9 drug samples; then add sodium citrate solution to the boiling chloroauric acid solution and stir to react. After cooling, the resulting solution is centrifuged and the precipitate obtained by centrifugation is redissolved in deionized water to obtain a colloidal SERS substrate containing gold nanoparticles; then, add the prepared drug samples (10 groups of each) into a 96-well plate, and add a fixed ratio of colloidal gold nanoparticles according to the ratio of 1:9 between the drug and gold nanoparticle solution; finally, incubate the mixed solution for 5 to 10 minutes and wait for the next step of SERS signal measurement.
[0069] Specifically, a total of 90 different sets of samples were prepared, with each drug having three different concentrations, and 10 sets of samples for each concentration.
[0070] S12. Data acquisition: Use a pipette to transfer the mixture of drugs and gold nanoparticles into a quartz tube and place it under a Raman spectrometer. Irradiate with a 785nm laser with a laser energy of 30mW and an integration time of 10s. Perform multiple measurements on each sample and save the spectral data to a computer.
[0071] Specifically, each sample must be measured 10 times, obtaining a total of 900 sets of SERS spectral data.
[0072] S13. Preprocess the collected spectral data, including baseline correction, denoising, and normalization.
[0073] S14. Construct nine original datasets of equal size: A, B, C, D, E, F, G, H, and I. Dataset A consists of SERS data from morphine at 0.25 ppm; Dataset B consists of SERS data from methadone at 10 ppm; Dataset C consists of SERS data from methadone at 12.5 ppm; Dataset D consists of SERS data from morphine at 0.5 ppm; Dataset E consists of SERS data from methadone at 20 ppm; Dataset F consists of SERS data from methadone at 25 ppm; Dataset G consists of SERS data from morphine at 1 ppm; Dataset H consists of SERS data from methadone at 40 ppm; and Dataset I consists of SERS data from methadone at 50 ppm. Each dataset contains 100 Raman spectra.
[0074] S15. Consider all the data in datasets A, B, and C as low-concentration drug datasets, with a sample size of 300. Consider all the data in datasets D, E, F, G, H, and I as high-concentration drug datasets, with a sample size of 600.
[0075] S2. Mix the low-concentration and high-concentration SERS spectral data in the data set to construct a mixed data set.
[0076] Specifically, step S2 includes:
[0077] S21. Select 200 data points from each of the two datasets A, B, and C (e.g., A and B), and then randomly select 70 SERS spectral data points from the remaining dataset (e.g., C) to form a mixed dataset J; among them, it must be ensured that data points are all present in datasets A, B, and C.
[0078] S22. Based on the drug type in the remaining dataset (e.g., C) from which 70 data points were extracted in S21, select a similar dataset (e.g., F) from datasets D, E, F, G, H, and I, containing the same drug type as C but at a different concentration. Then, randomly select 30 SERS spectral data points from this dataset and add them to mixed dataset J to enhance the feature diversity of the dataset. The SERS spectral data points extracted in S21 and S22 are then mixed to generate a mixed dataset.
[0079] S3. Based on the improved NSGA-II algorithm, the hyperparameters of the random forest model are optimized to construct an optimized classification model.
[0080] For details, see Figure 2 , step S3 includes:
[0081] S31. Obtain the F1-score objective function based on the random forest model.
[0082] Specifically, when using the random forest probability model, the F1-score is used as the main optimization target. The F1-score is a comprehensive consideration of precision and recall, effectively measuring the performance of the classification model when dealing with imbalanced data. Its calculation formula is:
[0083]
[0084] Among them, Precision is the proportion of samples predicted to be positive that are actually positive, and Recall is the proportion of samples that are correctly predicted to be positive.
[0085] S32. Set the value range of hyperparameters in the random forest model.
[0086] Specifically, set the value range of hyperparameters in the random forest model, such as the maximum depth of the tree, the minimum number of samples required for node splitting, the minimum number of samples for leaf nodes, etc.
[0087] S33. Configure the improved NSGA-II algorithm and initialize the population individuals.
[0088] Specifically, the population is initialized with a size of 100 to 200 individuals, and Latin hypercube sampling is used to ensure initial population diversity. The genetic operations are then set for crossover and mutation, using simulated binary crossover with a crossover probability of 0.8 to 0.9, and polynomial mutation with a mutation probability of 0.1 to 0.2. The top 20% of non-dominated solutions are retained in each generation to ensure optimal solution propagation. The system is terminated after a maximum number of iterations of 50 to 100 generations, or if there is no significant improvement on the Pareto frontier for 10 consecutive generations.
[0089] S34. Based on the F1-score objective function, combined with the computational efficiency objective function, a multi-objective function is obtained and the fitness of the individual is calculated.
[0090] Specifically, first, use cross-validation (such as 5-fold cross-validation) to calculate the F1-score, and combine it with the training time as the computational efficiency objective function. Then, the F1-score objective function is transformed into:
[0091] P train =1-F1_score(2);
[0092] T train =n estimate ×max depth (3);
[0093] Thus a multi-objective optimization function is generated:
[0094] multi object =[P train ,T train ] (4);
[0095] Among them, n estimate and max depth Represents the number of trees and the maximum depth of the tree respectively.
[0096] S35. Based on the multi-objective function and individual fitness, multi-objective optimization is performed to obtain the individual crowding degree and the hierarchical sorting of the population individuals.
[0097] First, a non-dominated sort is performed to hierarchically sort the individuals in the population and prioritize non-dominated solutions. These individuals perform well in multi-objective optimization, meaning they are not dominated by other individuals on all objectives. Next, the crowding degree of each individual is calculated to maintain the diversity of the solution set. The density of individuals in the objective space is calculated, and individuals with a higher crowding degree are prioritized. The crowding degree calculation formula is:
[0098]
[0099] Among them, f i+1 and f i-1 are the target values of adjacent individuals in the target space; and is the maximum and minimum value of target i; w i (x) is a weight function that is adaptively adjusted based on the position of individual x in the current generation.
[0100] S36. Based on the hierarchical sorting of population individuals, individuals at the lower levels with high crowding are prioritized to obtain the optimal hyperparameters in the optimized classification model.
[0101] Specifically, based on the results of the hierarchical sorting of individuals in the population, low-level, high-crowding solution individuals are prioritized until the required number of parent individuals is reached. During this process, individual selection follows a combination of non-dominated sorting and crowding, ensuring that the searched solutions are widely distributed in the target space, avoiding excessive concentration in a single region and thus avoiding local optimality. Through multiple generations of crossover and mutation operations, the improved NSGA-II continuously optimizes the solution set, ultimately converging to a set of Pareto-optimal solutions, i.e., the optimal hyperparameter combination for the random forest model.
[0102] S4. Optimize the classification model based on the mixed dataset to classify drugs.
[0103] Specifically, step S4 includes:
[0104] S41. A random forest model combined with the NSGA-II algorithm was used to classify the SERS spectral data of drug samples with different concentrations in the experiment.
[0105] See also Figure 3 As shown, 70% of the data (210 SERS spectral data) were randomly selected from the low-concentration drug datasets A, B, and C as training set 1, and the remaining 30% of the data (90 SERS spectral data) were used as test set 1. Under the random forest model combined with the NSGA-II algorithm, the classification accuracy of morphine, methamphetamine, and methadone reached 61%, 74%, and 64%, respectively.
[0106] See also Figure 4 As shown, 70% of the data (210 SERS spectral data) were randomly selected from the high-concentration drug datasets G, H, and I as training set 2, and the remaining 30% of the data (90 SERS spectral data) were used as test set 2. Under the random forest model, the classification accuracy of morphine, methamphetamine, and methadone can reach 88%, 86%, and 90%, respectively.
[0107] S42. Use the optimized machine learning model to classify or predict the SERS spectral data of drug samples in the mixed dataset in the experiment and evaluate the effect of the model improvement.
[0108] See also Figure 5-10 As shown, the ROC curve under mixed data under the random forest model, Figure 5-10 The composition of the training set and test set in each figure is shown in Table 1:
[0109] Table 1 shows the composition of the dataset under different mixing methods
[0110] Dataset composition Figure 5 Take 200 items from A and B, randomly select 70 items from C, and randomly select 30 items from F. Figure 6 Take 200 items from A and B, randomly select 70 items from C, and randomly select 30 items from I. Figure 7 200 items are randomly selected from A and C, 70 items are randomly selected from B, and 30 items are randomly selected from E. Figure 8 Take 200 items from A and C, randomly select 70 items from B, and randomly select 30 items from H. Figure 9 200 items are randomly selected from B and C, 70 items are randomly selected from A, and 30 items are randomly selected from D. Figure 10 200 items are randomly selected from B and C, 70 items are randomly selected from A, and 30 items are randomly selected from G.
[0111] from Figure 5-10 and Figure 2 A comparative analysis shows that if one drug is selected from the low-concentration data set and its high-concentration data is incorporated into the training set, the classification accuracy of the classifier for the remaining two low-concentration drugs will be significantly improved. Figure 5 In the training set, we incorporated some data of higher concentration (25ppm) of methamphetamine into the training set, and the classification accuracy of 0.25ppm of morphine and 10ppm of methadone were improved, from 61% to 77% and from 74% to 86%, respectively.
[0112] The present application provides a qualitative drug detection method based on a mixed SERS data set. This embodiment innovatively adopts a dynamic mixing mechanism (S1-S2) of high and low concentration spectral data, which effectively solves the problem that low concentration signals are easily interfered by noise in traditional detection methods. The classification accuracy obtained far exceeds the classification accuracy of using only low concentration SERS spectral data, so as to achieve reliable classification of drugs with lower concentrations. The NSGA-II algorithm (S3) is introduced to achieve an effective trade-off between classification performance and computational efficiency of the random forest model, thereby improving the optimization level of the overall system. Actual sample verification shows that as long as the SERS spectral data of a certain drug under higher concentration conditions are introduced into the model training set, the classification accuracy of the random forest model for the remaining two drugs can be significantly improved on the newly constructed mixed data set.
[0113] Example 2
[0114] This example provides a comparative experiment using a support vector machine algorithm, using the dataset described in Example 1. In step S3, the NSGA-II algorithm is used to optimize the hyperparameters of the support vector machine model to obtain the optimal regularization parameter, kernel function coefficient, and kernel function type. The optimized support vector machine model is then used to classify the mixed dataset, which is then compared with the classification results of a dataset previously classified using only a single concentration.
[0115] See also Figure 11 As shown in FIG, if training set 1 and test set 1 are used, the classification accuracy of morphine, methamphetamine and methadone can reach 53%, 72% and 66% respectively under the optimized support vector machine model.
[0116] See also Figure 12 As shown in FIG, if training set 2 and test set 2 are used, the classification accuracy of morphine, methamphetamine and methadone can reach 83%, 81% and 92% respectively under the optimized support vector machine model.
[0117] See also Figure 13-18 As shown, Figures 12 to 18 The ROC curves obtained based on the support vector machine model and the mixed data listed in Table 1 are shown respectively. Specifically, Figure 12 As shown in Figure 2, if we incorporate some data of higher concentration (25ppm) methamphetamine into the training set, the classification accuracy of 0.25ppm morphine and 10ppm methadone are improved, from 53% to 65% and from 72% to 87%, respectively.
[0118] To evaluate the effectiveness of this data blending method for other machine learning algorithms, Example 2 classified three drugs using the same dataset as Example 1 and obtained corresponding ROC curves. Comparison revealed that regardless of the machine learning model used, the accuracy of low-concentration drug classification was significantly improved by blending multiple concentration data, demonstrating the applicability of this method and its widespread application to other machine learning algorithms.
[0119] Example 3
[0120] This example conducted a comparative experiment based on a random forest model, using the dataset from Example 1. According to step S3, the hyperparameters of the random forest model were optimized using the improved NSGA-II algorithm, determining key parameters such as the optimal maximum tree depth, the minimum number of samples required for minimum node splitting, and the minimum number of samples for leaf nodes. Subsequently, the mixed dataset was classified based on the optimized random forest model, and the classification results were compared with those of the random forest model after adjusting the hyperparameters using the Bayesian optimization method, thereby verifying the effectiveness of the proposed optimization method in improving model performance.
[0121] See also Figure 19 Based on training set 1 and test set, after optimizing the hyperparameters of the random forest model using the improved NSGA-II algorithm, the classification accuracy for morphine, methamphetamine, and methadone reached 61%, 74%, and 64%, respectively. In contrast, the random forest model, after adjusting the hyperparameters using the Bayesian optimization method, achieved classification accuracies of only 54%, 63%, and 61%. Through the improved NSGA-II optimization, the classification accuracy of the three drug classes increased by 7%, 11%, and 3%, respectively, validating the effectiveness of the proposed method in improving model performance.
[0122] The above description is merely an embodiment of the present application and does not limit the patent scope of the present application. Any equivalent structure or equivalent process transformation made using the contents of the present application specification and drawings, or directly or indirectly applied in other related technical fields, are also included in the patent protection scope of the present application.
Claims
1. A qualitative drug detection method based on multi-concentration mixed SERS data, characterized in that: include: S1. Construct a dataset of drug concentrations based on the SERS spectra of the same drug at high and low concentrations. S2. mixing the low-concentration and high-concentration SERS spectral data in the data set to construct a mixed data set; S3. Based on the improved NSGA-II algorithm, hyperparameter optimization of the random forest model was performed to construct an optimized classification model. S4. Classify drugs based on the optimized classification model of the mixed data set.
2. The method according to claim 1, characterized in that The method of step S1 comprises: S11. Sample preparation: Morphine, methamphetamine, and methadone stock solutions were diluted according to three concentration gradients to obtain nine drug samples. A sodium citrate solution was then added to a boiling chloroauric acid solution and stirred for reaction. The resulting solution was cooled and centrifuged. The precipitate obtained by centrifugation was then redissolved in deionized water to obtain a colloidal SERS substrate containing gold nanoparticles. The drug samples were then added to a 96-well plate, and a fixed ratio of colloidal gold nanoparticles was added, with the drug and gold nanoparticle solution ratio being 1:
9. Finally, the mixed solution was incubated for 5–10 minutes to obtain a mixed solution. S12. Data acquisition: The mixed solution was transferred to a quartz tube using a pipette and placed under a Raman spectrometer. Irradiation was performed using a 785 nm laser with a laser power of 30 mW and an integration time of 10 s. Ten sets of SERS spectra were collected for each sample. S13. Preprocessing the collected SERS spectral data, including baseline correction, denoising, and normalization; S14. Based on the nine drug samples, nine original data sets A, B, C, D, E, F, G, H and I of equal size were constructed; data set A consisted of SERS data of morphine at 0.25 ppm; data set B consisted of SERS data of methadone at 10 ppm; data set C consisted of SERS data of methadone at 12.5 ppm; data set D consisted of SERS data of morphine at 0.5 ppm; data set E consisted of SERS data of methadone at 20 ppm; data set F consisted of SERS data of methadone at 25 ppm; data set G consisted of SERS data of morphine at 1 ppm; data set H consisted of SERS data of methadone at 40 ppm; and data set I consisted of SERS data of methadone at 50 ppm. The low-concentration data set refers to data sets A, B, and C; the high-concentration data set refers to data set DI; each of the data sets A, B, C, D, E, F, G, H, and I contains N SERS spectral data; S15. Datasets A, B, and C are low-concentration drug datasets, and data sets D, E, F, G, H, and I are high-concentration drug datasets. The sample capacity of each of the nine datasets is set to N.
3. The method according to claim 2, characterized in that The method of step S2 comprises: S21. Select 2N data points from each of the two datasets A, B, and C, and then randomly select 0.7N SERS spectral data points from the remaining dataset to form a mixed dataset J; wherein data points are extracted from each of the datasets A, B, and C; S22. Based on the drug types in the remaining data sets, select a similar data set from the data sets D, E, F, G, H, and I that is of the same type as the remaining data sets but has a different concentration, and randomly select 0.3N SERS spectral data from the similar data set and add it to the mixed data set J.
4. The method according to claim 1, wherein The method of step S3 comprises: S31. Based on the random forest model, obtain the F1-score objective function; S32. Set the value range of hyperparameters in the random forest model; S33. Configuring the improved NSGA-II algorithm to initialize the population individuals; S34. Based on the F1-score objective function, combined with the computational efficiency objective function, multi-objective optimization is performed to obtain the multi-objective function and calculate the fitness of the individual; S35. Based on the multi-objective function and the individual fitness, obtain the crowding degree of the individual and the hierarchical ranking of the population individuals; S36. Based on the hierarchical sorting of the population individuals, individuals at the lower levels with high crowding are preferentially selected to obtain the optimal hyperparameters in the optimized classification model.
5. The method according to claim 4, characterized in that The method for obtaining the F1-score objective function in step S31 includes: Based on formula (1), the F1-score model is obtained; wherein, the formula (1) is: Among them, Precision is the proportion of samples predicted to be positive that are actually positive, and Recall is the proportion of samples that are correctly predicted to be positive.
6. The method according to claim 4, characterized in that The method for obtaining the multi-objective function in step S34 includes: Based on formulas (2)-(4), the multi-objective function is obtained; wherein, the formulas (2)-(4) are: P train =1-F1_score (2); T train =n estimate ×max depth (3); multi object =[P train ,T train ] (4); Among them, n estimate and max depth Represents the number of trees and the maximum depth of the tree respectively.
7. The method according to claim 4, characterized in that The method for obtaining the crowding degree of the individual in step S35 includes: Based on formula (5), the crowding degree of the individual is obtained; wherein, the formula (5) is: Among them, f i+1 and f i-1 are the target values of adjacent individuals in the target space; and is the maximum and minimum value of target i; w i (x) is a weight function that is adaptively adjusted based on the position of individual x in the current generation.
Citation Information
Patent Citations
Method for simultaneously detecting two veterinary drugs based on SERS (Surface Enhanced Raman Scattering) marker detection and machine learning
CN117434045A
Amphetamine drug detection method based on combination of LPME and SERS
CN117929351A
Method and apparatus for rapid extraction and analysis, by SERS, of drugs in saliva
US20060084181A1