Cigarette authenticity identification method based on SERS and machine learning and application thereof

By combining SERS technology with machine learning, optimizing SERS substrate materials and sample processing conditions, and utilizing principal component analysis and quadratic kernel function SVM models, the accuracy and sensitivity issues of cigarette authenticity identification were resolved, achieving efficient and rapid brand cigarette identification.

CN122016757APending Publication Date: 2026-05-12GUIZHOU TOBACCO SCI RES INST
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
GUIZHOU TOBACCO SCI RES INST
Filing Date
2026-02-11
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

Existing technologies are insufficient to quickly and accurately identify the authenticity of different brands of cigarettes, especially when the composition is complex and the spectral features overlap. Traditional methods lack accuracy and sensitivity, making it difficult to meet the needs of tobacco market regulation.

Method used

By combining surface-enhanced Raman scattering (SERS) technology with machine learning, optimizing the SERS substrate material and sample processing conditions, and using principal component analysis (PCA) and quadratic kernel function support vector machine (SVM) models, efficient classification of cigarette spectra is achieved.

Benefits of technology

It achieves accurate classification and efficient identification of different brands of cigarettes, with a classification accuracy of up to 99.7%, sensitivity and specificity close to 100%, suitable for rapid on-site detection, and has robustness and generalization ability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122016757A_ABST
    Figure CN122016757A_ABST
Patent Text Reader

Abstract

The invention provides a cigarette authenticity identification method based on SERS (Surface Enhanced Raman Scattering) and machine learning. The cigarette authenticity identification method comprises the following steps: 1) pretreating a cigarette sample; 2) collecting SERS data; 3) data preprocessing and principal component analysis; and 4) machine learning modeling. The method belongs to the technical field of detection, and the spectrum identification capability of a tobacco sample is improved by optimizing an SERS substrate material and sample treatment conditions; on the basis of SERS (Surface Enhanced Raman Scattering) detection, the difference of spectrums of different cigarettes is further evaluated by using PCA (Principal Component Analysis), and a quadratic kernel function SVM (Support Vector Machine) is finally selected by screening various machine learning models to establish a true and false cigarette identification model, so that accurate classification and efficient identification of different brands of cigarettes are realized; powerful technical support is provided for rapid and high-reliability identification of cigarette authenticity.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of detection technology, and in particular relates to a method for identifying genuine and counterfeit cigarettes based on SERS and machine learning, and its application. Background Technology

[0002] Cigarettes have a complex composition, containing alkaloids, phenols, flavonoids, aromatic compounds, and various flavorings and additives. Nicotine, chlorogenic acid, scopolamine, hyoscyamine (a coumarin derivative), and rutin are among the main components. Different brands of cigarettes have varying market values ​​and tastes due to differences in tobacco grade, additive types, and proportions. Based on these differences in market value, many unscrupulous individuals alter the type, grade, and composition of tobacco to create counterfeit cigarettes or use low-quality cigarettes to impersonate high-quality ones, thereby achieving illegal profits. Therefore, accurately, quickly, and cost-effectively identifying different brands of cigarettes and their authenticity has become one of the key issues that urgently needs to be addressed in tobacco market regulation.

[0003] Currently, methods for identifying genuine and counterfeit cigarettes mainly include sensory observation, smoking tasting evaluation, and instrumental testing. Sensory observation and smoking tasting evaluation rely heavily on subjective experience and are significantly influenced by individual factors; their accuracy decreases as counterfeiting techniques improve. With the development of science and technology, instrumental testing is becoming increasingly widely used. Common detection methods include infrared spectroscopy, gas chromatography-mass spectrometry (GC-MS), and stable isotope labeling. While GC-MS and stable isotope labeling offer high accuracy, their high instrument costs, complex analysis processes, and long processing times make them unsuitable for rapid on-site detection. Infrared spectroscopy, despite its advantages of rapid analysis and simple operation, has limited identification capability and sensitivity for cigarette samples with complex compositions and significant overlap in spectral features. Therefore, there is an urgent need to develop a technology that combines high sensitivity, ease of operation, and suitability for rapid on-site detection of cigarette samples.

[0004] Surface-enhanced Raman scattering (SERS) is a rapid, simple, and sensitive analytical method that has been widely applied in food, biology, medicine, and environmental fields. SERS significantly enhances the Raman signal of the sample by adsorbing it onto the surface of nanomaterials such as Au and Ag, thereby improving detection sensitivity. Generally, pure silver nanomaterials lack ideal chemical stability, while pure gold nanomaterials do not provide ideal SERS enhancement. Therefore, the selection of the active substrate for SERS is crucial to obtaining a more ideal Raman enhanced signal. Existing research has shown that SERS technology can sensitively detect important components in tobacco such as nicotine, chlorogenic acid, coumarins, and flavonoids, confirming its application potential in the tobacco industry. Chinese patent application CN 112595702 A discloses a method for rapid detection of hexaconazole in tobacco using surface-enhanced Raman scattering (SERS). This method achieves highly sensitive detection of hexaconazole in tobacco through steps such as preparing mercapto-β-cyclodextrin-modified silver nanomaterials as Raman scattering enhancers and determining the characteristic peaks of hexaconazole SERS. Chinese patent application CN 113125409 A discloses a method for rapid detection of sec-butylamine in tobacco using SERS. This method achieves highly sensitive detection of sec-butylamine in tobacco through steps such as preparing a transparent and flexible gold nanoparticle tape SERS substrate and determining the characteristic peaks of sec-butylamine SERS.

[0005] However, tobacco samples are complex in composition and exhibit significant batch-to-batch variations, making it difficult to distinguish genuine from counterfeit tobacco brands using traditional SERS spectral analysis alone. This is especially true when the spectral characteristics of multiple components overlap considerably, rendering manual identification and simple statistical methods ineffective. In recent years, machine learning technology, with its powerful data processing and pattern recognition capabilities, has made groundbreaking progress in the field of complex spectral analysis. Strategies combining SERS technology with machine learning have been successfully applied to the identification of complex systems such as tea grading and pesticide residue detection, but their application in cigarette brand and genuine / counterfeit identification has not yet been reported.

[0006] Therefore, providing a method for identifying genuine and counterfeit cigarettes based on SERS and machine learning, and its application, is of great significance for regulating the tobacco market order and protecting consumer rights. Summary of the Invention

[0007] To address the problems existing in the prior art, this invention provides a method for identifying genuine and counterfeit cigarettes based on SERS and machine learning. By optimizing the SERS substrate material and sample processing conditions, the spectral identification capability of tobacco samples is improved. Building upon SERS detection, principal component analysis (PCA) is further used to evaluate the differences in the spectra of different cigarettes. Through screening various machine learning models, a quadratic kernel function (SVM) is ultimately selected to establish a model for identifying genuine and counterfeit cigarettes. This achieves accurate classification and efficient identification of different brands of cigarettes, providing strong technical support for rapid and highly reliable identification of genuine and counterfeit cigarettes. Based on the above findings, this invention is thus completed.

[0008] The objectives of this invention will be further explained by the following detailed description.

[0009] This invention provides a method for identifying genuine and counterfeit cigarettes based on SERS and machine learning, comprising the following steps: 1) Pretreatment of cigarette samples: Take the tobacco shreds of the cigarette sample, grind them, sieve them to obtain cigarette powder; add the cigarette powder to a centrifuge tube, add ultrapure water, sonicate, then extract in a water bath at 38-42℃, centrifuge, collect the supernatant to obtain the sample extract; 2) SERS data acquisition: Take the SERS substrate material, centrifuge, remove the supernatant to obtain a concentrated SERS substrate particle solution, mix it with the sample extract, sonicate, mix well, and then drop it onto Teflon tape for SERS detection to acquire SERS data; the SERS substrate material is prepared by reducing silver nitrate with sodium citrate; 3) Data preprocessing and principal component analysis: The SERS data are preprocessed and subjected to principal component analysis to obtain dimensionality-reduced data; the preprocessing includes: baseline correction to eliminate fluorescence background interference, Savitzky-Golay spectral smoothing to reduce noise, and vector normalization to eliminate the influence of signal intensity fluctuations; the principal component analysis includes dimensionality reduction. 4) Machine learning modeling: The data after dimensionality reduction is randomly divided into training and testing sets. The quadratic kernel function SVM is used as the model for training and performance evaluation to obtain the results of cigarette authenticity identification.

[0010] Preferably, the SERS substrate material is AgC NPs.

[0011] Preferably, the volume ratio of the concentrated SERS substrate particle solution to the sample extract is (1.8-2.2):1. More preferably, the volume ratio of the concentrated SERS substrate particle solution to the sample extract is 2:1.

[0012] Preferably, in step 1), the ultrasonic treatment conditions include: 35-45 kHz, 250-350 W, 4-6 min; in step 2), the ultrasonic treatment conditions include: 35-45 kHz, 250-350 W, 0.5-2 min. More preferably, in step 1), the ultrasonic treatment conditions include: 40 kHz, 300 W, 5 min; in step 2), the ultrasonic treatment conditions include: 40 kHz, 300 W, 1 min.

[0013] Preferably, the extraction time is 30-90 min. More preferably, the extraction time is 60 min.

[0014] Preferably, the SERS detection conditions include: using a Raman microscope, under a 50x objective lens, and excitation with a 532 nm laser. More preferably, the SERS detection conditions further include: laser power of 10 mW, integration time of 1 s, and acquisition of 200 spectral data points for each sample.

[0015] Preferably, the performance evaluation indicators include: true positive rate, true negative rate, false positive rate, false negative rate, overall accuracy, sensitivity, and specificity.

[0016] Preferably, the ratio of the training set to the test set is (65%-75%):(25%-35%). More preferably, the ratio of the training set to the test set is 70%:30%.

[0017] Preferably, the formula for the quadratic kernel function SVM includes: f(x)=∑i=1NαiyiK(xi,x)+bf(x)=\sum_{i=1}^{N} \alpha_i y_i K(x_i, x) + bf(x)=i=1∑NαiyiK(xi,x)+b ;where,xxx: spectral features to be judged (a certain sample),xix_ixi: support vector,yiy_iyi: category label (real cigarette or fake cigarette). Kernel functions (such as RBF).

[0018] SVM computes a discriminant function and outputs the following symbols: f(x)>0: classifies it as a certain type (real cigarette); f(x)<0: classifies it as another type (fake cigarette).

[0019] Furthermore, this invention also provides the application of the SERS-based and machine learning-based cigarette authenticity identification method in cigarette authenticity identification and / or cigarette brand recognition.

[0020] Compared with the prior art, the beneficial effects of the present invention include: (1) This invention provides a method for identifying genuine and counterfeit cigarettes based on SERS and machine learning. By optimizing the SERS substrate type and sample pretreatment conditions, AgC NPs were determined as the optimal substrate material, and 40℃ water bath extraction was determined as the optimal pretreatment temperature, which can significantly improve the stability of spectral signals and the distinguishability between genuine and counterfeit cigarette samples. Based on SERS spectral data acquisition and processing, a systematic comparative analysis of various machine learning models was conducted. The results show that the quadratic kernel function SVM model exhibits excellent classification performance on both the training and test sets, with classification accuracies of 99.6% and 99.7%, respectively, and sensitivity and specificity approaching 100.0%. The quadratic kernel function SVM model achieved an overall prediction accuracy of 92.3% for blind test samples from the same batch, and the accuracy for determining the authenticity and brand attribution of unknown cigarette samples from different batches both exceeded 80.0%, verifying the robustness and generalization ability of the method provided in this invention.

[0021] (2) The method of SERS technology coupled with quadratic kernel function SVM model established in this invention can accurately distinguish between genuine and counterfeit cigarettes of 17 brands without the need for complex preprocessing and expensive instruments. The overall accuracy of the test set is as high as 99.7%, realizing rapid, non-destructive and high-precision identification of genuine and counterfeit cigarettes of multiple brands, and providing technical support for tobacco market supervision, public safety law enforcement and quality control. Attached Figure Description

[0022] Figure 1 UV-Vis characteristic peak spectra of three nanoparticles, including AgC NPs.

[0023] Figure 2 SEM and EDS characterization spectra of three nanoparticles, including AgC NPs; among them... Figure 2 -A1 and Figure 2 -A2 are the detection spectra of AgC NPs. Figure 2 -B1 and Figure 2 -B2 are the detection spectra of Ag7Au3NPs. Figure 2 -C1 and Figure 2 -C2 are the detection spectra of AgA NPs.

[0024] Figure 3 Representative spectra of different cigarette samples obtained using AgC NPs as the SERS substrate.

[0025] Figure 4 SERS spectra and PCA results of cigarette samples at different temperatures.

[0026] Figure 5 SERS spectra of cigarette sample YYP1 at different temperatures, 1036 cm⁻¹-1 RSD detection results at wavenumber.

[0027] Figure 6 Background spectra of three nanoparticles and SERS spectra of different nanoparticles mixed with YYR1.

[0028] Figure 7 SERS spectra and PCA results of cigarette samples on AgA NPs and Ag7Au3NPs substrates.

[0029] Figure 8 The wavenumber in the SERS spectra of the three nanoparticles was 1036 cm⁻¹. -1 RSD plot at the location.

[0030] Figure 9 SERS spectra of concentrated nanoparticle solutions and sample extracts mixed at different ratios.

[0031] Figure 10 PCA score map of cigarette samples based on AgC NPs.

[0032] Figure 11 A comparison of the accuracy of eight candidate models in identifying cigarette types.

[0033] Figure 12 The secondary kernel SVM model constructed in this invention predicts test results for the same batch of blind test samples that were not trained. Detailed Implementation

[0034] The present invention will now be described in further detail with reference to the accompanying drawings and embodiments.

[0035] In this invention, the reagents, materials, and equipment involved are conventional commercially available products or can be obtained through conventional technical means in the field. For example: silver nitrate (AgNO3, 99.0%, Sigma-Aldrich), sodium citrate (C6H5Na3O7, 99.0%, Sigma-Aldrich), gold chloride trihydrate (Sigma-Aldrich), and ascorbic acid (>99.0%, Sigma-Aldrich).

[0036] The ultraviolet-visible spectra were obtained using a MAYA2000Pro ultraviolet-visible spectrometer from Marine Optics, Inc. Example 1: Cigarette Sample

[0037] A total of 34 cigarette samples were collected (as shown in Table 1), which came from the Guizhou Provincial Tobacco Quality Supervision and Testing Station and the Guizhou Provincial Tobacco Science Research Institute. The samples included one set each of genuine and counterfeit cigarettes from 17 brands, and were stored in sealed bags at room temperature.

[0038] Table 1. Authenticity information and abbreviation codes of 34 cigarette samples.

[0039] In addition, 10 different batches of cigarette samples (blind samples) were collected for subsequent verification of the generalization performance of the cigarette authenticity identification model. All cigarette samples were stored in sealed bags. Example 2: SERS spectral detection analysis and condition optimization

[0040] Raman spectroscopy acquisition was performed using a specially designed Raman microscope equipped with a Princeton Instruments Pixis-100BR CCD camera, an Acton SP-2500i spectrometer, and a 50 objective lens (NA = 0.60), acquiring SERS spectra under 532 nm laser excitation. 2.1 Preparation and Investigation of SERS-based Nanomaterials

[0041] To systematically compare the effects of different nanosubstrates on SERS detection performance, this study synthesized three types of nanoparticles: AgC NPs, Ag7Au3 NPs, and AgA NPs. The synthesis methods were based on relevant literature or optimized from existing literature. The preparation of AgC NPs was based on the method published by Wang, X. et al. Chemosphere The paper, "Multifunctional 3D magnetic carbon aerogel for adsorption separation and highly sensitive SERS detection of malachite green," specifically describes the preparation of malachite green by reducing silver nitrate with sodium citrate as a reducing agent. The reaction was carried out under reflux condensation at 100°C with continuous stirring for 1 hour to obtain a stable colloidal solution. Ag7Au3NPs were synthesized using a modified Turkevich method, specifically involving the sequential addition of chloroauric acid solution and silver nitrate solution to boiling deionized water, followed by the rapid addition of sodium citrate, and reaction under reflux for 30 minutes to form alloy nanoparticles. The preparation of AgA NPs is based on the work published by Liu, HQ et al. Analytical ChemistryThe paper, "In SituHot-Spot Assembly as a General Strategy for Probing Single Biomolecules," specifically describes the following steps: Sodium citrate, silver nitrate, and sodium chloride are premixed and rapidly injected into a preheated, boiling ascorbic acid solution. The mixture is reacted with stirring for 1 hour to obtain the product. The synthesized nanoparticle colloidal solution is then naturally cooled to room temperature after the reaction, and subsequently transferred to a light-protected environment and stored at 4°C until subsequent experimental use. The UV-Vis characteristic peak spectra of the three nanoparticles, including AgC NPs, are shown below. Figure 1 As shown in Figure 2.1, the SEM and EDS characterization spectra of the three nanoparticles, including AgC NPs, prepared in section 2.1 are as follows. Figure 2 As shown. 2.2 Sample Pretreatment

[0042] For the 34 cigarette samples in Example 1, tobacco shreds were taken and ground into a uniform powder using a grinder. The powder was then passed through a 20-mesh sieve to obtain cigarette powder. 0.10 g of the cigarette powder was weighed and added to a 10 mL centrifuge tube, along with 5 mL of ultrapure water. The mixture was ultrasonicated for 5 min (40 kHz, 300 W). Subsequently, the cigarette powder was extracted for 1 h in water baths at room temperature, 40℃, and 80℃, respectively. After treatment, the mixture was centrifuged at 10000 rpm for 10 min, and the supernatant was collected to obtain the sample extract, which was used for subsequent SERS testing. 2.3 SERS Data Acquisition and Processing

[0043] Take 1 mL of the nanoparticle colloidal solution prepared in step 2.1, centrifuge at 10000 rpm for 10 min, remove 980 μL of supernatant to obtain a concentrated nanoparticle solution. Mix 20 μL of the concentrated nanoparticle solution with 10 μL of the sample extract prepared in step 2.2, sonicate for 1 min, mix well, and then drop onto Teflon tape for SERS detection. Under a 50x objective lens, excitation was performed using a 532 nm laser with a laser power of 10 mW and an integration time of 1 s. 200 spectral data points were collected for each sample, resulting in a total of 6800 spectral data points. Simultaneously, 50 spectral data points were collected from the same batch of samples for external model validation (blind prediction experiment).

[0044] In addition, 50 spectral data points were collected from 10 different batches of cigarette samples (blind samples) to identify whether they were genuine or counterfeit.

[0045] All raw spectral data were subjected to a standardized preprocessing procedure before statistical analysis, including baseline correction to eliminate fluorescence background interference, Savitzky-Golay spectral smoothing to reduce noise, and vector normalization to eliminate the influence of signal intensity fluctuations. Subsequently, principal component analysis (PCA) was used to reduce the dimensionality of the preprocessed spectral data in order to extract key feature information and preliminarily assess the clustering distribution of different cigarette brands.

[0046] In addition, the 10 cigarette samples (blind samples) from different batches in Example 1 were preprocessed according to 2.2, and 50 spectral data points were collected and processed for each cigarette sample according to 2.3. 2.4 SERS spectral characteristic analysis of cigarette samples and optimization of SERS spectral detection conditions

[0047] To elucidate the main chemical characteristics of cigarette samples and explore the spectral differences between genuine and counterfeit cigarettes, this study performed SERS detection on 34 cigarette samples from Example 1 using the AgC NPs substrate prepared in Example 2. Fifty spectra were collected for each sample. After baseline correction, Savitzky-Golay smoothing, and normalization, representative spectra were obtained as follows: Figure 3 As shown in the figure. The results indicate that multiple characteristic Raman vibration peaks, at 1036, 1056, and 1195 cm⁻¹, can be detected in the cigarette samples. −1 The peaks at these locations are attributed to nicotine; 1117, 1165, 1329 cm −1 The peaks at these locations are attributed to chlorogenic acid; 610, 960, 1084, 1117, 1200, 1347, 1420, 1560 cm⁻¹. −1 The peaks at that location belong to Rutin; 736, 932, 1428, 1602 cm −1 The peak at 1272 cm is scopolamine; −1 The peaks are coumarin, and other peaks may be from flavorings or other additives. These peak positions are closely related to the vibrational modes of the main active ingredients in tobacco leaves and additives, confirming the high sensitivity and reliability of SERS technology in cigarette component detection.

[0048] To investigate the effect of sample pretreatment conditions on the discriminative power of SERS signals, this study pretreated samples by water bath extraction for 1 h at three different temperatures: room temperature, 40℃, and 80℃. The SERS spectra and PCA results of cigarette samples at different temperatures are shown below. Figure 4 As shown. Among them, Figure 4 -A1、 Figure 4 -B1、 Figure 4-C1 are the SERS spectra of five representative cigarette samples YYP1, YX1, HHL1, LQ1 and NJ1 at different temperatures, respectively. Figure 4 -A2、 Figure 4 -B2、 Figure 4 -C2 represents the corresponding PCA score distribution. The results show that moderate heating significantly improves spectral discrimination. At room temperature, most samples exhibit weak spectral signal intensity and significant overlap in their distributions, leading to PCA clustering's inability to effectively distinguish genuine from counterfeit samples. However, after treatment at 40℃, the overall signal-to-noise ratio of the spectrum slightly improves, and several key component peaks (such as 1036, 1165, and 1420 cm⁻¹) are distinguished. -1 The intensity and relative ratio of the ) were more stable, and the clustering distribution of different brands of genuine cigarette samples showed obvious boundaries; while the spectrum under 80℃ treatment conditions was not much different from that under 40℃, but from Figure 4 -C2 shows that the overlap between samples is more pronounced than at 40℃. The SERS spectra of cigarette sample YYP1 at three different temperatures (room temperature, 40℃, and 80℃) show a peak at 1036 cm⁻¹. -1 RSD detection results at wavenumber are as follows Figure 5 As shown, the RSD is less than 20% at all three temperatures, especially at 40 ℃ where the RSD is the lowest (less than 10%). This fully demonstrates that the SERS substrate and extraction in a 40℃ water bath used in this study are of great significance for improving the stability and accuracy of cigarette sample test results.

[0049] To investigate the effects of different nanomaterial substrates on the enhancement and discrimination performance of cigarette SERS signals, this study selected three substrates (AgC NPs, Ag7Au3 NPs, and AgA NPs) and processed the same batch of samples (five representative cigarette samples YYP1, YX1, HHL1, LQ1, and NJ1) at 40℃ according to steps 2.2 and 2.3. The background spectra of the three different nanoparticles and the SERS spectra of the different nanoparticles mixed with YYR1 are shown below. Figure 6 As shown. From Figure 6 It can be seen that the enhancement effect of different substrate nanoparticles on the Raman peaks of the samples is inconsistent. The results of AgC NPs are as follows: Figure 4 -B1、 Figure 4 -B2 shows the SERS spectra and PCA results of cigarette samples with AgA NPs and Ag7Au3NPs substrates. Figure 7 As shown, the results for AgA NPs are as follows: Figure 7 -A1、 Figure 7 -A2 is shown; the results for Ag7Au3NPs are as follows: Figure 7 -B1、 Figure 7 -B2 shows the SERS spectra of the three nanoparticles with a wavenumber of 1036 cm⁻¹. -1RSD plot at the location as shown Figure 8 As shown, the three substrates differ in particle size distribution, composition, and enhancement mechanism, resulting in varying spectral enhancement capabilities and signal stability. The PCA plot clearly shows that AgC NPs offer the best differentiation for cigarettes, while... Figure 7 -A2 shows that AgA NPs cannot distinguish between the five types of cigarettes, from... Figure 7 -B2 shows that the data of LQ1 and NJ1 overlap significantly, and the data of YYP1 and HHL1 overlap significantly. Furthermore, from... Figure 8 As can be seen, AgC NPs have the lowest RSD (8.31%) when used as the substrate. Therefore, considering both spectral enhancement and discrimination ability, AgC NPs are selected as the optimal SERS substrate for this invention. Furthermore, this invention also obtained corresponding SERS spectra by mixing the concentrated nanoparticle solution from step 2.3 with the sample extract obtained in step 2.2 at different ratios. Figure 9 As shown, the spectral information is richest when the volume ratio of the concentrated nanoparticle solution to the sample extract prepared in step 2.2 is 2:1. Therefore, this ratio was chosen for subsequent experiments.

[0050] Using AgC NPs as a substrate, extraction in a 40℃ water bath, and a 2:1 volume ratio of concentrated nanoparticle solution to sample extract, SERS spectral data from multiple groups of genuine and counterfeit cigarette samples underwent the aforementioned standardized preprocessing before further PCA analysis. The results are as follows: Figure 10 As shown. When analyzing 6 groups of genuine and counterfeit cigarettes (a total of 12 samples), genuine cigarettes and their corresponding counterfeit cigarettes could be distinguished relatively well, but there was still overlap between different brands. For example, NJ1 and YX2 overlapped. Figure 10 -A is shown; however, when analyzing all 17 types of cigarettes, most samples clustered together, indicating insufficient differentiation, such as... Figure 10 As shown in -B, this illustrates that while PCA, as an unsupervised method, can achieve preliminary visualization and dimensionality reduction, its classification ability is limited and cannot meet the requirements for high-precision identification of different batches of genuine and counterfeit cigarettes. Therefore, this invention, through experiments, found that PCA has limited ability to distinguish cigarette sample brands, necessitating the introduction of supervised machine learning methods to achieve more accurate classification and identification. Example 3: Construction of Machine Learning Model

[0051] To improve the generalization ability of the model, the labels of genuine and counterfeit cigarette samples were randomly shuffled before modeling. Specifically, the collected SERS data was preprocessed and subjected to PCA dimensionality reduction as described in section 2.3. The dataset was then randomly divided into training and test sets in a 7:3 ratio, and five-fold cross-validation was used to evaluate the model's performance. This experiment systematically evaluated the ability of various machine learning models to identify cigarette types. Through comparative analysis, eight machine learning models with high accuracy were selected as candidate models: SVM (Support Vector Machine) with different kernel functions (linear, quadratic, cubic kernel functions, and medium and coarse Gaussian kernel functions), Tree, k-Nearest Neighbors (KNN), and Naive Bayes (NB). Training and performance evaluation were then conducted. The model performance evaluation uses multiple indicators such as true positive rate (TPR), true negative rate (TNR), false positive rate (FPR), false negative rate (FNR), overall accuracy, sensitivity, and specificity. The classification results are visualized using a confusion matrix to comprehensively compare the recognition effects of different models and select the best-performing machine learning model for cigarette authenticity identification.

[0052] from Figure 11 As can be seen, all SVM classifiers exhibit high accuracy on both the training and test sets, outperforming Tree, KNN, and NB, demonstrating the advantages of Support Vector Machines in cigarette type recognition. It can be observed that the multinomial kernel function performs better than the Gaussian kernel function, while the quadratic and cubic kernel SVMs achieve overall accuracies of 99.6% and 99.7% on the training and test data, respectively. Considering the higher flexibility of the cubic kernel function, which may increase the risk of overfitting, the quadratic kernel SVM was ultimately chosen as the optimal model. Example 4: Verification of Cigarette Authenticity Identification Method and Classification Results Based on SERS and Machine Learning

[0053] The method for identifying genuine and counterfeit cigarettes based on SERS and machine learning includes the following steps: 1) Pretreatment of cigarette samples: Take the tobacco shreds from the cigarette sample, grind them into a uniform powder using a grinder, pass them through a 20-mesh sieve to obtain cigarette powder; add 0.10 g of cigarette powder to a 10 mL centrifuge tube, add 5 mL of ultrapure water, sonicate for 5 min (40 kHz, 300 W), then extract in a 40℃ water bath for 1 h, centrifuge at 10000 rpm for 10 min, collect the supernatant to obtain the sample extract; 2) SERS data acquisition: 1 mL of the AgC NPs nanoparticle colloidal solution prepared in section 2.1 was used as the SERS substrate material. It was centrifuged at 10000 rpm for 10 min, and 980 μL of supernatant was removed to obtain 20 μL of concentrated SERS substrate particle solution. This solution was mixed with 10 μL of sample extract and sonicated for 1 min (40 kHz, 300 W). After mixing, the mixture was dropped onto Teflon tape for SERS detection, and SERS data were acquired. The SERS detection conditions included: Raman microscopy, under a 50x objective lens, and excitation with a 532 nm laser; laser power of 10 mW, integration time of 1 s, and acquisition of 200 spectral data points for each sample.

[0054] 3) Data preprocessing and principal component analysis: The SERS data are preprocessed and subjected to principal component analysis to obtain dimensionality-reduced data; the preprocessing includes: baseline correction to eliminate fluorescence background interference, Savitzky-Golay spectral smoothing to reduce noise, and vector normalization to eliminate the influence of signal intensity fluctuations; the principal component analysis includes dimensionality reduction. 4) Machine learning modeling: The data after dimensionality reduction is randomly divided into a training set (70%) and a test set (30%). The quadratic kernel function SVM is used as the model for training and performance evaluation to obtain the results of cigarette authenticity identification.

[0055] Based on the SERS spectral data acquisition and processing results, the classification evaluation (accuracy, sensitivity, and specificity) of the quadratic kernel SVM model on genuine and counterfeit cigarette samples of different brands is shown in Table 2. Table 2 shows that the classification accuracy of the quadratic kernel SVM on the training set and the test set is 99.6% and 99.7%, respectively, verifying the reliability of the model. The sensitivity and specificity of the model on genuine and counterfeit cigarette samples of various brands are close to 100.0%, further demonstrating its strong discrimination ability and stability. Therefore, the quadratic kernel SVM was used as the classifier for genuine and counterfeit cigarettes in subsequent experiments.

[0056] Table 2. Classification evaluation results of the quadratic kernel SVM model on genuine and counterfeit cigarette samples of different brands.

[0057]

[0058] To evaluate the generalization ability and robustness of the constructed quadratic kernel SVM model on untrained samples, this study first conducted prediction tests on the same batch of blind test samples that were not used in training. Fifty spectra were independently collected for each blind test sample and input into the trained model for discrimination. The results are as follows: Figure 12As shown, the overall recognition accuracy was 92.3%, indicating that the model maintains high discrimination accuracy even when dealing with samples from the same batch, demonstrating excellent differentiation between genuine and counterfeit cigarettes. The prediction accuracy for most genuine and counterfeit cigarette samples approached 100.0%, demonstrating that this invention, utilizing AgC NPs as the SERS substrate and extraction conditions at 40℃ water bath, combined with SERS spectral feature acquisition and processing and a secondary kernel SVM model, can achieve the identification of genuine and counterfeit cigarettes.

[0059] To further verify the adaptability of the constructed quadratic kernel SVM model across batches of samples, this study conducted prediction tests on 10 cigarette samples (blind samples) from different batches in Example 1. The results are shown in Table 3. The model can accurately identify the authenticity and most likely brand category in the vast majority of cases; 80% of the samples achieved a prediction accuracy of 100.0%, and the accuracy of the remaining samples remained between 81.0% and 95.0%, indicating that the model possesses a certain degree of adaptability across batches and brand recognition capability. Considering the potential differences in formulation, tobacco grade, and additive type among different batches of cigarettes, the model still achieves a high recognition rate in real, complex samples, verifying the robustness and potential for wider application of this method.

[0060] As can be seen, by combining SERS-based rapid detection with an optimized quadratic kernel SVM algorithm, this invention can achieve rapid identification of genuine and counterfeit cigarettes and prediction of brand affiliation, and the model exhibits good generalization performance on new samples. In the future, by expanding the training sample size, introducing cigarette samples from more sources and batches, and combining feature selection or deep learning methods, it is expected that the model's cross-batch robustness and brand recognition accuracy can be further improved.

[0061] Table 3. Prediction results of unknown cigarette blind samples based on SERS spectral feature acquisition and processing combined with a secondary kernel SVM model.

[0062] In summary, this invention provides a rapid and high-precision method for identifying genuine and counterfeit cigarettes by combining surface-enhanced Raman spectroscopy (SERS) with machine learning algorithms, and systematically evaluates the impact of different experimental conditions on detection performance. By optimizing the SERS substrate type and sample pretreatment conditions, AgC NPs were determined to be the optimal substrate material, and 40℃ extraction was identified as the optimal pretreatment temperature, which significantly improves the stability of the spectral signal and the distinguishability between genuine and counterfeit cigarette samples. Based on SERS spectral data acquisition and processing, a systematic comparative analysis of various machine learning models was conducted. The results show that the quadratic kernel function SVM model exhibits excellent classification performance on both the training and test sets, with classification accuracies of 99.6% and 99.7%, respectively, and sensitivity and specificity approaching 100.0%. The quadratic kernel function SVM model achieved an overall prediction accuracy of 92.3% for blind test samples from the same batch, and the accuracy for determining the authenticity and brand attribution of unknown cigarette samples from different batches both exceeded 80.0%, verifying the robustness and generalization ability of the method. The method of SERS technology coupled with a quadratic kernel function SVM model established in this invention can achieve rapid, non-destructive, and high-precision identification of the authenticity of multiple brands of cigarettes without the need for complex preprocessing and expensive instruments, providing technical support for tobacco market supervision, public safety law enforcement, and quality control.

[0063] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the concept of the present invention, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. A method for identifying genuine and counterfeit cigarettes based on SERS and machine learning, characterized by: Includes the following steps: 1) Pretreatment of cigarette samples: Take the tobacco shreds of the cigarette sample, grind them, sieve them to obtain cigarette powder; add the cigarette powder to a centrifuge tube, add ultrapure water, sonicate, then extract in a water bath at 38-42℃, centrifuge, collect the supernatant to obtain the sample extract; 2) SERS data acquisition: Take the SERS substrate material, centrifuge, remove the supernatant to obtain a concentrated SERS substrate particle solution, mix it with the sample extract, sonicate, mix well, and then drop it onto Teflon tape for SERS detection to acquire SERS data. The SERS substrate material was prepared by reducing silver nitrate with sodium citrate. 3) Data preprocessing and principal component analysis: The SERS data are preprocessed and subjected to principal component analysis to obtain dimensionality-reduced data; the preprocessing includes: baseline correction to eliminate fluorescence background interference, Savitzky-Golay spectral smoothing to reduce noise, and vector normalization to eliminate the influence of signal intensity fluctuations; the principal component analysis includes dimensionality reduction. 4) Machine learning modeling: The data after dimensionality reduction is randomly divided into training and testing sets. The quadratic kernel function SVM is used as the model for training and performance evaluation to obtain the results of cigarette authenticity identification.

2. The cigarette authenticity identification method based on SERS and machine learning according to claim 1, characterized in that: The SERS substrate material is AgC NPs.

3. The cigarette authenticity identification method based on SERS and machine learning according to claim 1, characterized in that: The volume ratio of the concentrated SERS substrate particle solution to the sample extract is (1.8-2.2):

1.

4. The cigarette authenticity identification method based on SERS and machine learning according to any one of claims 1 to 3, characterized in that: In step 1), the conditions for ultrasonic treatment include: 35-45 kHz, 250-350 W, and 4-6 min; in step 2), the conditions for ultrasonic treatment include: 35-45 kHz, 250-350 W, and 0.5-2 min.

5. The cigarette authenticity identification method based on SERS and machine learning according to any one of claims 1 to 3, characterized in that: The extraction time is 30-90 minutes.

6. The cigarette authenticity identification method based on SERS and machine learning according to any one of claims 1 to 3, characterized in that: The conditions for SERS detection include: using a Raman microscope, under a 50x objective lens, and using a 532 nm laser for excitation.

7. The cigarette authenticity identification method based on SERS and machine learning according to any one of claims 1 to 3, characterized in that: The performance evaluation metrics include: true positive rate, true negative rate, false positive rate, false negative rate, overall accuracy, sensitivity, and specificity.

8. The cigarette authenticity identification method based on SERS and machine learning according to any one of claims 1 to 3, characterized in that: The ratio of the training set to the test set is (65%-75%): (25%-35%).

9. The cigarette authenticity identification method based on SERS and machine learning according to any one of claims 1 to 3, characterized in that: The formula for the quadratic kernel function SVM includes: ;in, Spectral features to be identified Support vectors, Category tags, Kernel function.

10. The application of the SERS and machine learning-based cigarette authenticity identification method according to any one of claims 1 to 9 in cigarette authenticity identification and / or cigarette brand identification.