A method for rapid and non-destructive detection of total sugar content in tobacco based on hyperspectral technology

CN122835972APending Publication Date: 2026-09-29LIUYANG BRANCH OF CHANGSHA COMPANY OF HUNAN TOBACCO
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611072931.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-20
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

[0004]铁氰化钾比色法烟草品质核心化学成分检测的传统基准方法,依托成熟的标准化体系与明确的技术原理,构建了覆盖低成本筛查、挥发性成分分离、极性成分精准定量的多元化检测体系,至今仍是实验室保障检测准确性与结果可比性的核心手段,这些传统方法存在显著局限,需复杂前处理、检测周期长且依赖精细操作控制,难以满足现代生产实时质量监控需求,亟需快速无损检测技术突破

Benefits of technology

[0005]为了解决上述的技术问题,本发明的目的是提供一种快速无损检测烟草总糖成分的高光谱方法,采用高光谱技术对样品直接进行检测;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122835972A_ABST
    Figure CN122835972A_ABST
Patent Text Reader

Abstract

A kind of tobacco total sugar component rapid nondestructive detection method based on hyperspectral technology, adopts tobacco total sugar component hyperspectral model, with the hyperspectral image of tobacco as input, the total sugar content of tobacco is obtained by prediction.The construction method of tobacco total sugar component hyperspectral model is:1) tobacco sample is dried, crushed and then stored;2) use hyperspectral device to collect the original spectral image of powder tobacco sample;3) divide ROI region from the original spectral image collected in step 2), and extract the original spectral data;4) the original spectral data extracted in step 3) is respectively preprocessed using different preprocessing methods;5) the spectral data processed in step 4) is respectively subjected to feature screening using different feature screening methods, and useless characteristic wavelength is filtered out;6) using the spectral data processed in step 5), a regression model is established using PLSR method.The optimal combination of preprocessing method and feature screening method constitutes the tobacco total sugar component hyperspectral model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of non-destructive testing technology for agricultural product components, and in particular to a rapid non-destructive testing method for total sugar content in tobacco based on hyperspectral technology. Background Technology

[0002] In tobacco quality grading, carbohydrates (such as total sugar) are a key parameter, and their accuracy directly affects the effective conversion of raw material value and the control of product risks.

[0003] Traditional methods for detecting total sugars in tobacco, such as the potassium ferricyanide colorimetric method, are classic chemical analysis techniques. Their core principle is based on redox reactions: in an alkaline environment, potassium ferricyanide reacts with reducing sugars such as glucose and fructose in tobacco, reducing the yellow potassium ferricyanide to colorless potassium ferrocyanide, while the reducing sugars themselves are oxidized. By detecting the remaining potassium ferricyanide in the reaction system (e.g., further reacting with ferric iron to produce Prussian blue and then colorimetrically measuring it) or monitoring color changes during the reaction process, the reducing sugar content in the sample can be calculated indirectly and accurately. This method plays a crucial role in the analysis of tobacco chemical components, especially in sugar detection.

[0004] The potassium ferricyanide colorimetric method is a traditional benchmark method for detecting core chemical components in tobacco quality. Relying on a mature standardization system and well-defined technical principles, it has built a diversified detection system covering low-cost screening, separation of volatile components, and precise quantification of polar components. To this day, it remains a core means for laboratories to ensure the accuracy of detection and the comparability of results. However, these traditional methods have significant limitations. They require complex pretreatment, have long detection cycles, and rely on precise operational control, making it difficult to meet the real-time quality monitoring needs of modern production. Breakthroughs in rapid and non-destructive testing technologies are urgently needed. Summary of the Invention

[0005] To address the aforementioned technical problems, the purpose of this invention is to provide a rapid and non-destructive hyperspectral method for detecting total sugar content in tobacco, which uses hyperspectral technology to directly detect the sample.

[0006] The spectral information of the samples was collected, and preprocessing methods such as multivariate scattering correction, standard normal variable transformation, and first derivative were used to eliminate spectral noise and baseline drift interference. The improvement effects of different preprocessing methods on data quality were compared. On this basis, the partial least squares regression (PLSR) method was used to construct a hyperspectral full-band prediction model. Furthermore, feature selection methods were used to extract feature variables of each quality parameter to improve model performance. Finally, a complete hyperspectral model of total sugar components in tobacco was established to achieve rapid and non-destructive detection of total sugar in tobacco.

[0007] During the detection process, a hyperspectral model of total sugar content in tobacco was used, with the hyperspectral image of tobacco (powder) as input, to predict the total sugar content of tobacco.

[0008] To achieve the above objectives, the present invention adopts the following technical solution:

[0009] A rapid and non-destructive method for detecting total sugar content in tobacco based on hyperspectral technology is characterized by using a hyperspectral model of total sugar content in tobacco, taking the hyperspectral image of tobacco as input, to predict the total sugar content of tobacco.

[0010] The method for constructing the hyperspectral model of total sugar components in tobacco is as follows:

[0011] 1) Sample preparation: The tobacco sample was dried, pulverized, and then sealed.

[0012] 2) Hyperspectral data acquisition: Raw spectral images of powdered tobacco samples were acquired using a hyperspectral device;

[0013] 3) Image data processing: Delineate the ROI region from the original spectral image acquired in step 2) and extract the original spectral data;

[0014] 4) Data preprocessing: The raw spectral data extracted in step 3) were preprocessed using different preprocessing methods;

[0015] 5) Feature filtering: Different feature filtering methods are used to filter out useless feature wavelengths from the spectral data processed in step 4).

[0016] 6) Modeling: Using the spectral data processed in step 5), a regression model is established using the partial least squares regression (PLSR) method;

[0017] The optimal combination of different preprocessing methods in step 4), different feature screening methods in step 5), and the PLSR method is used as the hyperspectral model of total sugar components in tobacco. Attached Figure Description

[0018] Figure 1 Hyperspectral ROI image of tobacco;

[0019] Figure 2 The raw spectral image extracted from tobacco;

[0020] Figure 3 This is the schematic diagram of a PLSR.

[0021] Figure 4(a) shows the comparison prediction results of the optimal model SG-CARS-PLSR training set;

[0022] Figure 4(b) shows the best prediction results of the optimal model SG-CARS-PLSR on the test set. Detailed Implementation

[0023] The technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the protection scope of the present invention.

[0024] I. Scheme Description

[0025] A rapid and non-destructive hyperspectral method for detecting total sugar content in tobacco includes the following steps:

[0026] 1) Sample preparation: The tobacco sample was dried in an oven, cooled and then pulverized. The powder was sieved and sealed for storage.

[0027] 2) Hyperspectral data acquisition: First, use a 17×5.5mm quartz cuvette to hold the tobacco powder processed in step 1), then use a SOC-710 hyperspectral device to acquire the raw data, and use a 75% reflectance gray card to correct the data.

[0028] 3) Image data processing: The raw hyperspectral image data acquired and corrected in step 2) is processed to divide the ROI region from the spectral image and extract the raw spectral data;

[0029] 4) Data preprocessing: For the raw spectral data extracted in step 3), seven preprocessing methods are used: Savitzky-Golay filtering (SG), wavelet transform, detrending algorithm (SNV), first derivative, multivariate scattering correction (MSC), first-order +SG and wavelet +SG processing.

[0030] 5) Feature selection: For the spectral data processed in step 4), competitive adaptive reweighted sampling (CARS) and recursive feature elimination (RFE) feature selection methods are used to filter out useless feature wavelengths and improve model performance;

[0031] 6) Modeling: For the spectral data processed in step 5), a regression model is established using the partial least squares regression (PLSR) method.

[0032] The optimal combination of preprocessing methods, feature selection methods, and PLSR was used to construct a hyperspectral model of total sugar components in tobacco.

[0033] Step 3) of dividing the ROI region includes: "threshold segmentation - finding the centroid - drawing a new ROI with the centroid as the center and a radius of five pixels"; the principle of threshold segmentation is to find the useful region by adjusting the three feature data of the image (RGB).

[0034] Step 4) employs seven preprocessing methods and their combinations, including five basic preprocessing methods and two combined preprocessing methods, namely:

[0035] (a)SG: A smoothing and denoising method based on local polynomial least squares fitting. By selecting spectral data points within a fixed window, a low-order polynomial (usually 2-3 order) is used to fit the data within the window, and the fitted values ​​replace the original data points. This method can remove high-frequency noise while preserving key features such as the peak position and shape of the spectrum. The window size and polynomial order need to be adjusted according to the data characteristics.

[0036] (b)SNV: Standardize the spectrum of each sample by first calculating the mean and standard deviation of the spectrum of the sample, then subtract the mean from the spectral value of each wavelength point and divide by the standard deviation. This can eliminate the spectral baseline drift and amplitude variation caused by differences in particle size and optical path between samples, and is suitable for reducing the influence of scattering effects.

[0037] (c)MSC: Using the average spectrum of all samples as the baseline spectrum, perform linear regression on the spectrum of each sample (fitting the slope a of the baseline spectrum). i and intercept b i Then, the original spectrum is corrected using formula (1):

[0038]

[0039] It can eliminate spectral baseline shift and scaling caused by differences in sample scattering, making the shape of different sample spectra more consistent.

[0040] (d) First derivative: The first derivative curve of the spectrum is obtained by calculating the slope of the spectral values ​​at adjacent wavelengths; it can eliminate baseline drift, smooth background interference, and enhance the resolution of spectral peaks (highlighting the differences of overlapping peaks). Usually, it needs to be combined with SG smoothing first to avoid noise amplification.

[0041] (e) Wavelet transform: A preprocessing method based on multi-scale time-frequency analysis. By scaling (controlling frequency) and translating (controlling time / wavelength) the mother wavelet, the spectral signal is decomposed into approximation coefficients (low frequency, overall trend) and detail coefficients (high frequency, noise / feature) at different scales. Noise in the detail coefficients can be removed by thresholding, and the spectrum can be reconstructed to achieve denoising or feature enhancement. It is suitable for processing non-stationary spectral signals.

[0042] Step 5) uses competitive adaptive reweighted sampling (CARS) or recursive feature elimination (RFE) feature selection methods, the principles of which include:

[0043] (a) CARS: A heuristic feature selection method that simulates "natural selection". Its core is adaptive reweighted iteration, and the steps are as follows:

[0044] Initially, all features are assigned equal weights, and then a subset of features is obtained by random sampling according to the weights.

[0045] Use this subset to build a PLSR model and evaluate its performance using the cross-validation error (RMSECV).

[0046] Update feature weights based on the absolute value of the model regression coefficients (increase the weights of features that contribute significantly).

[0047] Repeated sampling, modeling, and weight updates are performed, with an exponential decay function controlling the proportion of features retained in each round.

[0048] Finally, the feature combination with the smallest RMSECV among all iterations is selected as the optimal subset.

[0049] (b) RFE: A wrapper-style feature selection method based on "iterative removal of weak features", the steps of which are as follows:

[0050] Train the baseline model (Partial Least Squares Regression PLSR) using all features;

[0051] Evaluate the importance of all features and remove the least important features;

[0052] Retrain the model using the remaining features, and repeat the evaluation-removal operation;

[0053] Record the model performance of each feature subset in each round, and select the feature subset with the best performance;

[0054] It can effectively eliminate redundant / irrelevant features, alleviating the curse of dimensionality and multicollinearity problems.

[0055] The modeling method in step 6) is PLSR; PLSR is a regression method suitable for high-dimensional, collinear data, and its core is latent variable extraction. The steps include:

[0056] Standardize the input variable X (spectral / feature) and the response variable Y (target attribute / total sugar);

[0057] Iterative extraction of latent variables (t) that maximize the covariance of X and Y k =Xw k u k =Yc k w k and c k (as weight vector);

[0058] Calculate the load vector p corresponding to the latent variable k (X to t) k (regression coefficients) q k (Y to t) k The regression coefficients are used to update the residuals of X and Y (X=Xt).k p k T Y=Yt k q k T );

[0059] The optimal number of latent variables K was determined through cross-validation, and the regression equation Y=XB(B=W(P) was constructed. T W) -1 Q T W / P / Q (weight / loading matrix); it can simultaneously achieve dimensionality reduction, modeling, and variable interpretation, and is suitable for scenarios where the sample size is smaller than the number of variables.

[0060] Specifically:

[0061] Data standardization (preprocessing): Standardize each column (each variable) of the independent variable matrix Xn×p and the dependent variable matrix Yn×q. Obtain the standardized matrix E0 by standardizing the independent variables using formula (2) and obtain the standardized matrix F0 by formula (3).

[0062]

[0063]

[0064] Extract the first pair of "related components" (scores t1 and u1) using formulas (4)-(5): t1 is the score (component) of X, w1 is the weight vector of X (p-dimensional); u1 is the score (component) of Y, c1 is the weight vector of Y (q-dimensional). The optimization objective is: maxu1,c1Cov(t1,u1), that is, to make t1 as related to the information of Y as possible.

[0065]

[0066]

[0067] After extracting the components, the original matrix is ​​split into "component contribution part + residual part":

[0068] For X(E0), the load p is calculated using formula (6). 1, Formula (7) is obtained, where E1 is the residual matrix after extracting t1;

[0069] For Y(F0), the load r1 is calculated using formula (8), resulting in formula (9), where F1 is the residual matrix after extracting t1;

[0070]

[0071]

[0072]

[0073]

[0074] Replace the original matrices E0 and F0 with residual matrices E1 and F1, and repeat formulas (4) to (9) to extract the 2nd, 3rd...mth components (t). h ,u h ), until the number of components m satisfies the generalization accuracy of cross-validation. After extracting m components, the standardized matrices of X and Y can be decomposed into formulas (10) to (11).

[0075]

[0076]

[0078] Steps 4), 5), and 6) constructed a raw spectral data processing workflow of "7 preprocessing methods - 2 feature screening methods - PLSR". The final result showed that SG-CARS-PLSR performed best, with a test set correlation coefficient (R) of 0.8695 (as shown in Figure 4(b)), successfully achieving non-destructive detection of total sugar components in tobacco.

[0079] II. Experiment

[0080] 1. Experimental Materials

[0081] A total of 115 tobacco samples were collected in this study. The samples came from Fuchuan Yao Autonomous County, Jianghua County, Ningyuan County, Lanshan County, Jiahe County, Chenzhou area, Ningxiang, Jietou and Yongzhou. To ensure the consistency of the test results, all samples underwent the same pretreatment operations such as drying and grinding, and specific tests were carried out on the total sugar content, a core physicochemical indicator.

[0082] Experimental instruments and equipment

[0083] The SOC-710 hyperspectral analyzer, manufactured by Surface Optics Corporation in the United States, served as the core optical equipment for this testing. Its stable performance and core parameters are well-suited to the needs of tobacco sample detection: a spectral range covering the 400-1000nm visible to near-infrared band, a spectral resolution of 2nm, and 1920 spatial channels and 720 spectral channels, enabling precise capture of spectral information for each pixel. The device employs a Si CCD detector with a 14-bit dynamic range and a full-frame frame rate of 53–109fps, efficiently acquiring continuous spectral data. Its excellent spectral resolution and imaging stability provide accurate data support for the correlation analysis of physicochemical indicators such as total sugar in tobacco.

[0084] The 17×5.5mm quartz cuvette possesses excellent UV transmittance and chemical stability, providing reliable conditions for the colorimetric quantitative detection of total sugar content.

[0085] 2. Experimental Methods and Results

[0086] 2.1 Hyperspectral Image Acquisition and Correction

[0087] 1) Equipment preparation: Start the SOC-710 hyperspectral device in advance to preheat it and ensure that the instrument’s spectral response is stable. The preheating time must meet the requirements of the equipment operation procedure.

[0088] 2) Sample preparation: The tobacco powder sample was loaded into a 17×5.5mm quartz cuvette. After loading, the sample was gently pressed to ensure that its surface was flat and uniform, and to avoid gaps or uneven accumulation.

[0089] 3) Data acquisition: Place the cuvette containing the flat sample precisely below the camera of the hyperspectral device, adjust to the optimal imaging distance, and start the acquisition program to obtain the spectral image of the sample.

[0090] 4) Data correction: The raw image data after acquisition needs to be corrected using a 75% reflectivity standard gray card to eliminate ambient light and instrument system errors and ensure data accuracy.

[0091] 2.2 Hyperspectral Data Processing

[0092] 2.2.1 Hyperspectral Image Data Extraction

[0093] The core step in hyperspectral image data extraction is the accurate delineation of the region of interest (ROI), and the specific steps are as follows:

[0094] 1) Threshold segmentation: As a basic step in ROI region localization, its core principle is to effectively separate the target region from the background region by adjusting the numerical parameters of the three feature channels of the image: red (R), green (G), and blue (B), thereby filtering out the image region containing effective information of the tobacco sample.

[0095] 2) Find the centroid: Calculate the centroid coordinates of the effective region obtained after threshold segmentation, and use them as the reference point for subsequent ROI region delineation.

[0096] 3) Define the new ROI: Using the calculated centroid as the center, draw a circular area with a radius of 5 pixels. This circular area is the new ROI area that will be used for the final extraction of hyperspectral data.

[0097] Finally, the original spectral data was extracted based on the ROI region.

[0098] 2.2.2 Data Preprocessing

[0099] To eliminate noise, baseline drift, and system interference in the raw spectral data and improve data reliability and subsequent modeling accuracy, the extracted spectral data needs to be processed using seven preprocessing methods. The methods and their core functions are as follows:

[0100] 1) Savitzky-Golay filter (SG): Achieves spectral denoising through polynomial smoothing, preserving spectral features while reducing random noise interference;

[0101] 2) Wavelet transform: Based on the characteristics of multi-scale analysis, it effectively separates spectral signals from noise and accurately removes interference components of different frequencies;

[0102] 3) Detrending algorithm (SNV): Corrects baseline drift caused by differences in sample granularity and uneven light scattering, and enhances the consistency of spectral data;

[0103] 4) First derivative: Highlights the characteristic absorption peaks of the spectral curve, eliminates baseline drift and background interference, and improves spectral resolution;

[0104] 5) Multivariate Scattering Correction (MSC): Corrects for differences in scattering from the sample surface, reducing systematic errors caused by different sample packing densities;

[0105] 6) First-order + SG: Combining the feature enhancement advantages of the first-order derivative with the denoising capability of SG filtering, it achieves the dual effect of "denoising-feature enhancement";

[0106] 7) Wavelet + SG: Wavelet transform is used to remove high-frequency noise, and then SG filtering is used to smooth the spectrum to further optimize data quality.

[0107] 2.2.3 Feature Filtering

[0108] Hyperspectral data is characterized by high dimensionality and a lot of redundant information. To reduce the computational cost of subsequent modeling and improve the generalization ability of the model, the following two feature filtering methods are used to filter out useless feature wavelengths:

[0109] 1) Competitive Adaptive Reweighted Sampling (CARS): Through adaptive weighting iteration, feature wavelengths with high correlation to the total sugar index of tobacco are selected, noise and irrelevant information are removed, and core feature variables are retained;

[0110] 2) Recursive Feature Elimination (RFE): Based on the ranking of feature importance, features with low contribution are recursively deleted, and feature subsets are gradually optimized to improve the model's prediction accuracy for the target indicator.

[0111] 2.2.4 Regression Model

[0112] For the selected characteristic wavelength data and the physicochemical detection values ​​of total sugar in tobacco, a partial least squares regression (PLSR) model was used to construct the correlation. This model combines the advantages of multiple linear regression and principal component analysis, effectively handling the multicollinearity problem among characteristic variables, and is suitable for quantitative modeling of hyperspectral characteristics and physicochemical indicators.

[0113] 2.2.5 Cross-validation and evaluation metrics

[0114] To avoid overfitting and objectively evaluate its generalization performance, a five-fold cross-validation method was used to validate the PLSR model. The dataset was divided into a training set:validation set:test set = 5:1:1. The core performance metrics for the model were set as follows:

[0115] 1) Coefficient of determination (R): Reflects the degree of fit between the model's predicted values ​​and the actual values. The closer the value is to 1, the better the fit.

[0116] 2) Root Mean Square Error (RMSEP): Measures the average deviation between the model's predicted value and the actual value. The smaller the value, the higher the accuracy of the model's prediction.

[0117] The algorithm combination corresponding to Figure 4(a) and (b) is SG-CARS-PLRS.

Claims

1. A rapid and non-destructive method for detecting total sugar content in tobacco based on hyperspectral technology, characterized in that: A hyperspectral model of total sugar content in tobacco was used, with hyperspectral images of tobacco as input, to predict the total sugar content of tobacco. The method for constructing the hyperspectral model of total sugar components in tobacco is as follows: 1) Sample preparation: The tobacco sample was dried, pulverized, and then sealed. 2) Hyperspectral data acquisition: Raw spectral images of powdered tobacco samples were acquired using a hyperspectral device; 3) Image data processing: Delineate the ROI region from the original spectral image acquired in step 2) and extract the original spectral data; 4) Data preprocessing: Different preprocessing methods were used to preprocess the raw spectral data extracted in step 3); 5) Feature filtering: Different feature filtering methods are used to filter out useless feature wavelengths from the spectral data processed in step 4). 6) Modeling: Using the spectral data processed in step 5), a regression model is established using the partial least squares regression (PLSR) method; The optimal combination of different preprocessing methods in step 4), different feature screening methods in step 5), and the PLSR method is used as the hyperspectral model of total sugar components in tobacco.

2. The rapid and non-destructive method for detecting total sugar content in tobacco based on hyperspectral technology according to claim 1, characterized in that: In step 3), the steps for dividing the ROI region are as follows: first, perform thresholding on the image to find the centroid; then, draw a new ROI with the centroid as the center and a radius of five pixels. Thresholding segmentation finds useful regions by adjusting the RGB three feature data of an image.

3. The rapid and non-destructive method for detecting total sugar content in tobacco based on hyperspectral technology according to claim 1, characterized in that: Step 4) uses five basic preprocessing methods and two combined preprocessing methods; the five basic preprocessing methods are as follows: (a) Savitzky-Golay filter SG: A smoothing and denoising method based on local polynomial least squares fitting. It selects spectral data points within a fixed window, fits the data within that window with a low-order polynomial, and replaces the original data points with the fitted values. (b) Detrending algorithm SNV: Standardize the spectrum of each sample by first calculating the mean and standard deviation of the spectrum of the sample, and then subtract the mean from the spectral value of each wavelength point and divide by the standard deviation. (c) Multivariate scattering correction (MSC): Using the average spectrum of all samples as the baseline spectrum, a linear regression is performed on the spectrum of each sample to fit the slope α of the baseline spectrum. i and intercept b i Then, the original spectrum is corrected using formula (1) to obtain: , (d) First derivative: The first derivative curve of the spectrum is obtained by calculating the slope of the spectral values ​​at adjacent wavelengths; (e) Wavelet transform: A preprocessing method based on multi-scale time-frequency analysis. The frequency and wavelength are controlled by the scaling and translation operations of the mother wavelet, respectively, and the spectral signal is decomposed into approximation coefficients and detail coefficients at different scales. The noise in the detail coefficients is removed by the thresholding method, and the spectrum is reconstructed to achieve denoising or feature enhancement. The approximation coefficients are used to characterize the overall trend of low frequency, and the detail coefficients are used to characterize the noise characteristics of high frequency. The two combined preprocessing methods are the combination of first derivative and SG, and the combination of wavelet transform and SG.

4. The rapid and non-destructive method for detecting total sugar content in tobacco based on hyperspectral technology according to claim 1, characterized in that: Step 5) Use competitive adaptive reweighted sampling (CARS) or recursive feature elimination (RFE) for feature selection, respectively: (a) CARS: A heuristic feature selection method that simulates "natural selection". Its core is adaptive reweighted iteration, and the steps are as follows: First, all features are initially assigned equal weights, and then a subset of features is obtained by random sampling according to the weights. Then, a partial least squares (PLSR) model is built using feature subsets, and the performance of the PLSR model is evaluated using cross-validation error RMSECV. Next, the feature weights are updated based on the absolute values ​​of the regression coefficients of the PLSR model; among them, the weights of features that contribute more are increased. Repeat the above three steps of sampling, modeling, and updating weights, and control the proportion of features retained in each round through an exponential decay function; Finally, the feature combination with the smallest RMSECV corresponding to the subset in all iterations is selected as the optimal feature subset; (b) RFE: A wrapper-style feature selection method based on "iterative removal of weak features", the steps of which are as follows: First, a baseline model is trained using all features; the baseline model is a partial least squares regression (PLSR) model. Then, assess the importance of all features and remove the features with the lowest importance; Next, the PLSR model is retrained using the remaining features, and the evaluation and elimination operations of the previous step are repeated. Finally, record the PLSR model performance for each feature subset and select the feature subset with the best performance.

5. The rapid and non-destructive method for detecting total sugar content in tobacco based on hyperspectral technology according to claim 1, characterized in that: Step 6) uses the Partial Least Squares Regression (PLSR) method, and the steps include: S1 Data Standardization: Standardizing the independent variable X and its matrix X n×p The dependent variable Y matrix n×q Each column is standardized, and the standardized matrix E0 is obtained by standardizing the independent variables using formula (2), and the standardized matrix F0 is obtained by standardizing the dependent variables using formula (3): , , In the formula, i and j represent the row and column indices in the matrix, respectively; X represents the spectral features, and Y represents the target attribute, total sugar. S2 extracts the first pair of related components (t1, u1) using formulas (4) to (5); t1 is the score of X, w1 is the p-dimensional weight vector of X; u1 is the score of Y, and c1 is the q-dimensional weight vector of Y. The optimization objective is max u1,c1 Cov(t1,u1) represents the method of associating t1 with information from Y as much as possible. , , After S3 extracts the components, E0 and F0 are separated into "component contribution part + residual part": For E0, the load p1 is calculated using formula (6), resulting in formula (7), where E1 is the residual matrix after extracting t1; , , For F0, the load r1 is calculated using formula (8), resulting in formula (9), where F1 is the residual matrix after extracting t1; , , S4 replaces the original matrices E0 and F0 with residual matrices E1 and F1 respectively, and repeats formulas (4) to (9) to extract the second and third pairs of related components (t). h ,u h ), until the number of components m satisfies the generalization accuracy of cross-validation; after extracting m components, the standardized matrices of X and Y are decomposed into formulas (10) to (11): , 。 6. The rapid and non-destructive method for detecting total sugar content in tobacco based on hyperspectral technology according to claim 1, characterized in that: In step 2), the spectral acquisition range is 400–1000 nm, and the data is corrected using a 75% reflectance gray card before being processed in step 3).

7. The rapid and non-destructive method for detecting total sugar content in tobacco based on hyperspectral technology according to claim 1, characterized in that: In the hyperspectral model of total sugar components in tobacco, the preprocessing method in step 4) is SG filtering; the feature selection method in step 5) is competitive adaptive reweighted sampling CARS; and the optimal evaluation method is the optimal correlation coefficient R of the test set.