SVM (Support Vector Machine) model-based lamina alcoholization quality discrimination method
Through the image analysis method based on the SVM model, the problem of difficulty in judging the suitability of tobacco leaves for aging was solved, rapid non-destructive detection was achieved, and the efficiency and quality of the tobacco leaf aging process were improved.
Patent Information
- Application Number
- CN202510882793.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-28
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies are unable to quickly and non-destructively determine the suitability of tobacco leaf aging, resulting in a time-consuming and labor-intensive tobacco leaf aging process that lacks a scientific basis.
An image analysis method based on the SVM model is used to establish a mapping relationship between tobacco leaf color characteristics and aging grades through image acquisition, preprocessing, feature extraction and model training, thereby realizing rapid and non-destructive detection of tobacco leaf aging process.
It realizes the rapid and non-destructive detection of tobacco leaf aging process, optimizes the aging process, improves tobacco leaf quality, and provides a scientific basis for quality control.
Smart Images

Figure CN120765597A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the tobacco industry tobacco sheet alcoholization process detection technology, and relates to a tobacco sheet alcoholization process discrimination method based on image analysis, in particular to a tobacco sheet alcoholization quality discrimination method based on an SVM model. BACKGROUND
[0002] The quality of tobacco raw materials is a key factor determining the quality of cigarette products. During the process of tobacco from harvesting to final processing into cigarette products, there are multiple links such as harvesting, curing, redrying, and alcoholization. The alcoholization process is an important link for improving the quality of tobacco. Alcoholization can reduce the green and harsh odor of tobacco itself, increase the aroma quality and quantity, reduce the content of harmful substances, and improve the smoking quality. Judging the alcoholization process and determining the appropriate period of alcoholization are of great significance for improving the utilization efficiency and industrial value of tobacco.
[0003] Due to the diversity of geographical location, climate conditions and warehouse types faced by each cigarette enterprise, there are obvious differences in the alcoholization conditions of tobacco, and the appropriateness of alcoholization cannot be simply judged by the storage time, so it is urgent to establish a rapid and non-destructive method for judging the appropriateness of tobacco alcoholization. At present, in order to determine the best alcoholization period of tobacco, enterprises need to continuously detect the appearance, chemical composition and sensory quality, which is time-consuming and laborious.
[0004] During the alcoholization process of tobacco, the degradation of plastid pigments, enzyme browning and non-enzyme browning reactions will cause the surface color of tobacco to gradually deepen, which provides a theoretical basis for us to construct a tobacco alcoholization process judgment based on image recognition. Through image acquisition and processing, a tobacco sheet color representation model is constructed to realize the quantification of the color depth of tobacco sheet, which can objectively evaluate the color of tobacco. At the same time, combined with machine vision technology, a large number of samples can be processed to improve efficiency and repeatability. SUMMARY
[0005] The purpose of the present application is to provide a tobacco sheet alcoholization process discrimination method based on image analysis, which is used for rapid and non-destructive detection of the alcoholization process of tobacco, so as to optimize the alcoholization process and improve the quality of tobacco.
[0006] The purpose of the present application is achieved by the following technical solutions:
[0007] A tobacco sheet alcoholization quality discrimination method based on an SVM model, comprising the following steps:
[0008] S1, under uniform illumination conditions, an image of tobacco to be detected is photographed, and a standard color card is placed in the same picture, the correspondence between the standard color value of the color card and the pixel value in the image is identified, the tobacco image is color-mapped and corrected, and the interference of illumination and imaging equipment on subsequent analysis is eliminated;
[0009] The collected images must be taken under 6500K LED light, using the Datacolor SpyderCHECKR24 color card and the same type of color card.
[0010] S2, threshold segmentation and edge contour extraction based on HSV color space: First, the image is converted to HSV color space, and the threshold range is set based on the hue (H), saturation (S), and brightness (V) characteristics of the tobacco leaves to generate an initial tobacco leaf area mask. Subsequently, all contours in the image are extracted using this mask, and the contour with the largest area is selected as the target tobacco leaf area. Based on the minimum circumscribed rectangle of this contour, the original image is accurately cropped to preliminarily separate the tobacco leaf image area. To further improve the regional purity, the cropped image is grayed, and brightness threshold segmentation is applied to eliminate the remaining black background pixels, ultimately obtaining a true pixel area containing only tobacco leaves.
[0011] S3: Divide the segmented tobacco leaf region into several subregions. Calculate the average value of the RGB channels for each subregion to obtain a preliminary color feature vector. After calculating the average value of the RGB channels, arrange each region in order and use it as an explanatory variable. Standardize all subregion features to ensure that different features have consistent dimensions and reduce errors caused by different numerical scales.
[0012] The standardization method is Z-score standardization, and its formula is:
[0013]
[0014] Where: x is the original data value, μ is the mean of the data, and σ is the standard deviation of the data.
[0015] In S4, use the train_test_split function to split the feature data and the alcoholization level labels into a training set and a test set in a ratio of 7:3. The training set is used for model parameter learning, and the test set is used to evaluate classification accuracy. The model recognition effect is measured using indicators such as accuracy, precision, recall, and F1 index.
[0016] In the training set, the train_X variable represents the explanatory variable, i.e., the tobacco leaf RGB value vector extracted from the image; and the train_Y variable represents the explained variable, i.e., the tobacco leaf aging degree label.
[0017] S5. Build a judgment model. Import the standardized tobacco leaf color features and corresponding aging grade labels into the Python environment and implement classification modeling using the SVC class in the scikit-learn library. Select an appropriate kernel function (such as RBF) and set hyperparameters (C, gamma, etc.) as needed to ensure that the model has good recognition ability for different aging grades.
[0018] Furthermore, the SVM model is modeled using radial basis function (RBF kernel function), and the values of hyperparameters C and gamma are automatically optimized through grid search (GridSearchCV). Specifically, grid search is performed on the hyperparameters C and gamma of the radial basis function kernel SVM model, and 5-fold cross validation is performed through GridSearchCV. The parameter search range is C∈{0.1,1,10,100,1000}, gamma∈{10 -3 ,10 -2 ,10 -1 ,1,10}, and use accuracy as the evaluation index, and finally select C_best and gamma_best with the highest cross-validation accuracy for model training.
[0019] S6: Model saving. After completing model training and preliminary evaluation, the parameters of the final trained model are saved so that it can be directly called later without retraining. The joblib library is used to save the trained model in .pkl format for subsequent calling in different environments.
[0020] S7: Model trial run. After the model is saved, it can be tested in a small-scale environment to verify the performance of the model under conditions similar to production or testing.
[0021] The advantages of the present invention are that it can realize rapid and non-destructive detection of the aging process of tobacco leaves, thereby optimizing the aging process and improving the quality of tobacco leaves. The present invention can effectively predict the aging degree of tobacco leaves and provide a scientific basis for tobacco quality control. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] Figure 1 It is a flow chart of the technical solution of the present invention. DETAILED DESCRIPTION
[0023] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the accompanying drawings. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention fall within the scope of protection of the present invention.
[0024] According to the present invention, a method for determining the aging quality progress of tobacco strips based on a Support Vector Machine (SVM) model includes capturing images of tobacco leaves to be tested and obtaining corresponding aging grade data. The image acquisition method involves photographing tobacco strip samples under 6500K LED lighting every three months after entering the aging warehouse, with a Datacolor SpyderCHECKR 24 standard colorimetric chart placed within the same frame. The corresponding aging grade data is obtained by determining the aging grade data for each image based on the physical, chemical, and sensory qualities of the tobacco strips according to established evaluation criteria. The aging grade data is generated by creating a dummy variable, using a scale of 0-2 to distinguish between "0" and "2" ("0" represents insufficient aging), "1" represents appropriate aging, and "2" represents excessive aging), forming a "aging progress" label.
[0025] The collected image data is then preprocessed by importing the open source computer vision library (opencv) in Python, finding the color correction card datacolor SpyderCHECKR24 as a reference image through the find_color_card function, inputting the reference image and the image to be corrected, applying the histogram of the color card in the reference image to the image to be corrected, and finally outputting the corrected image.
[0026] The corrected image is then separated from the background using edge detection to obtain a pure tobacco leaf area. The specific method is as follows:
[0027] This method uses an image segmentation strategy based on the HSV color space for the automatic extraction of tobacco leaf regions. First, the input image is converted from BGR space to HSV color space to more accurately distinguish color attributes. In HSV space, appropriate hue (H), saturation (S), and value (V) threshold ranges are set for the tobacco leaf color characteristics to construct a color mask, thereby extracting pixel regions that may belong to tobacco leaves. The mask is then used for contour extraction. The contour with the largest area is selected as the target region, and its minimum bounding rectangle is calculated to accurately crop the tobacco leaf image from the original image. To ensure the validity of the extracted region, the cropped image is further grayscaled, and a threshold is set to remove black pixels in the background, retaining only the actual tobacco leaf region for subsequent analysis.
[0028] Because the tobacco aging process involves mixing samples from different locations, in order to make the R, G, and B component values extracted from the image better represent the characteristics of the image at that time, the segmented tobacco leaf area is divided into 10 sub-areas, and the average value of the R, G, and B channels of each sub-area is calculated to obtain the preliminary color feature vector. The features of all sub-areas are normalized to ensure that different features have consistent dimensions and reduce errors caused by different numerical scales. The normalization formula is:
[0029]
[0030] x is the original data value. μ is the mean of the data. σ is the standard deviation of the data.
[0031] After constructing the model variables, we established a judgment model. We imported the standardized tobacco leaf color feature values and the corresponding aging grade labels into the Python environment and implemented classification modeling using the SVC class in the scikit-learn library. We selected an appropriate kernel function (RBF, Sigmoid) and set hyperparameters (C, gamma, etc.) as needed to ensure the model's ability to discriminate between different aging grades.
[0032] The training set used in the modeling is the data set from 2021-2022, and the test set can be the data set from 2023. The training set to test set ratio is 7:3. Random partitioning can also be used.
[0033] In the training set, a train_X variable is required to represent the explanatory variable, that is, the RGB value vector of the tobacco leaf extracted from the image, and a train_Y variable is required to represent the explained variable, that is, the degree of aging of the tobacco leaf. It should be noted that the train_Y variable can only be a single column of data, while the train_X variable can be multiple columns of data (each column represents the explanatory variable of a sub-region), and the number of rows of the train_X and train_Y variables should be the same. After training is completed, the trained model is saved as a .pkl format file using the joblib library and named model_test so that it can be called repeatedly in the future.
[0034] Next, we use the established SVM model model_test to test the test set. The scikit-learn package is used for testing.
[0035] The evaluation metrics are as follows:
[0036] Accuracy, Precision, Recall, and F1 Score are calculated using the following formulas:
[0037]
[0038] TP (True Positive): True positive examples, the number of positive examples that are predicted to be positive
[0039] TN (True Negative): True negative examples, the number of negative examples that are predicted to be negative and are actually negative
[0040] FP (False Positive): False positive examples, the number of false positives that are predicted to be positive but are actually negative
[0041] FN (False Negative): False negative examples, the number of false negative examples predicted to be negative but actually positive
[0042] The accuracy rate represents the proportion of correctly predicted samples to all samples.
[0043] Precision represents how many of the samples predicted to be positive are actually positive.
[0044] The recall rate represents how many samples that are actually positive are successfully identified as positive by the model.
[0045] The F1 score is the harmonic mean of precision and recall, and is used to comprehensively evaluate model performance.
[0046] The model evaluation results are as follows:
[0047]
[0048] This method performed stably on the training set, with all metrics around 80%, including an accuracy of 80.67%. Precision, recall, and F1 scores were also close, indicating that the model had sufficiently learned the training data. On the test set, the model maintained strong performance, with accuracy of 71.15%, precision of 70.91%, recall of 71.15%, and F1 score of 70.92%. The differences between these metrics were relatively small, demonstrating the model's robustness and stability in practical applications.
[0049] In summary, with the help of the above-mentioned technical solution of the present invention, the aging process of tobacco leaves in the aging warehouse can be quickly judged through image collection and machine learning technology, providing technical support for guiding the next step of use of the tobacco leaves.
[0050] The above descriptions are merely embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A method for distinguishing the aging quality of tobacco strips based on the SVM model, characterized in that: The following steps are involved: S1, captures an image of the tobacco leaf to be tested under uniform lighting conditions, and places a standard colorimetric chart on the same screen. By identifying the correspondence between the standard color values of the colorimetric chart and the pixel values in the image, color mapping correction is performed on the tobacco leaf image to eliminate interference from differences in lighting and imaging equipment on subsequent analysis; S2, threshold segmentation and edge contour extraction based on HSV color space: First, the image is converted to HSV color space, and the threshold range is set based on the hue (H), saturation (S), and brightness (V) characteristics of the tobacco leaves to generate an initial tobacco leaf area mask. Subsequently, all contours in the image are extracted using this mask, and the contour with the largest area is selected as the target tobacco leaf area. Based on the minimum circumscribed rectangle of this contour, the original image is accurately cropped to preliminarily separate the tobacco leaf image area. To further improve the regional purity, the cropped image is grayed, and brightness threshold segmentation is applied to eliminate the remaining black background pixels, ultimately obtaining a true pixel area containing only tobacco leaves. S3, the segmented tobacco leaf area is divided into several sub-areas, and the average value of the RGB three channels of each sub-area is calculated to obtain the preliminary color feature vector; all sub-area features are normalized to ensure that different features have consistent dimensions and reduce errors caused by different numerical scales; In S4, use the train_test_split function to split the feature data and the alcoholization level labels into a training set and a test set in a ratio of 7:
3. The training set is used for model parameter learning, and the test set is used to evaluate classification accuracy. The model recognition effect is measured using indicators such as accuracy, precision, recall, and F1 index. S5. Build a judgment model. Import the standardized tobacco leaf color features and corresponding aging grade labels into the Python environment and implement classification modeling using the SVC class in the scikit-learn library. Select an appropriate kernel function (such as RBF) and set hyperparameters (C, gamma, etc.) as needed to ensure that the model has good recognition ability for different aging grades. S6: Model saving: After completing model training and preliminary evaluation, the parameters of the final trained model are saved so that it can be directly called later without retraining; S7: Model trial run. After the model is saved, it can be tested in a small-scale environment to verify the performance of the model under conditions similar to production or testing.
2. The method for distinguishing the aging quality of tobacco strips based on the SVM model according to claim 1, characterized in that: The images collected in step S1 must be taken under 6500K LED light, using a datacolor SpyderCHECKR 24 color chart and a color chart of the same type.
3. The method for distinguishing the aging quality of tobacco strips based on the SVM model according to claim 1, characterized in that: The standardization method in step S3 is Z-score standardization, and its formula is: Where: x is the original data value, μ is the mean of the data, and σ is the standard deviation of the data.
4. The method for distinguishing the aging quality of tobacco strips based on the SVM model according to claim 1, characterized in that: After the average values of the three RGB channels are calculated in step S3, each region is arranged in sequence and used as an explanatory variable.
5. The method for distinguishing the aging quality of tobacco strips based on the SVM model according to claim 1, characterized in that: In step S5, the SVM model is modeled using a radial basis function (RBF kernel function), and the values of the hyperparameters C and gamma are automatically optimized through grid search (GridSearchCV).
6. The method for distinguishing the aging quality of tobacco strips based on the SVM model according to claim 1, characterized in that: In step S4, in the training set: the train_X variable represents the explanatory variable, that is, the tobacco leaf RGB value vector extracted from the image; the train_Y variable represents the explained variable, that is, the tobacco leaf aging degree label.
7. The method for distinguishing the aging quality of tobacco strips based on the SVM model according to claim 1, characterized in that: In step S6, the model is saved using the joblib library to save the trained model in .pkl format for subsequent use in different environments.
8. The method for determining the quality of aged tobacco strips according to claim 1, wherein: In step S5, grid search is performed on the hyperparameters C and gamma of the radial basis function kernel SVM model, and 5-fold cross validation is performed through GridSearchCV. The parameter search range is C∈{0.1,1,10,100,1000}, gamma∈{10 -3 ,10 -2 ,10 -1 ,1,10}, and use accuracy as the evaluation index, and finally select C_best and gamma_best with the highest cross-validation accuracy for model training.