Method for rapidly determining the particle size of a tobacco sample

By constructing a particle size discrimination model for tobacco leaf samples based on the XGBoost machine algorithm, the problem of inaccurate particle size assessment of unknown tobacco samples in near-infrared spectroscopy was solved, ensuring the accuracy of chemical composition prediction.

CN119738321BActive Publication Date: 2025-11-21CHINA TOBACCO YUNNAN IND
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411786810.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-12-06
Publication Date
2025-11-21
Estimated Expiration
2044-12-06

AI Technical Summary

Technical Problem

Existing technologies, when using near-infrared spectroscopy to analyze the particle size of unknown tobacco samples, cannot quickly and accurately assess whether they meet the reference samples for establishing prediction models, resulting in low accuracy in predicting chemical composition.

Method used

A particle size discrimination model for tobacco leaf samples was constructed using the XGBoost machine learning algorithm and cross-validation technology. By collecting, standardizing, and dimensionality-expanding the spectral data of tobacco leaf samples, a stable particle size discrimination model was established to quickly identify the particle size of unknown tobacco samples.

Benefits of technology

This technology enables rapid and accurate identification of tobacco sample particle size before near-infrared spectroscopy analysis, ensuring the accuracy of subsequent chemical composition prediction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119738321B_ABST
    Figure CN119738321B_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of tobacco processing, and particularly relates to a method for rapidly judging granularity of a tobacco sample, comprising the following steps: obtaining a tobacco sample particle spectrum and obtaining standardized tobacco sample particle spectrum data; performing dimension expansion processing on the tobacco sample particle spectrum to obtain different forms of tobacco sample particle spectrum data, splicing the same tobacco sample particle spectrum data based on a tobacco sample particle category to obtain tobacco sample particle spectrum dimension expansion data; based on an XGBoost machine algorithm and the tobacco sample particle spectrum dimension expansion data, constructing a tobacco sample granularity discrimination model; inputting unknown tobacco sample particle data into the stable tobacco sample granularity discrimination model to obtain a preliminary attribution category of the unknown tobacco sample particle; before analyzing the sample by using near-infrared spectroscopy technology, the granularity of the sample can be accurately identified to determine whether the granularity meets the requirements of the prediction model, thereby guaranteeing the accuracy of subsequent prediction of the chemical composition of the sample by using near-infrared spectroscopy technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of tobacco processing technology, and in particular to a method for rapidly determining the particle size of tobacco samples. Background Technology

[0002] Near-infrared spectroscopy is a rapid and non-destructive analytical method that primarily infers the physical and chemical properties of a sample (including, but not limited to, important indicators such as moisture content, protein content, and fat content) by measuring the sample's spectral response in the near-infrared region. Because near-infrared spectroscopy requires no sample preparation and can be performed directly, it not only improves detection efficiency but also reduces the use of chemical reagents, making it an important tool in modern quality control and process monitoring.

[0003] However, near-infrared spectroscopy is a technique that relies on the optical properties of the sample. The physical morphology of the sample, especially its particle size, directly affects the scattering and absorption characteristics of near-infrared light, thus impacting the quality of the spectral data and the accuracy of the results derived from it. Therefore, to improve the accuracy of sample analysis results, before analyzing the chemical composition of the sample, a model is typically used to predict and assess the sample's particle size to determine whether it conforms to the reference sample used to establish the prediction model.

[0004] Currently, when predicting and determining sample particle size, a calibration model is typically built using samples of a specific particle size. However, this model is usually only applicable to samples with the same or similar particle sizes. If the sample particle size differs significantly from the particle size used in the modeling, the predicted results will be biased. If the original model is used to directly predict the particle size of an unknown sample, the predicted results will be distorted, thus affecting the final analytical results. This is because different sample particle sizes result in different scattering effects, which in turn change the near-infrared spectral characteristics, affecting the applicability of the model.

[0005] Therefore, it is necessary to design a method for determining the particle size of tobacco samples to address the problem that before using near-infrared spectroscopy to analyze the particle size of unknown samples, it is not possible to quickly and accurately assess whether the particle size of tobacco samples conforms to the reference sample used to establish a prediction model, which leads to low accuracy in predicting the chemical composition of samples using near-infrared spectroscopy. Summary of the Invention

[0006] The novel objective of this invention is to propose a method for rapidly determining the particle size of tobacco samples. This addresses the problem that before using near-infrared spectroscopy to analyze the particle size of unknown samples, it is impossible to quickly and accurately assess whether the particle size of tobacco samples conforms to the reference sample used to establish a prediction model, which leads to low accuracy in subsequent predictions of the chemical composition of samples using near-infrared spectroscopy.

[0007] To solve the above-mentioned technical problems, the present invention adopts the following technical solution:

[0008] A method for rapidly determining the particle size of tobacco samples includes the following steps:

[0009] Step 1: Collect different types of tobacco leaf sample particles after drying, perform infrared spectral scanning on different types of tobacco leaf sample particles to obtain the spectra of tobacco leaf sample particles, and then standardize the spectra of tobacco leaf sample particles to obtain standardized spectral data of tobacco leaf sample particles.

[0010] Step 2: Dimension expansion processing of the tobacco sample particle spectrum is performed to obtain different forms of tobacco sample particle spectral data. Based on the tobacco sample particle category, the spectral data of the same tobacco sample particles are spliced ​​together to obtain the tobacco sample particle spectral expansion data.

[0011] Step 3: Based on the XGBoost machine algorithm and the expanded spectral data of tobacco leaf sample particles, construct a tobacco leaf sample particle size discrimination model, optimize the tobacco leaf sample particle size discrimination model using cross-validation technology, obtain the optimized tobacco leaf sample particle size discrimination model, train the optimized tobacco leaf sample particle size discrimination model, and obtain a stable tobacco leaf sample particle size discrimination model.

[0012] Step 4: Input the unknown tobacco sample particle data into a stable tobacco sample particle size discrimination model to obtain the preliminary classification of the unknown tobacco sample particles. Then, combine the spectral data of the unknown tobacco sample particles to accurately predict the preliminary classification of the unknown tobacco sample particles and obtain the final classification of the unknown tobacco sample particles.

[0013] Prioritizes Step 1, which involves collecting different types of tobacco leaf sample particles after drying, performing infrared spectral scanning on the different types of tobacco leaf sample particles to obtain the spectra of the tobacco leaf sample particles, and then standardizing the tobacco leaf sample particle spectra to obtain standardized tobacco leaf sample particle spectral data. Specifically, this includes the following steps:

[0014] Step 1.1: Collect tobacco leaf samples from different varieties and grades from different tobacco-producing areas to obtain different categories of tobacco leaf samples. Place the different categories of tobacco leaf samples in a constant temperature and humidity chamber at 22℃ and 35% humidity for 36 hours to obtain different categories of dried tobacco leaf samples.

[0015] Step 1.2: Select dry tobacco samples of the same category from different categories of dry tobacco samples and mix them. Divide the evenly mixed dry tobacco samples of the same category into three equal parts. Then, use a 1095Cyclotec (XF-98B) cyclone precision pulverizer to pulverize and screen the three equal parts of the same category of dry tobacco samples in sequence to obtain 0.425mm dry tobacco sample particles, 1mm dry tobacco sample particles and 3mm dry tobacco sample particles.

[0016] Step 1.3: Based on the different scanning parameters of different models of near-infrared spectrometers, the different models of near-infrared spectrometers are divided into nine different combinations. Then, the near-infrared spectrometers with nine different scanning parameter combinations are used to scan and collect 0.425mm dry tobacco leaf sample particles, 1mm dry tobacco leaf sample particles and 3mm dry tobacco leaf sample particles respectively to obtain the spectra of different tobacco leaf sample particles.

[0017] Step 1.4: Using linear interpolation, the particle spectra of each different tobacco leaf sample are converted into corresponding data to obtain standardized particle spectra of tobacco leaf samples under different scanning parameters.

[0018] Preferably, the different types of near-infrared spectrometers are the Matrix-I type Fourier transform near-infrared spectrometer, the MPA type Fourier transform near-infrared spectrometer, and the Antaris II type Fourier transform near-infrared spectrometer.

[0019] Prior to this, Step 2, which involves expanding the spectral dimensions of tobacco leaf sample particles to obtain spectral data of tobacco leaf sample particles in different forms, and then stitching together spectral data of the same type of tobacco leaf sample particles based on the particle category, to obtain expanded spectral data of tobacco leaf sample particles, specifically includes the following steps:

[0020] Step 2.1: Perform maximum and minimum normalization on the particle spectra of different tobacco leaf samples from Step 1.3 to obtain the particle spectra Spc of different tobacco leaf samples;

[0021] Step 2.2: After performing SNV correction on the particle spectra of different tobacco leaf samples in Step 1.3, perform maximum and minimum normalization to obtain the particle spectra of different tobacco leaf samples Spc_0.

[0022] Step 2.3: After performing first-order differential correction on the particle spectra of different tobacco leaf samples in Step 1.3, perform maximum-minimum normalization to obtain the particle spectra Spc_1 of different tobacco leaf samples.

[0023] Step 2.4: After performing second-order differential correction on the particle spectra of different tobacco leaf samples in Step 1.3, perform maximum-minimum normalization to obtain the particle spectra of different tobacco leaf samples Spc_2.

[0024] Step 2.5: By matching the particle names of different tobacco leaf samples, find the same spectral data of tobacco leaf sample particles from the different tobacco leaf sample particle spectra Spc in Step 2.1, the different tobacco leaf sample particle spectra Spc_0 in Step 2.2, the different tobacco leaf sample particle spectra Spc_1 in Step 2.3, and the different tobacco leaf sample particle spectra Spc_2 in Step 2.4, and stitch them together to obtain the expanded spectral data of tobacco leaf sample particles.

[0025] Prior to this, Step 3, based on the XGBoost machine learning algorithm and expanded spectral data of tobacco leaf samples, constructs a particle size discrimination model for tobacco leaf samples. Cross-validation is then used to optimize this model, resulting in an optimized particle size discrimination model. The optimized model is then trained to obtain a stable particle size discrimination model. Specifically, this includes the following steps:

[0026] Step 3.1: Input the expanded spectral data of tobacco leaf sample particles into the XGBoost machine algorithm to construct a particle size discrimination model for tobacco leaf samples;

[0027] Step 3.2: Using cross-validation technology, the learning rate, maximum tree depth, and regularization parameters in the tobacco sample particle size discrimination model are optimized and adjusted to obtain the optimized tobacco sample particle size discrimination model.

[0028] Step 3.3: Irradiate the same tobacco sample particles with different near-infrared spectrometers to obtain infrared spectral data of the tobacco sample particles. Input the infrared spectral data of the tobacco sample particles as a training set into the optimized tobacco sample particle size discrimination model in Step 3.2 to obtain a stable tobacco sample particle size discrimination model.

[0029] Prioritize that Step 4, which involves inputting the unknown tobacco sample particle data into a stable tobacco leaf sample particle size discrimination model to obtain the preliminary classification of the unknown tobacco sample particles, and then combining the spectral data of the unknown tobacco sample particles to accurately predict the preliminary classification of the unknown tobacco sample particles and obtain the final classification of the unknown tobacco sample particles, specifically includes the following steps:

[0030] Step 4.1: Input the unknown tobacco sample particle data into the stable tobacco sample particle size discrimination model. The stable tobacco sample particle size discrimination model judges the particle size category of the unknown tobacco sample and obtains the preliminary classification of the unknown tobacco sample particles.

[0031] Step 4.2: Use an infrared spectrometer to perform infrared spectral scanning on the unknown tobacco sample particles to obtain spectral data of the unknown tobacco sample particles;

[0032] Step 4.3: Based on the spectral data of unknown tobacco sample particles, accurately predict the preliminary classification of unknown tobacco sample particles to obtain the final classification of unknown tobacco sample particles.

[0033] Prior to this, Step 3 involves constructing a tobacco sample particle size discrimination model based on the XGBoost machine algorithm and expanded spectral data of tobacco sample particles. The model is then optimized using cross-validation technology to obtain an optimized tobacco sample particle size discrimination model. After training the optimized model to obtain a stable model, the process further includes periodically updating the optimized model to obtain an updated model.

[0034] Prioritize periodically updating the optimized tobacco sample particle size discrimination model to obtain an updated tobacco sample particle size discrimination model, including the following steps:

[0035] Input the data of tobacco leaf samples of the same category with unknown particle size into a stable tobacco leaf sample particle size discrimination model, and make a discrimination on the tobacco leaf samples of the same category with unknown particle size to obtain the predicted particle size category of the tobacco leaf samples of the same category.

[0036] The predicted particle size category of the same type of tobacco leaf sample is compared with the known particle size category of the same type of tobacco leaf sample. If the predicted particle size category of the same type of tobacco leaf sample differs significantly from the known particle size category of the same type of tobacco leaf sample, the stable tobacco leaf sample particle size discrimination model is updated periodically to obtain the updated tobacco leaf sample particle size discrimination model.

[0037] Compared with the prior art, the present invention has at least one of the following beneficial effects:

[0038] This invention provides a method for rapidly determining the particle size of tobacco samples. Before analyzing the sample using near-infrared spectroscopy, it can quickly and accurately identify whether the particle size of the sample meets the requirements of the prediction model. This ensures the accuracy of subsequent predictions of the chemical composition of the sample using near-infrared spectroscopy, thus solving the problem that before analyzing the particle size of unknown samples using near-infrared spectroscopy, it is impossible to quickly and accurately assess whether the particle size of tobacco samples meets the reference sample used to establish the prediction model, which leads to low accuracy in subsequent predictions of the chemical composition of the sample using near-infrared spectroscopy. Attached Figure Description

[0039] Figure 1This is a schematic diagram of a method for rapidly determining the particle size of tobacco samples according to the present invention.

[0040] Figure 2 This is the original Spc spectrum in this invention.

[0041] Figure 3 This is the spectral data standardization process and sample data dimension expansion spectrum in this invention.

[0042] Figure 4 This is a schematic diagram of the particle size discrimination model for tobacco leaf samples in the structure of the present invention.

[0043] Table 1 shows the different scanning parameters of the near-infrared spectrometer in this invention.

[0044] Table 2 is a table of spectral data in this invention. Detailed Implementation

[0045] like Figure 1-4 As shown, to make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the present invention and are not intended to limit the present invention.

[0046] Example

[0047] Currently, when predicting and determining sample particle size, a calibration model is typically built using samples of a specific particle size. However, this model is usually only applicable to samples with the same or similar particle sizes. If the sample particle size differs significantly from the particle size used in the modeling, the predicted results will be biased. If the original model is used to directly predict the particle size of an unknown sample, the predicted results will be distorted, thus affecting the final analytical results. This is because different sample particle sizes result in different scattering effects, which in turn change the near-infrared spectral characteristics, affecting the applicability of the model.

[0048] Therefore, this application proposes a method for rapidly determining the particle size of tobacco samples, which solves the problem that before using near-infrared spectroscopy to analyze the particle size of unknown samples, it is impossible to quickly and accurately assess whether the particle size of tobacco samples conforms to the reference sample used to establish a prediction model, thus leading to low accuracy in subsequent prediction of the chemical composition of samples using near-infrared spectroscopy.

[0049] For details, please refer to Figure 1 A method for rapidly determining the particle size of tobacco samples includes the following steps:

[0050] The first step involves collecting different types of tobacco leaf sample particles after drying, performing infrared spectral scanning on the different types of tobacco leaf sample particles to obtain the spectra of the tobacco leaf sample particles, and then standardizing the tobacco leaf sample particle spectra to obtain standardized tobacco leaf sample particle spectral data. This process includes the following steps:

[0051] 1) 29 tobacco leaf samples of different varieties and grades from different tobacco-producing areas in Yunnan Province were collected from 2020 to 2024. 29 tobacco leaf samples of different categories were obtained. The different categories of tobacco leaf samples were placed in a constant temperature and humidity chamber at 22℃ and 35% humidity for 36 hours to obtain 87 dry tobacco leaf samples of different categories.

[0052] 2) Select dry tobacco samples of the same category from different categories of dry tobacco samples and mix them. Divide the evenly mixed dry tobacco samples of the same category into three equal parts. Then, use a 1095Cyclotec (XF-98B) cyclone precision pulverizer to pulverize and screen the three equal parts of dry tobacco samples of the same category in sequence to obtain a total of 87 dry tobacco sample particles of 0.425mm, 1mm and 3mm.

[0053] It should be noted that when screening 0.425mm, 1mm, and 3mm dry tobacco leaf sample particles, the 1095Cyclotec(XF-98B) cyclone precision pulverizer can sequentially switch between three sieves with apertures of 0.425mm (40 mesh), 1mm, and 3mm to complete the screening of different dry tobacco leaf sample particles.

[0054] Specifically, when screening 3mm dry tobacco leaf sample particles, the 3mm sieve is transferred to a 1095Cyclotec (XF-98B) cyclone precision pulverizer. After screening the 3mm dry tobacco leaf sample particles, the 3mm sieve is removed from the 1095Cyclotec (XF-98B) cyclone precision pulverizer.

[0055] When screening 1mm dry tobacco leaf sample particles, the 1mm sieve is transferred to the 1095Cyclotec(XF-98B) cyclone precision pulverizer. After screening the 1mm dry tobacco leaf sample particles, the 1mm sieve is removed from the 1095Cyclotec(XF-98B) cyclone precision pulverizer.

[0056] When screening 0.425mm (40 mesh) dry tobacco leaf sample particles, the 0.425mm (40 mesh) sieve is switched to a 1095Cyclotec (XF-98B) cyclone precision pulverizer. After screening 0.425mm (40 mesh) 1mm dry tobacco leaf sample particles, the 1mm sieve is removed from the 1095Cyclotec (XF-98B) cyclone precision pulverizer.

[0057] 3) Based on the different scanning parameters (different scanning parameters: different resolutions and scan numbers of the infrared spectrometers) of 3 Matrix-I type Fourier transform near-infrared spectrometers, 2 MPA type Fourier transform near-infrared spectrometers, and 3 Antaris II type Fourier transform near-infrared spectrometers, the 3 Matrix-I type Fourier transform near-infrared spectrometers, 2 MPA type Fourier transform near-infrared spectrometers, and 3 Antaris II type Fourier transform near-infrared spectrometers were divided into nine different combinations. The 3 Matrix-I type Fourier transform near-infrared spectrometers were then used in nine different combinations. Nine different scanning parameter combinations were used with an infrared spectrometer, two MPA-type Fourier transform near-infrared spectrometers, and three Antaris II-type Fourier transform near-infrared spectrometers to scan and acquire 87 dry tobacco leaf samples, including 0.425 mm, 1 mm, and 3 mm dry tobacco leaf particles. The scanning wavenumber range was 4000-10000 cm⁻¹. The scanning parameters are shown in Table 1. After scanning, a total of 6264 spectral data points of different tobacco leaf sample particles were obtained, as shown in Table 2.

[0058] Table 1. Near-infrared spectrometer scanning parameters

[0059] serial number resolution Number of scans C1 4 32 C2 4 64 C3 4 128 C4 8 32 C5 8 64 C6 8 128 C7 16 32 C8 16 64 C9 16 128

[0060] Table 2 Spectral Data

[0061]

[0062] 4) By using linear interpolation, the particle spectra of each different tobacco leaf sample are converted into data of 4200-9800 cm⁻¹ with a resolution of 4 cm⁻¹, thus obtaining standardized particle spectra of tobacco leaf samples under different scanning parameters.

[0063] Step 2, please refer to Figure 2 and Figure 3 The process involves expanding the spectral dimensions of tobacco leaf samples to obtain spectral data of different types of tobacco leaf samples. Based on the particle type of the tobacco leaf samples, spectral data of the same type of tobacco leaf samples are stitched together to obtain expanded spectral data of tobacco leaf samples. The specific steps include:

[0064] 1) Perform maximum and minimum normalization on the particle spectra of different tobacco leaf samples to obtain the particle spectra Spc of different tobacco leaf samples;

[0065] 2) After performing SNV correction on the particle spectra of different tobacco leaf samples, the maximum and minimum normalization processes were then applied to obtain the particle spectra of different tobacco leaf samples, Spc_0.

[0066] 3) After performing first-order differential correction on the particle spectra of different tobacco leaf samples, the maximum-minimum normalization was performed to obtain the particle spectra Spc_1 of different tobacco leaf samples.

[0067] 4) After performing second-order differential correction on the particle spectra of different tobacco leaf samples in Step 1.3, perform maximum-minimum normalization to obtain the particle spectra of different tobacco leaf samples Spc_2.

[0068] 5) By matching the names of different tobacco leaf sample particles, find the same tobacco leaf sample particle spectral data from the different tobacco leaf sample particle spectral data Spc in Step 2.1, the different tobacco leaf sample particle spectral data Spc_0 in Step 2.2, the different tobacco leaf sample particle spectral data Spc_1 in Step 2.3, and the different tobacco leaf sample particle spectral data Spc_2 in Step 2.4, and stitch them together to obtain the expanded dimension data of tobacco leaf sample particle spectral data.

[0069] The third step involves constructing a particle size discrimination model for tobacco samples based on the XGBoost machine learning algorithm and expanded spectral data of tobacco leaf samples. This model is then optimized using cross-validation techniques to obtain an optimized particle size discrimination model. Finally, the optimized model is trained to obtain a stable particle size discrimination model. This process includes the following steps:

[0070] 1) Input the expanded spectral data of tobacco leaf samples into the XGBoost machine algorithm to construct a particle size discrimination model for tobacco leaf samples;

[0071] 2) Cross-validation technology was used to optimize and adjust the learning rate, maximum tree depth, and regularization parameters in the tobacco sample particle size discrimination model to obtain an optimized tobacco sample particle size discrimination model.

[0072] 3) Irradiate the same tobacco sample particles with different near-infrared spectrometers to obtain infrared spectral data of the tobacco sample particles. Input the infrared spectral data of the tobacco sample particles as the training set into the optimized tobacco sample particle size discrimination model to obtain a stable tobacco sample particle size discrimination model.

[0073] The fourth step involves inputting the unknown tobacco sample particle data into a stable tobacco sample particle size discrimination model to obtain the preliminary classification of the unknown tobacco sample particles. Then, by combining the spectral data of the unknown tobacco sample particles, the preliminary classification of the unknown tobacco sample particles is accurately predicted to obtain the final classification of the unknown tobacco sample particles.

[0074] 1) Input the unknown tobacco sample particle data into a stable tobacco sample particle size discrimination model. The stable tobacco sample particle size discrimination model judges the particle size category of the unknown tobacco sample and obtains the preliminary classification of the unknown tobacco sample particles.

[0075] 2) An infrared spectrometer was used to scan the unknown tobacco sample particles to obtain spectral data of the unknown tobacco sample particles;

[0076] 3) Based on the spectral data of unknown tobacco sample particles, the preliminary classification of unknown tobacco sample particles is accurately predicted, and the final classification of unknown tobacco sample particles is obtained.

[0077] It should be noted that in the third step, based on the XGBoost machine learning algorithm and the expanded spectral data of tobacco leaf samples, a particle size discrimination model for tobacco leaf samples is constructed. This model is then optimized using cross-validation techniques to obtain an optimized particle size discrimination model. After training this optimized model to obtain a stable particle size discrimination model, the process also includes periodically updating the optimized model to obtain an updated particle size discrimination model. This process specifically includes the following steps:

[0078] Input the data of the same type of tobacco leaf samples with unknown particle size into a stable tobacco leaf sample particle size discrimination model, and then discriminate the same type of tobacco leaf samples with unknown particle size to obtain the predicted particle size category of the same type of tobacco leaf samples.

[0079] The predicted particle size category of the same type of tobacco leaf sample is compared with the known particle size category of the same type of tobacco leaf sample. If the predicted particle size category of the same type of tobacco leaf sample differs significantly from the known particle size category of the same type of tobacco leaf sample, the stable tobacco leaf sample particle size discrimination model is updated periodically to obtain the updated tobacco leaf sample particle size discrimination model.

[0080] This application discloses a method for rapidly determining the particle size of tobacco samples. This method can quickly and accurately identify whether the particle size of a sample meets the requirements of a prediction model before analysis using near-infrared spectroscopy. This ensures the accuracy of subsequent predictions of the chemical composition of the sample using near-infrared spectroscopy. It addresses the problem that before analyzing the particle size of unknown samples using near-infrared spectroscopy, it is impossible to quickly and accurately assess whether the particle size of tobacco samples meets the reference sample used to establish the prediction model, leading to low accuracy in subsequent predictions of the chemical composition of the sample using near-infrared spectroscopy.

[0081] Although the invention has been described herein with reference to several illustrative embodiments, it should be understood that many other modifications and implementations can be devised by those skilled in the art, which will fall within the scope and spirit of the principles disclosed herein. More specifically, various variations and modifications can be made to the components and / or layout of the subject matter arrangement within the scope of the disclosure, drawings, and claims. Besides variations and modifications to the components and / or layout, other uses will be apparent to those skilled in the art.

Claims

1. A method for rapidly determining the particle size of tobacco samples, characterized in that: Includes the following steps: Step 1: Collect different types of tobacco leaf sample particles after drying, perform infrared spectral scanning on different types of tobacco leaf sample particles to obtain the spectra of tobacco leaf sample particles, and then perform standardization processing on the spectra of tobacco leaf sample particles to obtain standardized spectral data of tobacco leaf sample particles. Step 2: Dimensional expansion processing of the tobacco leaf sample particle spectra is performed to obtain spectral data of tobacco leaf sample particles in different forms. Based on the particle type of the tobacco leaf sample, spectral data of the same type of tobacco leaf sample particles are stitched together to obtain dimensional expansion data of the tobacco leaf sample particle spectra. This includes the following steps: Step 2.1: Perform maximum and minimum normalization on the particle spectra of different tobacco leaf samples from Step 1.3 to obtain the particle spectra Spc of different tobacco leaf samples; Step 2.2: After performing SNV correction on the particle spectra of different tobacco leaf samples in Step 1.3, perform maximum and minimum normalization to obtain the particle spectra of different tobacco leaf samples Spc_0. Step 2.3: After performing first-order differential correction on the particle spectra of different tobacco leaf samples in Step 1.3, perform maximum-minimum normalization to obtain the particle spectra Spc_1 of different tobacco leaf samples. Step 2.4: After performing second-order differential correction on the particle spectra of different tobacco leaf samples in Step 1.3, perform maximum-minimum normalization to obtain the particle spectra of different tobacco leaf samples Spc_2. Step 2.5: By matching the particle names of different tobacco leaf samples, find the same spectral data of tobacco leaf sample particles from the different tobacco leaf sample particle spectra Spc in Step 2.1, the different tobacco leaf sample particle spectra Spc_0 in Step 2.2, the different tobacco leaf sample particle spectra Spc_1 in Step 2.3, and the different tobacco leaf sample particle spectra Spc_2 in Step 2.4, and stitch them together to obtain the expanded spectral data of tobacco leaf sample particles. Step 3: Based on the XGBoost machine learning algorithm and expanded spectral data of tobacco leaf samples, a particle size discrimination model for tobacco leaf samples is constructed. Cross-validation is used to optimize the model, resulting in an optimized model. The optimized model is then trained to obtain a stable particle size discrimination model, including the following steps: Step 3.1: Input the expanded spectral data of tobacco leaf sample particles into the XGBoost machine algorithm to construct a particle size discrimination model for tobacco leaf samples; Step 3.2: Using cross-validation technology, the learning rate, maximum tree depth, and regularization parameters in the tobacco sample particle size discrimination model are optimized and adjusted to obtain the optimized tobacco sample particle size discrimination model. Step 3.3: Irradiate the same tobacco sample particles with different near-infrared spectrometers to obtain infrared spectral data of the tobacco sample particles. Input the infrared spectral data of the tobacco sample particles as a training set into the optimized tobacco sample particle size discrimination model in Step 3.2 to obtain a stable tobacco sample particle size discrimination model. Step 4: Input the unknown tobacco sample particle data into a stable tobacco sample particle size discrimination model to obtain the particle size of the unknown tobacco sample. Then, combine the unknown tobacco sample particle spectral data to accurately predict the preliminary classification of the unknown tobacco sample particles and obtain the final classification of the unknown tobacco sample particles.

2. The method for rapidly determining the particle size of tobacco samples according to claim 1, characterized in that: Step 1, collecting different types of tobacco leaf sample particles after drying, performing infrared spectral scanning on the different types of tobacco leaf sample particles to obtain the spectra of the tobacco leaf sample particles, and then standardizing the tobacco leaf sample particle spectra to obtain standardized tobacco leaf sample particle spectral data, specifically includes the following steps: Step 1.1: Collect tobacco leaf samples from different varieties and grades from different tobacco-producing areas to obtain different categories of tobacco leaf samples. Place the different categories of tobacco leaf samples in a constant temperature and humidity chamber at 22℃ and 35% humidity for 36 hours to obtain different categories of dried tobacco leaf samples. Step 1.2: Select dry tobacco samples of the same category from different categories of dry tobacco samples and mix them. Divide the evenly mixed dry tobacco samples of the same category into three equal parts. Then, use a 1095Cyclotec XF-98B cyclone precision pulverizer to pulverize and screen the three equal parts of the same category of dry tobacco samples in sequence to obtain 0.425mm dry tobacco sample particles, 1mm dry tobacco sample particles and 3mm dry tobacco sample particles. Step 1.3: Based on the different scanning parameters of different models of near-infrared spectrometers, the different models of near-infrared spectrometers are divided into nine different combinations. Then, the near-infrared spectrometers with nine different scanning parameter combinations are used to scan and collect 0.425mm dry tobacco leaf sample particles, 1mm dry tobacco leaf sample particles and 3mm dry tobacco leaf sample particles respectively to obtain the spectra of different tobacco leaf sample particles. Step 1.4: Using linear interpolation, the particle spectra of each different tobacco leaf sample are converted into corresponding data to obtain standardized particle spectra of tobacco leaf samples under different scanning parameters.

3. The method for rapidly determining the particle size of tobacco samples according to claim 2, characterized in that: The different models of near-infrared spectrometers are the Matrix-I type Fourier transform near-infrared spectrometer, the MPA type Fourier transform near-infrared spectrometer, and the Antaris II type Fourier transform near-infrared spectrometer.

4. The method for rapidly determining the particle size of tobacco samples according to claim 1, characterized in that: Step 4 involves inputting the unknown tobacco sample particle data into a stable tobacco sample particle size discrimination model to obtain a preliminary classification of the unknown tobacco sample particles. Then, by combining the spectral data of the unknown tobacco sample particles, a precise prediction is made of the preliminary classification of the unknown tobacco sample particles to obtain the final classification of the unknown tobacco sample particles. This specifically includes the following steps: Step 4.1: Input the unknown tobacco sample particle data into the stable tobacco sample particle size discrimination model. The stable tobacco sample particle size discrimination model judges the particle size category of the unknown tobacco sample and obtains the preliminary classification of the unknown tobacco sample particles. Step 4.2: Use an infrared spectrometer to perform infrared spectral scanning on the unknown tobacco sample particles to obtain spectral data of the unknown tobacco sample particles; Step 4.3: Based on the spectral data of unknown tobacco sample particles, accurately predict the preliminary classification of unknown tobacco sample particles to obtain the final classification of unknown tobacco sample particles.

5. The method for rapidly determining the particle size of tobacco samples according to claim 1, characterized in that: Step 3 involves constructing a particle size discrimination model for tobacco samples based on the XGBoost machine learning algorithm and expanded spectral data of tobacco sample particles. This model is then optimized using cross-validation techniques to obtain an optimized particle size discrimination model. After training the optimized model to obtain a stable particle size discrimination model, the process further includes periodically updating the optimized model to obtain an updated particle size discrimination model.

6. The method for rapidly determining the particle size of a tobacco sample according to claim 5, characterized in that: The process of periodically updating the optimized tobacco sample particle size discrimination model to obtain an updated tobacco sample particle size discrimination model includes the following steps: Input the data of the same type of tobacco leaf samples with unknown particle size into a stable tobacco leaf sample particle size discrimination model, and then discriminate the same type of tobacco leaf samples with unknown particle size to obtain the predicted particle size category of the same type of tobacco leaf samples. The predicted particle size category of the same type of tobacco leaf sample is compared with the known particle size category of the same type of tobacco leaf sample. If the predicted particle size category of the same type of tobacco leaf sample differs significantly from the known particle size category of the same type of tobacco leaf sample, the stable tobacco leaf sample particle size discrimination model is updated periodically to obtain the updated tobacco leaf sample particle size discrimination model.

Citation Information

Patent Citations

  • Method for detecting particle size distribution based on near infrared spectrum technology

    CN109738342A