A method for analyzing the availability of tobacco leaf near-infrared spectrum data
By preprocessing, selecting characteristic bands, and using multi-criteria similarity measurement on near-infrared spectral data of tobacco leaves, highly representative and low-redundancy data were selected, solving the problem of abnormal spectral data affecting the applicability of the model and improving the efficiency and accuracy of model training and detection.
Patent Information
- Application Number
- CN202411436819.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-15
- Publication Date
- 2025-12-16
- Estimated Expiration
- 2044-10-15
AI Technical Summary
In existing technologies, abnormal spectral data affects the applicability and accuracy of tobacco quality assessment models, and there is a lack of effective and universal methods for evaluating the availability of model data.
By performing correction preprocessing, characteristic band selection, interpolation, and multi-criteria similarity measurement on near-infrared spectral data of tobacco leaves, highly representative and low-redundancy data were selected. Usability analysis was then conducted in conjunction with historical data to eliminate unusable data.
This improves the training efficiency and accuracy of the data processing model, enhances the efficiency and accuracy of the detection process, and ensures the applicability and reliability of the model.
Smart Images

Figure CN119312105B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of tobacco leaf data availability analysis, and in particular to a method for availability analysis of near-infrared spectral data for general models. Background Technology
[0002] With the accelerating consolidation of the tobacco industry and major brands' focus on high-quality tobacco leaf resources, the integration and efficient utilization of tobacco leaf resources have become core challenges for the industry's development. Traditional tobacco leaf formulation research and quality control models are increasingly unable to meet the demands of modern production. Domestically and internationally, digital detection, intelligent control, and lean management technologies based on tobacco leaf quality indicators are being gradually adopted to address the complexity and diversity of the production process.
[0003] Near-infrared spectroscopy, due to its accuracy, stability, and reliability, is widely used in laboratory testing of chemical indicators in tobacco leaves and in online monitoring on production lines. This technology provides crucial support for the detection of aroma and quality grades in tobacco leaves. However, the effectiveness of spectral detection methods is determined by both data quality and the scientific validity of the algorithm. Abnormal spectral data will severely impact the applicability and accuracy of the model, thereby affecting the assessment of tobacco leaf quality.
[0004] Currently, practitioners have conducted extensive research on anomaly sample screening, but research on data usability evaluation for general models is relatively limited. Therefore, how to combine techniques such as feature band selection, spectral feature extraction, and spectral similarity matching to complete the usability evaluation of spectral data has become an urgent technical challenge to be solved. Summary of the Invention
[0005] To address the problems in existing technologies, this invention proposes a usability analysis method for near-infrared spectral data of tobacco leaves. Considering that abnormal tobacco leaf spectral data can severely impact the training efficiency and accuracy of application-side data processing models, this invention proposes to first perform usability analysis on the tobacco leaf spectral data to filter out problematic data, thereby improving the training efficiency and accuracy of the application-side data processing models. Simultaneously, for the application phase of the trained model, this invention's method can eliminate worthless data to be detected, improving the efficiency of the detection process.
[0006] The method for analyzing the usability of near-infrared spectral data of tobacco leaves according to the present invention includes the following steps:
[0007] Step 1: Perform calibration preprocessing on the near-infrared spectral data of the tobacco leaves to be evaluated;
[0008] Step 2: Select characteristic bands from the preprocessed data to identify bands with high representativeness and low redundancy;
[0009] Step 3: Perform interpolation on the standard spectral data of tobacco leaves stored in the spectral library to match the dimensionality of the data after band selection;
[0010] Step 4: Calculate the multi-criteria similarity measure between the data after band selection in Step 2 and the standard spectral data of tobacco leaves after interpolation in Step 3, and obtain the mean and variance of the multi-criteria similarity measure;
[0011] Step 5: Based on the results obtained in Step 4, perform a usability analysis of the near-infrared spectral data of tobacco leaves and remove unusable data.
[0012] Compared with the prior art, the present invention has the following beneficial effects:
[0013] 1) This invention preprocesses the near-infrared spectral data of tobacco leaves to be evaluated to correct for errors that may occur during the acquisition process (such as mechanical errors caused by the equipment), and selects a subset of characteristic bands from the preprocessed data to retain the physical information in the data to the greatest extent possible, thereby reducing data redundancy. The number of bands selected is generally preset by the virtual dimension in the data.
[0014] 2) This invention employs a multi-criteria fusion similarity measurement method to evaluate the usability of data. It comprehensively considers the autocorrelation coefficient, spectral angular distance, spectral information divergence, and Euclidean distance between data points, enabling it to effectively identify anomalous samples and complete the usability evaluation of spectral data.
[0015] 3) This invention incorporates historical data for usability evaluation. This historical data consists of near-infrared spectral data of tobacco leaves that were previously considered usable. Utilizing historical data improves the accuracy of usability evaluation, avoids potential overall data bias in a single batch of data to be evaluated, and fully reuses existing data resources. This invention can effectively eliminate abnormal tobacco leaf spectral data, complete the screening of usable spectral data, reduce or avoid unusable data, and improve the accuracy and applicability of subsequent general models during training and application. Attached Figure Description
[0016] Figure 1 This is a flowchart of the near-infrared spectral data availability analysis method of the present invention;
[0017] Figure 2 This is a spectral graph of three unavailable data points; the horizontal axis in the graph represents wavenumber (cm). -1 The vertical axis represents absorbance;
[0018] Figure 3 A spectral curve of tobacco leaf data;
[0019] Figure 4 This is a histogram distribution of the similarity measures between mixed tobacco samples and non-tobacco samples. Detailed Implementation
[0020] The present invention will be further described and illustrated below with reference to specific embodiments. The embodiments described are merely examples of the content of this disclosure and do not limit the scope of the invention. The technical features of each embodiment in the present invention can be combined accordingly, provided that there is no mutual conflict.
[0021] like Figure 1 As shown, the embodiments of the present invention mainly include the following steps:
[0022] Step 1: Near-infrared spectral data of the tobacco leaves to be evaluated Perform correction preprocessing; N represents the number of samples in the data, and D represents the number of bands.
[0023] The purpose of preprocessing is to correct for errors that may occur during the acquisition of spectra, including but not limited to geometric correction, atmospheric correction, radiometric calibration, and radiometric normalization.
[0024] Step 2: Select characteristic bands from the preprocessed data to identify bands with high representativeness and low redundancy. L represents the number of bands in the near-infrared spectral data after band selection; the purpose of characteristic band selection is to obtain the essential characteristic components of the data.
[0025] In some embodiments of the present invention, feature band selection refers to selecting a subset of features from the band set of the original data to retain the physical information in the image to the greatest extent possible, thereby reducing data redundancy. The number of bands selected is generally determined by the virtual dimension of the image. The physical information is selected from the generalized Rayleigh entropy of the bands, and the redundancy is selected from the covariance or cross-correlation matrix between the bands.
[0026] Step 3: Perform interpolation on the standard spectral data of tobacco leaves stored in the spectral library to match the dimensionality of the data after band selection;
[0027] Step 4: Calculate the multi-criteria similarity measure between the data after band selection in Step 2 and the standard spectral data of tobacco leaves after interpolation in Step 3, and obtain the mean and variance of the multi-criteria similarity measure;
[0028] Step 5: Based on the results obtained in Step 4, perform a usability analysis of the near-infrared spectral data of tobacco leaves and remove unusable data.
[0029] In a specific embodiment of the present invention, step 2 is implemented using the following method:
[0030] 2.1) Initialize the selected band set X S Empty; set counter k = 1; set the number of bands to be selected to L;
[0031] 2.2) From the candidate band set X UTraverse all unselected bands and calculate the score of the \(t\)-th band \(x\) t as \(score(x\) t ):
[0032]
[0033] where \(corr(x,x\) t ) represents the correlation between band \(x\) and the \(t\)-th band \(x\) t ;
[0034] 2.3) Select the band with the maximum score from the current set of candidate bands \(X\) U as the target band and add it to the set of selected bands \(X\) S , and delete this band from the set of candidate bands \(X\) U ;
[0035] 2.4) If \(k < L\) at this time, then increment \(k\) by 1 and return to step 2.2); otherwise, output the set of selected bands \(X\) S , and the data after feature band selection is represented as
[0036] In a specific embodiment of the present invention, step 4 is implemented by the following method:
[0037] Input the \(N\) samples \(g\) i , \(i = 1,2,\cdots,N\) contained in the near-infrared spectral data to be evaluated after band selection. Each sample \(g\) i is an \(L\)-dimensional vector containing the spectral reflectance of \(L\) bands. Calculate the four similarity measures between each sample \(g\) i and the tobacco leaf standard reference spectral data \(c\) p . Among them, the four similarity measures are the correlation coefficient, spectral angle distance, spectral information divergence, and Euclidean distance;
[0038] The specific calculation methods of the four similarity measures are as follows:
[0039] ① Correlation coefficient \(COR(g\) <00000
[0043] ② Spectral angular distance SAD(g) i ,c p ):
[0044]
[0045] ③ Spectral Information Divergence SID(g) i ,c p ):
[0046]
[0047] in They represent g after standardization. i and c p The spectral value in band l; and
[0048] ④ Euclidean distance Eud(g) i ,c p ):
[0049]
[0050] Based on four similarity measures, a multi-criteria fusion similarity measure is obtained; further, the mean and variance of the multi-criteria similarity measure of this batch of near-infrared spectral data are obtained.
[0051] The similarity measure for multi-criteria fusion is:
[0052]
[0053] This similarity metric effectively combines multiple distance metrics such as spectral angular distance, spectral information divergence, Euclidean distance, and correlation coefficient, improving the robustness and accuracy of the algorithm. The mean of the multi-criteria similarity metric is:
[0054]
[0055] In a specific embodiment of the present invention, step 5 is implemented using the following method:
[0056] The mean and variance of the multi-criteria similarity measure of the current batch of data are then averaged with the mean and variance of the multi-criteria similarity measure of the historical tobacco near-infrared spectral data to obtain the mean μ. p and variance mean σ p
[0057]
[0058] Where M represents the total number of data batches, and the total number of batches of historical tobacco near-infrared spectral data is M-1; M=1 means there is no historical data, and the historical data are tobacco near-infrared spectral data that have been previously screened by humans and algorithms and are considered to be usable for subsequent models. Let be the mean of the multi-criteria similarity of the i-th batch of data. Utilizing historical data can improve the accuracy of usability assessment and also fully reuse existing data resources. The near-infrared spectral data of tobacco leaves after removing unusable samples using the method of this invention will also be used as historical data storage after recalculating the mean and variance of the multi-criteria similarity measure, and will assist in the usability analysis of subsequent batches of data.
[0059] This invention utilizes the 3 sigma principle of the normal distribution, which, based on the assumption of a normal distribution, considers that most variation in the data is caused by random error, while outliers exceeding a certain range are considered gross errors. Specifically, this invention calculates the usability of the near-infrared spectral data of tobacco leaves to be evaluated, using the following evaluation function:
[0060] Γ min =μ p -3σ p
[0061] Γ max =μ p +3σ p
[0062]
[0063] Among them, S(g i ,c p Let Γ be the similarity measure of sample i fused by multiple criteria, and let Γ be the usability assessment threshold of the data to be evaluated. min Γ represents the minimum boundary threshold. max denoted as the maximum boundary threshold; F(i) represents the availability evaluation function for the i-th sample in the data to be evaluated, used to evaluate whether there is abnormal data in the data to be evaluated, i = 1, 2, ..., N, where 1 indicates that the data is available and 0 indicates that it is not available;
[0064] When F(i) = 0, it means that the sample detection error does not meet the standard, the data to be detected is unusable data, and it is removed from the data to be evaluated;
[0065] When F(i) = 1, it means that the sample is usable data, and the larger F(i) is, the stronger the usability of the data.
[0066] Embodiments of the present invention also provide a system for analyzing the availability of near-infrared spectral data of tobacco leaves, comprising:
[0067] The data preprocessing module is used to perform correction preprocessing on the near-infrared spectral data of the tobacco leaves to be evaluated;
[0068] The feature band selection module selects feature bands from the preprocessed data to select band data with high representativeness and low redundancy.
[0069] The spectral library preprocessing module performs interpolation on the standard spectral data of tobacco leaves stored in the spectral library to match the dimensionality of the data after band selection.
[0070] The spectral matching module is used to calculate the similarity measure of the multi-criteria fusion between the band-selected data and the interpolated standard spectral data of tobacco leaves, and to obtain the mean and variance of the multi-criteria similarity measure.
[0071] The data availability evaluation module is used for availability analysis of near-infrared spectral data of tobacco leaves, and to remove unusable data.
[0072] In one specific embodiment of the present invention, the input near-infrared spectral data includes 202 samples g. i i = 1, 2, ..., 202, each sample g i This is a 1609-dimensional vector containing spectral reflectance for 1609 bands. In this example, the histogram distribution of the similarity metric values between the mixed tobacco sample and the non-tobacco sample is as follows. Figure 4 As shown, after querying the sample labels, samples with values less than -2 are all non-tobacco samples. The histogram results show that the similarity measure is effective in distinguishing between tobacco and non-tobacco samples. Based on the calculated F(i) value, three data points were identified as unusable, and their spectral curves are shown below. Figure 2 As shown, querying the data labels reveals that the data consists of plastic, hemp rope, and weed powder, not tobacco leaf samples. All other valid samples in the dataset are tobacco leaf samples, and the spectral curves of the tobacco leaf samples are shown below. Figure 3 As shown, different colors represent different samples, by Figure 2 and Figure 3 It is evident that the spectral curves of tobacco leaves are significantly different from those of non-tobacco leaf samples, thus validating the effectiveness of the method.
[0073] The above-described embodiments are merely illustrative of several implementations of the present invention, and while the descriptions are specific and detailed, they should not be construed as limiting the scope of the present invention. Those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.
Claims
1. A method for usability analysis of near-infrared spectral data of tobacco leaves, characterized by comprising the following steps: Step 1: Perform calibration preprocessing on the near-infrared spectral data of the tobacco leaves to be evaluated; Step 2: Perform feature band selection on the preprocessed data to select bands with high representativeness and low redundancy; the feature band selection in Step 2 specifically includes: 21) Initialize the selected band set Empty; set counter k=1; set the number of bands to be selected to L; 22) From the set of candidate bands Iterate through all unselected bands and calculate the value of band t. fractions : ; in, Indicates band The magnitude of the correlation; 23) Select the current set of candidate bands The band with the highest score is added to the selected band set as the target band. In, and from the candidate band set Delete that band; 24) If at this time , then let And return to step 22); otherwise, output the set of selected bands. The data after selecting the characteristic bands is represented as follows: ; Step 3: Perform interpolation on the standard spectral data of tobacco leaves stored in the spectral library to match the dimensionality of the data after band selection; Step 4: Calculate the multi-criteria similarity measure between the data after band selection in Step 2 and the standard spectral data of tobacco leaves after interpolation in Step 3, and obtain the mean and variance of the multi-criteria similarity measure; Step 4 includes: The near-infrared spectral data to be evaluated after input band selection contains... Sample , Each sample For inclusion Spectral reflectance of each band 3D vector, calculate each sample To tobacco standard reference spectral data The four similarity measures include correlation coefficients. spectral angular distance Spectral information divergence and European distance ; A multi-criteria fusion similarity measure is obtained based on four similarity measures; further, the mean and variance of the multi-criteria similarity measure of this batch of near-infrared spectral data are obtained; The similarity measure obtained based on four similarity measures and multi-criteria fusion is as follows: ; The mean of the multi-criteria similarity measure is: ; Step 5: Based on the results obtained in Step 4, perform a usability analysis of the near-infrared spectral data of tobacco leaves and remove unusable data; Step 5 includes: The mean and variance of the multi-criteria similarity measure of the current batch of data are then averaged with the mean and variance of the multi-criteria similarity measure of the historical tobacco near-infrared spectral data to obtain the average mean. and variance mean ; ; ; Where M represents the total number of data batches, and the total number of batches of historical tobacco near-infrared spectral data is M-1; , Let be the mean of the multi-criteria similarity of the i-th batch of data; The usability of the near-infrared spectral data of the tobacco leaves to be evaluated is calculated using the following evaluation function: ; ; in, This represents the usability assessment threshold for the data to be evaluated, where Indicates the minimum boundary threshold. Indicates the maximum boundary threshold; Let represent the usability evaluation function for the i-th sample in the data to be evaluated, used to assess whether there are outliers in the data to be evaluated. Where 1 indicates that the data is available and 0 indicates that it is not available; when When the error rate is below the standard, it indicates that the data to be tested is unusable and should be removed from the data to be evaluated. when When, it indicates that the sample is usable data, and The larger the value, the greater the availability of the data.
2. The method according to claim 1, characterized in that, The correction preprocessing includes one or more of geometric correction, atmospheric correction, radiometric calibration, or radiometric normalization.
3. The method according to claim 1, characterized in that, The near-infrared spectral data of tobacco leaves to be evaluated in step 1 simultaneously includes multiple near-infrared spectral data samples of tobacco leaves, denoted as... N represents the number of samples in the data, and D represents the number of bands.
4. The method according to claim 1, characterized in that, Step 3) specifically involves interpolating the standard reference spectral data of tobacco leaves in the spectral library to obtain the standard reference spectral curve of tobacco leaves, so as to match the number of bands L of the near-infrared spectral data after the feature band selection.
5. The method according to claim 1, characterized in that, The specific calculation methods for the four similarity measures are as follows: ① Correlation coefficient : ; in and They are respectively and In the Band data, , The mean values of the near-infrared spectral data to be evaluated and the data in the spectral library in each band; ② Spectral angular distance : ; ③ Spectral information divergence : ; in , They represent the results after standardization. and The spectral value in band l; and , ; ④ European distance : 。 6. A system for analyzing the availability of near-infrared spectral data of tobacco leaves, implementing the method according to any one of claims 1-5, characterized in that... include: The data preprocessing module is used to perform correction preprocessing on the near-infrared spectral data of the tobacco leaves to be evaluated; The feature band selection module selects feature bands from the preprocessed data to select band data with high representativeness and low redundancy. The spectral library preprocessing module performs interpolation on the standard spectral data of tobacco leaves stored in the spectral library to match the dimensionality of the data after band selection. The spectral matching module is used to calculate the similarity measure of the multi-criteria fusion between the band-selected data and the interpolated standard spectral data of tobacco leaves, and to obtain the mean and variance of the multi-criteria similarity measure. The data availability evaluation module is used for availability analysis of near-infrared spectral data of tobacco leaves, and to remove unusable data.
Citation Information
Patent Citations
Alternative method for tobacco leaf and cigarette leaf group formula based on near infrared spectrum
CN109975238A
Tobacco leaf quality evaluation and comparison system and method
CN117990646A