Method for intelligently and accurately judging color, shape and state of tobacco leaves in baking process

Through the dual-modal fusion of near-infrared spectrum and image data and the feature attention model, the problem of insufficient accuracy of multimodal information fusion of tobacco leaf maturity was solved, the precise judgment of tobacco leaf maturity was achieved, the quality fluctuations of different modal data were adapted, and the accuracy and robustness of the judgment were improved.

CN120629064APending Publication Date: 2025-09-12KUNMING UNIV OF SCI & TECH

Patent Information

Application Number
CN202510903988.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-01
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

The accuracy of multimodal information fusion for tobacco leaf maturity in existing technologies is insufficient. Especially when tobacco varieties are diverse and the growing environment is complex, a single modality discrimination method is difficult to fully and accurately reflect the comprehensive characteristics of tobacco leaves. In addition, existing multimodal fusion methods lack an adaptive mechanism and cannot handle the problems of inconsistent data quality and unbalanced feature contributions between modalities.

Method used

A bimodal fusion method of near-infrared spectroscopy and image data, combined with principal component analysis (PCA) dimensionality reduction and a bimodal feature attention model, combined with a feature attention mechanism and dynamic weight adjustment, enables accurate identification of tobacco leaf maturity. The specific steps include collecting tobacco leaf near-infrared spectroscopy and image data, extracting internal chemical composition and external color and shape features, applying PCA dimensionality reduction, and using a bimodal feature attention model to weight feature importance. Finally, a tobacco leaf state identification model is constructed to output maturity grades.

Benefits of technology

It achieves accurate identification of tobacco leaf maturity, overcomes the limitations of traditional single-modal identification methods, and comprehensively characterizes the comprehensive state of tobacco leaves by integrating internal chemical composition and external color and shape characteristics, thereby improving the accuracy and robustness of identification and adapting to identification under conditions of quality fluctuations in different modal data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120629064A_ABST
    Figure CN120629064A_ABST
Patent Text Reader

Abstract

The invention provides a method for intelligently and accurately judging the color shape state of tobacco leaves in the curing process, and belongs to the technical field of tobacco leaf curing. The method comprises the steps that firstly, near infrared spectrum data and image data of the tobacco leaves are collected to establish a maturity grading standard, and internal chemical component characteristics and external chromaticity morphological characteristics of the tobacco leaves are extracted through preprocessing; respectively carrying out dimensionality reduction on the two types of data by utilizing a principal component analysis method to obtain principal component characteristics of the near infrared spectrum and principal component characteristics of the image; secondly, performing data fusion on the two types of features, performing feature importance weighting by applying a bimodal feature attention model, and learning different modal feature association and mapping relationships through a multi-head self-attention and mutual attention mechanism by the model; and according to the signal-to-noise ratio of the spectral data, the image definition score and the data consistency evaluation value, dynamically adjusting the modal weight, and finally constructing a tobacco state discrimination model based on weighted fusion data to realize accurate discrimination of the maturity grade of the tobacco.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of tobacco leaf baking, and specifically relates to a method for intelligently distinguishing the color and shape status of tobacco leaves during an intelligent and precise baking process. Background Art

[0002] Tobacco leaf curing is a critical step in tobacco production, and determining leaf maturity directly impacts curing results and leaf quality. Traditionally, tobacco leaf maturity assessment relies primarily on manual experience, using single indicators such as visual inspection of leaf color, palpation of leaf texture, or moisture content. With technological advancements, automated methods such as near-infrared spectroscopy and machine vision have been introduced to tobacco leaf maturity assessment. Near-infrared spectroscopy can rapidly detect the internal chemical composition of tobacco leaves, while machine vision technology can objectively quantify the appearance of tobacco leaves. However, single-modality methods have significant limitations. Near-infrared spectroscopy focuses on detecting the internal chemical composition of tobacco leaves, but is susceptible to environmental interference such as humidity fluctuations and lighting conditions, resulting in unstable spectral signal-to-noise ratios. Machine vision technology primarily relies on identifying external features of tobacco leaves, but struggles to capture changes in internal composition, and image acquisition quality is significantly affected by imaging conditions. These single-modality methods struggle to fully and accurately reflect the comprehensive characteristics of tobacco leaves, especially when tobacco varieties are diverse and growing in complex environments. Single-modality information accuracy is significantly limited. At present, there is no effective solution for the multimodal information fusion and discrimination technology of tobacco leaf maturity. Existing multimodal fusion methods mostly use mechanical fusion strategies such as simple feature splicing or weighted averaging. They lack an in-depth understanding of the characteristics of different modal data and adaptive fusion mechanisms, and are difficult to deal with the problems of inconsistent data quality and unbalanced feature contributions between modalities. Especially in the actual tobacco leaf baking production environment, due to unstable collection conditions, the quality of different modal data fluctuates. How to dynamically adjust the fusion strategy according to the data quality to achieve accurate discrimination of tobacco leaf maturity is a core technical problem that the existing technology urgently needs to solve. In other words, there is a technical problem in the existing technology that the accuracy of multimodal information fusion discrimination of tobacco leaf maturity is insufficient. Summary of the Invention

[0003] In view of this, the present invention provides a method for intelligently distinguishing the color and shape status of tobacco leaves during an intelligent and precise baking process, which can solve the technical problem in the prior art of insufficient accuracy in distinguishing the maturity of tobacco leaves by fusion of multimodal information.

[0004] The present invention is achieved as follows: the present invention provides a method for intelligently distinguishing the color and shape status of tobacco leaves during an intelligent and precise baking process, comprising: collecting tobacco leaf near-infrared spectral data and tobacco leaf image data, establishing a tobacco leaf maturity grading standard; pre-processing the tobacco leaf near-infrared spectral data to obtain pre-processed tobacco leaf near-infrared spectral data, extracting characteristic information of the internal chemical components of the tobacco leaves from the pre-processed tobacco leaf near-infrared spectral data; performing image processing on the tobacco leaf image data to extract tobacco leaf chromaticity parameters and tobacco leaf morphological characteristic parameters, and quantifying the external characteristics of the tobacco leaves; and using principal component analysis to compare the pre-processed tobacco leaf near-infrared spectral data with the tobacco leaf chromaticity parameters. The principal component characteristics of tobacco leaf near-infrared spectrum and tobacco leaf image are respectively subjected to dimensionality reduction processing to obtain principal component characteristics of tobacco leaf near-infrared spectrum and principal component characteristics of tobacco leaf image; the principal component characteristics of tobacco leaf near-infrared spectrum and tobacco leaf image are fused to obtain fused data, and the bimodal feature attention model is applied to weight the feature importance of the fused data to obtain weighted fused data; a tobacco leaf state discrimination model is constructed based on the weighted fused data, and the fusion weight adjustment function is used to optimize the parameters of the tobacco leaf state discrimination model; the tobacco leaf data to be tested is input into the tobacco leaf state discrimination model to realize the color and shape state discrimination of the tobacco leaf, and the tobacco leaf maturity grade discrimination result is output.

[0005] Among them, tobacco leaf near-infrared spectral data refers to the spectral reflectance data obtained by scanning tobacco leaf samples through a near-infrared spectrometer within the near-infrared band, which mainly reflects the absorption characteristics of groups such as NH bonds, CH bonds and OH bonds inside the tobacco leaves.

[0006] Among them, tobacco leaf image data refers to the RGB color images of tobacco leaves collected under standard light source conditions by a high-resolution digital camera, which contains visual feature information such as tobacco leaf color, tobacco leaf texture and tobacco leaf morphology.

[0007] Among them, preprocessing refers to the mathematical processing of tobacco leaf near-infrared spectral data such as standard normal variable transformation, multivariate scattering correction or first-order derivative to remove non-sample information interference such as baseline drift, scattering and noise.

[0008] The tobacco leaf chromaticity parameters refer to the hue, saturation, and brightness values ​​in the HSV color space extracted from the tobacco leaf RGB color image, which are used to quantitatively characterize the color characteristics of the tobacco leaf.

[0009] Among them, tobacco leaf morphological characteristic parameters refer to geometric morphological characteristics such as tobacco leaf area, tobacco leaf circumference, tobacco leaf aspect ratio, tobacco leaf roundness, tobacco leaf rectangularity and tobacco leaf vein distribution, which are used to quantitatively characterize tobacco leaf shape characteristics.

[0010] Among them, principal component analysis refers to a statistical analysis method that reduces high-dimensional data to low-dimensional data. It converts the original features into linearly independent new features through orthogonal transformation, called principal components, which are used to reduce data redundancy and retain key information.

[0011] Among them, data fusion refers to the splicing and combination of the principal component features of the tobacco leaf near-infrared spectrum and the principal component features of the tobacco leaf image at the feature level to form a new feature set that can comprehensively characterize the internal and external characteristics of the tobacco leaf.

[0012] Among them, the bimodal feature attention model refers to a deep learning model based on the attention mechanism, which realizes the adaptive fusion and weight distribution of multi-source data by learning the correlation and importance between different modal features.

[0013] Among them, the tobacco leaf state discrimination model refers to a classification model based on a deep convolutional neural network structure, which is used to receive weighted fusion data and output the tobacco leaf maturity grade discrimination results.

[0014] The bimodal feature attention model is structured as a multi-layer perceptron, comprising an input layer, multiple hidden layers, and an output layer. The hidden layers consist of an interactive attention module, a feature extraction module, and a feature fusion module. The interactive attention module utilizes multi-head self-attention and mutual attention mechanisms, the feature extraction module comprises a fully connected layer and an activation function, and the feature fusion module utilizes spatial attention and channel attention mechanisms. The training dataset for the bimodal feature attention model was established by collecting 3,000 tobacco leaf samples from various varieties, including Honghua Dajinyuan, K326, and Yunyan 87. These samples were categorized into three maturity levels: immature, moderately mature, and overmature, based on expert ratings of leaf maturity.

[0015] The fusion weight adjustment function is a mathematical function that dynamically adjusts the weight distribution of the interactive attention module in the dual-modal feature attention model based on the signal-to-noise ratio of the tobacco leaf near-infrared spectral data, the clarity score of the tobacco leaf image data, and the data consistency assessment value. The fusion weight adjustment function is calculated based on three parameters: the signal-to-noise ratio of the tobacco leaf near-infrared spectral data, the clarity score of the tobacco leaf image data, and the data consistency assessment value. It obtains a balance value between 0 and 1. When the balance value is in different ranges, different weight adjustment functions are used to adjust the modal feature weights.

[0016] Among them, the tobacco leaf maturity grade discrimination result refers to the tobacco leaf maturity classification result output by the tobacco leaf state discrimination model after processing the input tobacco leaf data to be tested, including three grades: immature, moderately mature and over-mature.

[0017] Furthermore, the fusion weight adjustment function adjusts the parameters of the interactive attention module in the dual-modal feature attention model, specifically adjusting the weight distribution of each attention head in the multi-head attention mechanism, and realizing feature selection and information fusion of different modal data by dynamically adjusting the attention score.

[0018] Among them, the signal-to-noise ratio of tobacco leaf near-infrared spectral data refers to the ratio of effective signal to noise in tobacco leaf near-infrared spectral data, which reflects the quality of spectral data. The calculation method is the effective signal intensity divided by the standard deviation of background noise.

[0019] Among them, the tobacco leaf image data clarity score refers to the quality evaluation index of tobacco leaf image data, which is obtained by calculating the weighted average of the sum of the image gradient amplitude and the sum of the Laplace transform amplitude, reflecting the texture details and edge clarity of the tobacco leaf image.

[0020] Among them, the data consistency assessment value refers to the correlation measurement between the two modal data. It is obtained by calculating the mutual information of the principal component characteristics of the tobacco leaf near-infrared spectrum and the principal component characteristics of the tobacco leaf image, reflecting the consistency of the two data sources in describing the characteristics of the same tobacco leaf sample.

[0021] Among them, when the balance value is in the range of 0 to 0.3, the logarithmic weight adjustment function is used to enhance the weight of the principal component features of the near-infrared spectrum of tobacco leaves; when the balance value is in the range of 0.3 to 0.7, the linear weight adjustment function is used to maintain the balanced weight of the principal component features of the near-infrared spectrum of tobacco leaves and the principal component features of the tobacco leaf image; when the balance value is in the range of 0.7 to 1, the exponential weight adjustment function is used to enhance the weight of the principal component features of the tobacco leaf image.

[0022] Among them, the tobacco leaf data to be tested refers to the comprehensive feature data of the tobacco leaf near-infrared spectrum data and tobacco leaf image data of the tobacco leaf samples that need to be judged for maturity after preprocessing, feature extraction, dimensionality reduction processing and data fusion.

[0023] This invention achieves accurate identification of tobacco leaf maturity through bimodal fusion of near-infrared spectroscopy and image data, combined with a feature attention mechanism and dynamic weight adjustment. The method first collects tobacco leaf near-infrared spectral and image data, extracts internal and external feature information through preprocessing, then fuses them using principal component analysis for dimensionality reduction. A bimodal feature attention model is then applied to weight the feature importance. Finally, a tobacco leaf state discrimination model is constructed to output the maturity grade discrimination result. This invention overcomes the limitations of traditional single-modal discrimination methods by fusing the internal chemical composition characteristics of tobacco leaves with their external color and shape features to comprehensively characterize the comprehensive state of tobacco leaves. In particular, the bimodal feature attention model employed is capable of adaptively learning the correlation and importance between features of different modalities, achieving deep cross-modal feature fusion. More importantly, the fusion weight adjustment function designed in this invention dynamically adjusts the modal weights based on the signal-to-noise ratio of the tobacco leaf near-infrared spectral data, the clarity score of the tobacco leaf image data, and the data consistency assessment value, effectively solving the discrimination problem under conditions of fluctuating quality of different modal data. BRIEF DESCRIPTION OF THE DRAWINGS

[0024] Figure 1 is a flow chart of the method of the present invention.

[0025] Figure 2 Schematic diagram of the tobacco leaf state discrimination model.

[0026] Figure 3 Schematic diagram of the bimodal feature attention model. DETAILED DESCRIPTION

[0027] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.

[0028] like Figure 1 FIG. 1 is a flow chart of a method for intelligently distinguishing the color and shape of tobacco leaves during an intelligent and precise baking process provided by the present invention. The method comprises the following steps:

[0029] S01. Collect tobacco leaf near-infrared spectral data and tobacco leaf image data to establish a tobacco leaf maturity grading standard;

[0030] S02, preprocessing the tobacco leaf near-infrared spectral data to obtain preprocessed tobacco leaf near-infrared spectral data, and extracting characteristic information of internal chemical components of the tobacco leaf from the preprocessed tobacco leaf near-infrared spectral data;

[0031] S03, performing image processing on the tobacco leaf image data to extract tobacco leaf colorimetric parameters and tobacco leaf morphological characteristic parameters, and quantifying tobacco leaf external characteristics;

[0032] S04, using principal component analysis to perform dimensionality reduction processing on the pre-processed tobacco leaf near-infrared spectrum data, the tobacco leaf chromaticity parameters, and the tobacco leaf morphological characteristic parameters, respectively, to obtain principal component characteristics of the tobacco leaf near-infrared spectrum and principal component characteristics of the tobacco leaf image;

[0033] S05, fusing the principal component features of the tobacco leaf near-infrared spectrum with the principal component features of the tobacco leaf image to obtain fused data, calculating the signal-to-noise ratio of the tobacco leaf near-infrared spectrum data, the clarity score of the tobacco leaf image data, and the data consistency evaluation value, and applying a bimodal feature attention model to weight the feature importance of the fused data to obtain weighted fused data;

[0034] S06. Constructing a tobacco leaf state discrimination model based on the weighted fusion data, and optimizing the tobacco leaf state discrimination model parameters using a fusion weight adjustment function;

[0035] S07. Input the tobacco leaf data to be tested into the tobacco leaf state discrimination model to discriminate the color and shape of the tobacco leaves, and output the tobacco leaf maturity grade discrimination result.

[0036] Among them, tobacco leaf near-infrared spectral data refers to the spectral reflectance data obtained by scanning tobacco leaf samples through a near-infrared spectrometer within the near-infrared band, which mainly reflects the absorption characteristics of groups such as NH bonds, CH bonds and OH bonds inside the tobacco leaves.

[0037] Among them, tobacco leaf image data refers to the RGB color image of tobacco leaves collected under standard light source conditions by a high-resolution digital camera, which contains visual feature information such as tobacco leaf color, tobacco leaf texture and tobacco leaf morphology.

[0038] Among them, preprocessing refers to the mathematical processing of tobacco leaf near-infrared spectral data such as standard normal variable transformation, multivariate scattering correction or first-order derivative to remove non-sample information interference such as baseline drift, scattering and noise.

[0039] The tobacco leaf chromaticity parameters refer to the hue, saturation, and lightness values ​​in the HSV color space extracted from the tobacco leaf RGB color image, which are used to quantitatively characterize the color characteristics of the tobacco leaf.

[0040] Among them, tobacco leaf morphological characteristic parameters refer to geometric morphological characteristics such as tobacco leaf area, tobacco leaf circumference, tobacco leaf aspect ratio, tobacco leaf roundness, tobacco leaf rectangularity and tobacco leaf vein distribution, which are used to quantitatively characterize tobacco leaf shape characteristics.

[0041] Among them, principal component analysis refers to a statistical analysis method that reduces high-dimensional data to low-dimensional data. It converts the original features into linearly independent new features through orthogonal transformation, called principal components, which are used to reduce data redundancy and retain key information.

[0042] Among them, data fusion refers to the splicing and combination of the principal component features of the tobacco leaf near-infrared spectrum and the principal component features of the tobacco leaf image at the feature level to form a new feature set that can comprehensively characterize the internal and external characteristics of the tobacco leaf.

[0043] Among them, the bimodal feature attention model refers to a deep learning model based on the attention mechanism, which realizes the adaptive fusion and weight distribution of multi-source data by learning the correlation and importance between different modal features.

[0044] Among them, the specific structure of the dual-modal feature attention model is a multi-layer perceptron structure, which includes an input layer, multiple hidden layers and an output layer. The hidden layer is composed of an interactive attention module, a feature extraction module and a feature fusion module. The interactive attention module is composed of multi-head self-attention and mutual attention mechanisms. The feature extraction module is composed of a fully connected layer and an activation function. The feature fusion module is composed of spatial attention and channel attention mechanisms. The input layer receives the principal component features of the near-infrared spectrum of tobacco leaves and the principal component features of the tobacco leaf image. The output layer generates weighted fusion data and feature weight distribution. The dual-modal feature attention model extracts the internal feature association of the single modality through the multi-head self-attention mechanism, learns the cross-modal feature mapping relationship through the mutual attention mechanism, and finally realizes the deep fusion and weight optimization of multimodal features through the feature fusion module.

[0045] Among them, the steps for establishing the training data set of the dual-modal feature attention model specifically include collecting a total of 3,000 tobacco leaf samples of different varieties, covering Honghua Dajinyuan, K326, and Yunyan 87 tobacco varieties, and dividing them into three levels: immature, moderately mature, and over-mature according to the expert rating of tobacco leaf maturity. For each sample, tobacco leaf near-infrared spectral data and tobacco leaf image data are collected at the same time. The collected tobacco leaf near-infrared spectral data and tobacco leaf image data are preprocessed and standardized, features are extracted and data cleaning is performed, abnormal samples and noise data are eliminated, and the training set and validation set are divided into 7 to 3 ratios to ensure that samples of each variety and maturity level are balanced, and a paired training sample set containing the principal component features of tobacco leaf near-infrared spectra, principal component features of tobacco leaf images and label information is constructed.

[0046] Among them, the steps of bimodal feature attention model training specifically include setting the number of iterations epoch to 200, the learning rate to 0.001, the batch size to 32, using the Adam optimizer for parameter optimization, and the loss function using a weighted combination of cross-entropy loss and contrastive learning loss. An early stopping mechanism is introduced to prevent overfitting, and the verification frequency is set to once every 5 epochs. When the accuracy does not improve after 10 consecutive verifications, the training is stopped. The optimal model parameters are saved according to the performance of the verification set. The learning rate decay strategy is adopted during training, and the learning rate is reduced to 0.8 times the original every 50 epochs. Finally, the model with the best performance on the verification set is selected as the pre-training model.

[0047] Among them, the fusion weight adjustment function refers to a mathematical function that dynamically adjusts the weight distribution of the interactive attention module in the dual-modal feature attention model according to the signal-to-noise ratio of the tobacco leaf near-infrared spectral data, the clarity score of the tobacco leaf image data and the data consistency evaluation value. It is used to optimize the dependence of the tobacco leaf state discrimination model on different modal data and improve the generalization ability and discrimination accuracy of the tobacco leaf state discrimination model.

[0048] Among them, the signal-to-noise ratio of tobacco leaf near-infrared spectral data refers to the ratio of effective signal to noise in tobacco leaf near-infrared spectral data, which reflects the quality of spectral data. The calculation method is the effective signal intensity divided by the standard deviation of background noise.

[0049] Among them, the tobacco leaf image data clarity score refers to the quality evaluation index of tobacco leaf image data, which is obtained by calculating the weighted average of the sum of the image gradient amplitude and the sum of the Laplace transform amplitude, reflecting the texture details and edge clarity of the tobacco leaf image.

[0050] Among them, the data consistency assessment value refers to the correlation measurement between the two modal data. It is obtained by calculating the mutual information of the principal component characteristics of the tobacco leaf near-infrared spectrum and the principal component characteristics of the tobacco leaf image, reflecting the consistency of the two data sources in describing the characteristics of the same tobacco leaf sample.

[0051] Among them, the fusion weight adjustment function is calculated based on three parameters: the signal-to-noise ratio of tobacco leaf near-infrared spectrum data, the clarity score of tobacco leaf image data, and the data consistency evaluation value, to obtain a balance value ranging from 0 to 1. When the balance value is in the range of 0 to 0.3, the logarithmic weight adjustment function is used to enhance the weight of the principal component features of the tobacco leaf near-infrared spectrum; when the balance value is in the range of 0.3 to 0.7, the linear weight adjustment function is used to maintain the balanced weight of the principal component features of the tobacco leaf near-infrared spectrum and the principal component features of the tobacco leaf image; when the balance value is in the range of 0.7 to 1, the exponential weight adjustment function is used to enhance the weight of the principal component features of the tobacco leaf image, and the adaptability and discrimination robustness of the tobacco leaf state discrimination model to data of different qualities are improved through dynamic weight distribution.

[0052] Among them, the fusion weight adjustment function adjusts the parameters of the interactive attention module in the dual-modal feature attention model, specifically adjusts the weight distribution of each attention head in the multi-head attention mechanism, and realizes feature selection and information fusion of different modal data by dynamically adjusting the attention score. When the balance value is low, the attention weight corresponding to the principal component characteristics of the tobacco leaf near-infrared spectrum is enhanced, and the influence of the principal component characteristics of the tobacco leaf image is weakened; when the balance value is high, the attention weight corresponding to the principal component characteristics of the tobacco leaf image is enhanced, and the influence of the principal component characteristics of the tobacco leaf near-infrared spectrum is weakened; when the balance value is in the middle area, the balanced contribution of the principal component characteristics of the tobacco leaf near-infrared spectrum and the principal component characteristics of the tobacco leaf image is maintained, and dynamic optimization of feature selection is achieved by adjusting the weight coefficient of the attention head.

[0053] Among them, the tobacco leaf state discrimination model refers to a classification model based on a deep convolutional neural network structure, which is used to receive weighted fusion data and output the tobacco leaf maturity grade discrimination results.

[0054] Among them, the tobacco leaf data to be tested refers to the comprehensive feature data of the tobacco leaf near-infrared spectrum data and tobacco leaf image data of the tobacco leaf samples that need to be judged for maturity after preprocessing, feature extraction, dimensionality reduction processing and data fusion.

[0055] Among them, the tobacco leaf maturity grade discrimination result refers to the tobacco leaf maturity classification result output by the tobacco leaf state discrimination model after processing the input tobacco leaf data to be tested, including three grades: immature, moderately mature and over-mature.

[0056] The specific implementation of the above steps is described in detail below.

[0057] The specific implementation of step S01 involves collecting data from tobacco leaf samples using a near-infrared spectrometer and a high-resolution digital camera, and establishing a tobacco leaf maturity grading standard. The specific implementation process involves first scanning tobacco leaf samples using a near-infrared spectrometer with a wavelength range of 800 to 2500 nm and a sampling interval of 2 nm. Spectra are collected 32 times for each sample and averaged to eliminate random errors, thereby obtaining tobacco leaf near-infrared spectral data. Simultaneously, a high-resolution digital camera with a resolution of at least 3000 × 2000 pixels is used to capture color images of the tobacco leaves at a vertical distance of 30 cm under standard D65 light conditions with an illumination intensity of 800 Lux and a color temperature of 5500 K. Images of both the front and back of each sample are captured. By hiring at least five experts with at least 10 years of tobacco leaf grading experience, a three-level tobacco leaf maturity grading standard was established: immature, moderately mature, and overmature, based on leaf color, aroma, oil content, tissue structure, and elasticity. The purpose of this step is to obtain multimodal data that comprehensively characterizes tobacco leaf characteristics, providing basic data support for subsequent analysis.

[0058] The specific implementation method of step S02 is to preprocess the tobacco leaf near-infrared spectral data to obtain preprocessed tobacco leaf near-infrared spectral data and extract the characteristic information of the internal chemical components of the tobacco leaf. The specific implementation process is to first perform a standard normal variable transformation on the original spectral data to eliminate the influence of baseline drift. Then, the spectral data is smoothed using the Savitzky-Golay smoothing algorithm, with the window width set to 15 wavelength points and the polynomial order to 3 to reduce spectral noise. Then, a multivariate scattering correction algorithm is used to eliminate the scattering effect caused by the uneven sample particle size. The correction parameters are selected from the optimal values ​​obtained by iterative optimization. Then, a first-order derivative transformation is applied to enhance the spectral characteristics, and the first-order derivative of the spectrum is calculated using the difference method with a differential interval of 5nm. Finally, an interval selection algorithm is used to screen out wavelength ranges with high correlation with the chemical components of the tobacco leaf. These ranges mainly include three ranges: 1100 to 1200nm, 1400 to 1600nm, and 1900 to 2100nm. These ranges correspond to the absorption characteristics of groups such as NH bonds, CH bonds, and OH bonds in the tobacco leaf, respectively. The purpose of this step is to eliminate non-sample information interference, enhance the characteristic information of the internal chemical components of tobacco leaves, and improve the accuracy of subsequent analysis.

[0059] The specific implementation method of step S03 is to perform image processing on the tobacco leaf image data to extract tobacco leaf color parameters and tobacco leaf morphological characteristic parameters, and quantify the external characteristics of the tobacco leaves. The specific implementation process is to first pre-process the original image using the Gaussian filter algorithm, with a filter kernel size of 5×5 and a standard deviation of 1.5 to remove image noise. Then, the Otsu threshold segmentation algorithm is applied to extract the tobacco leaf contour. The segmentation threshold is automatically calculated, generally between grayscale values ​​45 and 65, to separate the tobacco leaves from the background. Then, the RGB color space is converted to the HSV color space, and the values ​​of the three channels of hue, saturation, and lightness are extracted. The mean, standard deviation, skewness, and kurtosis of each channel in the tobacco leaf area are calculated to form a 12-dimensional chromaticity feature vector. Subsequently, based on the tobacco leaf contour obtained by segmentation, the geometric morphological features such as tobacco leaf area, perimeter, aspect ratio, roundness, and rectangularity are calculated, where the roundness threshold is set to 0.5. A value lower than this indicates that the tobacco leaf shape is irregular. Finally, the Goberta algorithm was used to extract the vein distribution characteristics of the tobacco leaves. Parameters such as the number of main veins, the number of lateral veins, the vein density, and the vein orientation angle were extracted to form an 8-dimensional morphological feature vector. The purpose of this step was to quantitatively characterize the external color and shape of the tobacco leaves and obtain characteristic parameters that reflect the appearance quality of the tobacco leaves.

[0060] The specific implementation of step S04 is to use principal component analysis to perform dimensionality reduction on the pre-processed tobacco leaf near-infrared spectral data, tobacco leaf colorimetric parameters, and tobacco leaf morphological characteristic parameters, respectively, to obtain the principal component features of the tobacco leaf near-infrared spectrum and the principal component features of the tobacco leaf image. The specific implementation process is to first perform principal component analysis on the pre-processed tobacco leaf near-infrared spectral data, calculate the eigenvalues ​​and eigenvectors, sort them in descending order of eigenvalue, and select the top K principal components with a cumulative contribution rate of 95%, typically with a K value of 8 to 12, as the principal component features of the tobacco leaf near-infrared spectrum. Then, the tobacco leaf colorimetric parameters and tobacco leaf morphological characteristic parameters are combined to form a 20-dimensional feature vector. Dimensionality reduction is also performed using principal component analysis, and the top M principal components with a cumulative contribution rate of 90%, typically with an M value of 5 to 8, are selected as the principal component features of the tobacco leaf image. Finally, the obtained principal component features are normalized to unify the data range of each principal component to between 0 and 1, facilitating subsequent data fusion. The purpose of this step is to reduce data redundancy through dimensionality reduction, extract the most representative internal and external feature information of the tobacco leaf, improve computational efficiency, and avoid the curse of dimensionality problem.

[0061] The specific implementation of step S05 is to fuse the principal component features of the tobacco leaf near-infrared spectrum with the principal component features of the tobacco leaf image to obtain fused data, calculate data quality assessment indicators, and apply a bimodal feature attention model to weight the feature importance of the fused data. The specific implementation process is to first use a feature cascade method to connect the principal component features of the tobacco leaf near-infrared spectrum with the principal component features of the tobacco leaf image in the feature dimension to form a K+M dimensional fused feature vector. Then, the signal-to-noise ratio of the tobacco leaf near-infrared spectrum data is calculated by dividing the average signal intensity of the main spectral band by the standard deviation of the background noise. The signal-to-noise ratio threshold is set to 15, and values ​​below this value indicate poor data quality. Next, the clarity score of the tobacco leaf image data is calculated using the Laplace operator. The clarity score threshold is set to 0.6, and values ​​below this value indicate poor image quality. Then, the data consistency assessment value is calculated. The correlation between the two modal features is calculated using the normalized mutual information method. The consistency threshold is set to 0.4, and values ​​below this value indicate inconsistent representations of the two data sources. Finally, the fused data and three evaluation metrics are fed into a pre-trained bimodal feature attention model. This model learns the correlation and importance between features from different modalities and adaptively adjusts feature weights to generate weighted fused data. This step aims to effectively fuse multi-source tobacco leaf data, dynamically adjusting the contribution ratio of different modal data based on data quality, and improving the expressiveness and discriminative performance of the fused features.

[0062] The specific implementation of step S06 is to construct a tobacco leaf state discrimination model based on weighted fusion data and optimize the parameters of the tobacco leaf state discrimination model using a fusion weight adjustment function. The specific implementation process is to first use the weighted fusion data as input to construct a tobacco leaf state discrimination model based on a deep convolutional neural network. The network structure includes an input layer, four convolutional layers, two fully connected layers, and an output layer. The convolution kernel size is 3×3, the number of convolution kernels in each layer is 32, 64, 128, and 256, respectively, the number of neurons in the fully connected layers is 1024 and 512, and the number of neurons in the output layer is 3, corresponding to the three maturity levels. Then, a balance value is calculated based on the signal-to-noise ratio of the tobacco leaf near-infrared spectral data, the clarity score of the tobacco leaf image data, and the data consistency assessment value. The balance value calculation formula is the weighted sum of the three indicators divided by the weight sum, with the weights set to 0.4, 0.4, and 0.2, respectively. Next, a corresponding weight adjustment function is selected based on the balance value: when the balance value is between 0 and 0.3, a logarithmic weight adjustment function is used to enhance the weights of the principal component features of the tobacco leaf near-infrared spectrum; when the balance value is between 0.3 and 0.7, a linear weight adjustment function is used to maintain balanced weights; and when the balance value is between 0.7 and 1, an exponential weight adjustment function is used to enhance the weights of the principal component features of the tobacco leaf image. Finally, the weight adjustment function is used to adjust the parameters of the interactive attention module in the bimodal feature attention model to optimize the tobacco leaf state discrimination model's dependence on data from different modalities. The goal of this step is to build a high-performance tobacco leaf state discrimination model and improve its adaptability and robustness through dynamic weight adjustment.

[0063] The specific implementation method of step S07 is to input the tobacco leaf data to be tested into the tobacco leaf state discrimination model, realize the color and shape state discrimination of the tobacco leaf, and output the tobacco leaf maturity grade discrimination result. The specific implementation process is to first collect near-infrared spectral data and image data of the tobacco leaf sample to be tested, using the same equipment and parameters as step S01. Then, according to the processing flow of steps S02 to S05, the collected data is preprocessed, feature extracted, dimension reduced, data fused and feature weighted to obtain weighted fusion data of the tobacco leaf to be tested. Then, the weighted fusion data is input into the optimized tobacco leaf state discrimination model, and the probability distribution of the three maturity grades is obtained by forward propagation calculation. Finally, the grade with the highest probability is selected as the discrimination result. When the highest probability exceeds 85%, the discrimination result has high credibility. When the highest probability is less than 65%, re-collection of data or manual intervention is required. The purpose of this step is to apply the constructed tobacco leaf state discrimination model to perform maturity grade discrimination on the tobacco leaf to be tested, providing accurate guidance for the tobacco leaf baking process.

[0064] like Figure 2As shown in the figure, the detailed structure of the tobacco leaf state discrimination model adopts an improved deep convolutional neural network design and consists of three parts: a feature extraction module, a feature fusion module, and a classification module. The feature extraction module uses a two-stream structure to process the principal component features of the tobacco leaf near-infrared spectrum and the principal component features of the tobacco leaf image, respectively. Each stream contains two one-dimensional convolutional layers with a convolution kernel size of 3, a stride of 1, and a padding of 1. Each layer is followed by a batch normalization layer and a ReLU activation function, as well as a maximum pooling layer with a pooling kernel size of 2 and a stride of 2. The feature fusion module uses an attention mechanism to achieve adaptive fusion of the two features. It includes two sub-modules: channel attention and spatial attention. Channel attention extracts channel descriptors through global average pooling and global maximum pooling, and generates channel weights through a shared multi-layer perceptron. Spatial attention extracts spatial feature maps through convolution operations and generates a spatial weight matrix. The classification module consists of two fully connected layers, with 1024 neurons in the first layer and 512 in the second. Each layer is followed by a dropout layer with a dropout rate of 0.5 to prevent overfitting. The final layer, the output layer, has 3 neurons corresponding to the three maturity levels and uses a softmax activation function to output a probability distribution. The training dataset was constructed by collecting 3,000 tobacco leaf samples from various varieties, including Honghua Dajinyuan, K326, and Yunyan 87. Near-infrared spectral and image data were first collected for each sample, with spectral data acquired at 851 wavelengths within the standard wavelength range. Five tobacco leaf grading experts with over 10 years of experience were then invited to grade the samples based on their appearance and internal quality. A consensus rating was reached using the Delphi method, resulting in three grades: immature, moderately mature, and overmature. The collected data was then preprocessed to remove anomalous samples. Outliers were identified using boxplots. Spectral data with a signal-to-noise ratio below 10 and image data with a clarity score below 0.5 were removed. Features were then extracted and normalized using the z-score method. Finally, the training and validation sets were randomly divided into a 7:3 ratio, ensuring a balanced distribution of samples across varieties and maturity levels, resulting in a training dataset consisting of 2,100 training samples and 900 validation samples.

[0065] The bimodal feature attention model adopts a multi-layer perceptron structure, such as Figure 3 As shown in Figure 1, it mainly consists of an input layer, multiple hidden layers, and an output layer. The hidden layer includes an interactive attention module, a feature extraction module, and a feature fusion module. The detailed structure of the model is as follows:

[0066] The input layer receives two modal feature data, namely the principal component eigenvectors of tobacco leaf near-infrared spectrum and and the principal component eigenvector of tobacco leaf image K is typically 8 to 12, and M is typically 5 to 8. These two parameters are determined based on the cumulative contribution rate of principal component analysis. The input layer normalizes the eigenvectors, using the min-max normalization method to map the eigenvalues ​​to a range between 0 and 1, ensuring that the different modal features are within the same numerical range.

[0067] The interactive attention module consists of two parts: the multi-head self-attention mechanism and the mutual attention mechanism. The multi-head self-attention mechanism processes the feature association within a single modality separately. For each modality feature X, it first passes three different linear transformations W Q 、W K 、W V Generate query matrix Q, key matrix K and value matrix V, all with dimension d k , usually set to 64. Then calculate the attention score, which is calculated in the form of dot product and scaled by the factor Normalize to prevent gradient disappearance. Then apply Softmax function to the attention score to obtain weight coefficient, and finally multiply the weight coefficient with the value matrix V to obtain weighted features. The multi-head attention mechanism uses 8 attention heads for parallel calculation, and the output dimension of each head is d k / 8, and finally, the outputs of the eight heads are concatenated and linearly transformed to obtain the final output. The mutual attention mechanism is used to learn the mapping relationship between cross-modal features. Features from one modality are used as queries, and features from the other modality are used as keys and values. The computation is similar to the self-attention mechanism, but the inputs come from different modalities. The mutual attention module also includes residual connections and layer normalization operations. Residual connections add the input and the attention output, and layer normalization standardizes the result to ensure training stability.

[0068] The feature extraction module, consisting of fully connected layers and activation functions, is used to further extract high-level representations of each modal feature. This module employs a two-layer fully connected network structure, with 256 neurons in the first layer and 128 neurons in the second layer. Each layer is followed by batch normalization and a Reluctant Unit (ReLU) activation function. To prevent overfitting, a Dropout layer is added between the fully connected layers, with a dropout rate set to 0.3. The weights of the fully connected layers are initialized using the He initialization method, with biases initialized to 0. The feature extraction module processes the outputs of self-attention and mutual-attention, generating four feature vectors: the self-attention representation of the tobacco leaf's near-infrared spectral features, the self-attention representation of the tobacco leaf's image features, the mutual-attention representation of the tobacco leaf's near-infrared spectral features with respect to the tobacco leaf's image features, and the mutual-attention representation of the tobacco leaf's image features with respect to the tobacco leaf's near-infrared spectral features.

[0069] The feature fusion module consists of spatial attention and channel attention mechanisms to integrate multimodal features. The spatial attention mechanism first concatenates the four feature vectors to form a feature map. It then extracts the spatial attention map through two convolutional layers. The first convolutional layer has a kernel size of 7×7, a stride of 1, a padding of 3, and an output channel of 32. The second convolutional layer has a kernel size of 5×5, a stride of 1, a padding of 2, and an output channel of 1. Finally, a sigmoid function is applied to map the attention map to a value between 0 and 1, which serves as a spatial weight coefficient. The channel attention mechanism extracts channel descriptors through global average pooling and global maximum pooling, respectively. The two pooling results are fed into a shared multilayer perceptron (MLP). The perceptron consists of two fully connected layers. The number of neurons in the first layer is 1 / 16 of the number of input channels, and the number of neurons in the second layer is equal to the number of input channels. The outputs of the two branches are summed and then applied to the sigmoid function to obtain the channel weight coefficient. The feature fusion module multiplies the outputs of the spatial and channel attention algorithms and then adds them to the original features to achieve adaptive feature fusion.

[0070] The output layer receives the fused features and generates the final weighted fused data and feature weight distribution through a fully connected layer. The output layer consists of a fully connected layer with K+M neurons, which generates weighted fused data with the same dimensions as the input, and a fully connected layer with 2 neurons, which outputs the weight distribution of tobacco leaf near-infrared spectral features and tobacco leaf image features. The weighted fused data uses a linear activation function to maintain the original distribution of features, and the feature weight distribution uses a softmax activation function to ensure that the sum of the weights of the two modalities is 1.

[0071] The total number of parameters in the bimodal feature attention model is approximately 500,000 to 1,000,000, and the model depth is 8 layers, including 1 input layer, 5 hidden layers, and 2 output layers. The model uses an end-to-end training approach, and the loss function consists of three parts: reconstruction loss, contrastive loss, and classification loss. The reconstruction loss uses mean squared error to measure the difference between the weighted fused data and the original features; the contrastive loss uses InfoNCE loss to enhance the mutual information between different modalities; and the classification loss uses cross-entropy loss to ensure good discriminative power of the weighted fused data. The weights of the three losses are set to 0.4, 0.3, and 0.3, respectively, and the total loss function is the weighted sum of the three.

[0072] The steps for establishing a training dataset for the bimodal feature attention model include sample collection, data acquisition, expert rating, data preprocessing, feature extraction, and dataset partitioning. The detailed steps are as follows:

[0073] The sample collection phase first selected representative tobacco varieties, including the three mainstream varieties of Honghua Dajinyuan, K326 and Yunyan 87, covering tobacco leaf samples from different climatic regions and cultivation conditions. A total of 3,000 tobacco leaf samples were collected, including 1,000 samples of Honghua Dajinyuan, 1,000 samples of K326 varieties, and 1,000 samples of Yunyan 87 varieties. The sample collection time spanned a full growing season, including tobacco leaves at different maturity stages, ensuring that the data set covered various maturity states. Sample collection followed the principle of random stratified sampling. Tobacco plants were randomly selected from different plots, and the 8th to 16th tobacco leaves were taken from each plant to ensure consistency in sample location. The collected tobacco leaf samples were immediately sealed and stored in a constant temperature and humidity environment, with the temperature controlled at 22±2°C and the relative humidity controlled at 65±5% to prevent changes in sample characteristics.

[0074] During the data collection phase, near-infrared spectral data and image data were collected simultaneously for each tobacco leaf sample. Near-infrared spectral data were collected using a near-infrared spectrometer with a wavelength range of 800 to 2500 nm and a resolution better than 2 nm. The sampling interval was set to 2 nm. Each sample was scanned from the center, covering a 20 mm diameter circular area. To reduce random error, each sample was scanned 32 times, and the average value was used as the final spectral data. Image data were collected using a digital camera with a resolution of 4000 × 3000 pixels, equipped with a standard D65 light source with a color temperature of 5500K and an illumination intensity of 800 Lux. The camera was held 30 cm vertically from the tobacco leaf surface to ensure consistent image size and resolution. Images of both the front and back sides of each sample were collected, and a standard color chart and scale were placed at the edge of the image for subsequent color correction and size calibration. During the data collection process, the ambient temperature was maintained at 22 ± 2°C and the relative humidity at 65 ± 5% to prevent environmental factors from affecting data collection.

[0075] During the expert rating phase, five experts with more than 10 years of experience in tobacco leaf grading were invited to grade the maturity of the collected tobacco leaf samples. The rating adopted the Delphi method, in which the experts first independently rated the samples, then summarized their opinions and discussed the divergent samples until a consensus was reached. The rating criteria were based on five indicators: tobacco leaf color, aroma, oil content, tissue structure, and elasticity. Each indicator was scored on a scale of 1 to 5, with a total score between 5 and 25 points. Based on the total score, the maturity of the tobacco leaves was divided into three levels: immature (5 to 12 points), moderately mature (13 to 20 points), and overmature (21 to 25 points). The rating process was carried out under standard light conditions to ensure the consistency of visual assessment. In order to test the reliability of the rating, 10% of the samples were randomly selected for repeated rating, and the rating consistency was calculated. The consistency requirement was to reach more than 90%. After expert rating, the final label distribution was: 1,000 immature samples, 1,200 moderately mature samples, and 800 overmature samples.

[0076] The data preprocessing phase first involved outlier detection on the near-infrared spectral data. The Mahalanobis distance method was used to identify multivariate outliers, with a threshold of 3.0. Samples exceeding this threshold were labeled as outliers. The spectral signal-to-noise ratio (SNR) was then calculated by dividing the signal intensity of the main absorption peak by the standard deviation of the non-absorbing region. The SNR threshold was set to 10, and samples below this value were eliminated. The spectral data were then baseline corrected, smoothed, and normalized. Baseline correction was performed using polynomial fitting, smoothing using the Savitzky-Golay algorithm, and normalization using the standard normal variate transformation. Image data were first color corrected, with white balance adjusted using a standard color chart in the image to ensure color consistency. Image segmentation was then performed, using the Otsu thresholding method to separate the tobacco leaves from the background and extract the leaf regions. Image sharpness was then calculated using a Laplace transform, with a threshold of 0.5 set, and samples below this threshold were eliminated. Finally, image standardization was performed, with all images resized to a uniform resolution of 2000 × 1500 pixels and brightness normalized. Through data preprocessing, about 150 abnormal samples were eliminated, leaving 2,850 valid samples.

[0077] The feature extraction phase first performs principal component analysis on the preprocessed near-infrared spectral data to extract the principal component features of the spectrum. The covariance matrix is ​​calculated, and the eigenvalues ​​and eigenvectors are solved. The top 10 principal components with a cumulative contribution rate of 95% are selected as the principal component features of the tobacco leaf near-infrared spectrum. HSV color space features are then extracted from the image data. The mean, standard deviation, skewness, and kurtosis of the hue, saturation, and lightness channels are calculated to form a 12-dimensional chromaticity feature vector. Morphological parameters, including geometric parameters such as leaf area, perimeter, aspect ratio, roundness, and rectangularity, as well as vein distribution, are then extracted to form an 8-dimensional morphological feature vector. The chromaticity and morphological feature vectors are merged into a 20-dimensional feature vector, which is then reduced in dimension using principal component analysis. The top six principal components with a cumulative contribution rate of 90% are selected as the principal component features of the tobacco leaf image. Finally, the two principal component features are normalized to ensure consistent numerical ranges for subsequent fusion processing. After feature extraction, the 10-dimensional principal component features of tobacco leaf near-infrared spectrum and the 6-dimensional principal component features of tobacco leaf image are obtained to form the feature vector of each sample.

[0078] Stratified random sampling was used to partition the dataset into training and validation sets in a ratio of 7 to 3. To ensure a balanced distribution of samples across varieties and maturity levels, the samples were first stratified by tobacco variety and maturity level, ensuring that samples within each subgroup were equally distributed. The Honghua Dajinyuan variety was divided into a training set of 665 samples and a validation set of 285; the K326 variety was divided into a training set of 672 samples and a validation set of 288; and the Yunyan 87 variety was divided into a training set of 658 samples and a validation set of 282. Furthermore, immature samples were divided into a training set of 686 samples and a validation set of 294; moderately mature samples were divided into a training set of 826 samples and a validation set of 354; and over-mature samples were divided into a training set of 483 samples and a validation set of 207. A fixed random seed was used during the partitioning process to ensure reproducible results. The training set was used for model parameter learning, and the validation set was used for hyperparameter tuning and model performance evaluation. Finally, a paired training sample set containing the principal component features of tobacco leaf near-infrared spectra, principal component features of tobacco leaf images and label information is constructed. The data format is (X NIR , X IMG , y), where X NIR is the principal component eigenvector of the 10-dimensional tobacco leaf near-infrared spectrum, X IMG is the principal component eigenvector of the 6-dimensional tobacco leaf image, and y is the maturity grade label.

[0079] The bimodal feature attention model was trained using a batch gradient descent algorithm with 200 epochs, a learning rate of 0.001, and a batch size of 32. Parameter optimization was performed using the Adam optimizer, with initial momentum parameters β1 set to 0.9, β2 set to 0.999, and a weight decay coefficient of 0.0001. The loss function was a weighted combination of cross-entropy loss and contrastive learning loss, with weights of 0.7 and 0.3, respectively. An early stopping mechanism was introduced during training to prevent overfitting. Validation was performed every five epochs, and training was terminated after ten consecutive validation cycles without improvement. A learning rate decay strategy was also implemented, reducing the learning rate by a factor of 0.8 every 50 epochs. To enhance model generalization, data augmentation techniques were applied during training, including Gaussian noise, random feature masking, and eigenvalue perturbation. The best-performing model on the validation set was selected as the pretrained model for the subsequent tobacco leaf state discrimination task.

[0080] It should be noted that the core technical ideas of the present invention mainly include two aspects: bimodal feature fusion and bimodal feature attention model.

[0081] The dual-modal feature fusion technology approach achieves a comprehensive characterization of the internal chemical composition and external color and shape characteristics of tobacco leaves by simultaneously collecting near-infrared spectral data and tobacco leaf image data. Unlike traditional single-modal identification methods, this approach overcomes the limitations of a single data source, enabling the system to fully perceive the multi-dimensional characteristics of tobacco leaves. Near-infrared spectral data can reflect the absorption characteristics of groups such as NH bonds, CH bonds, and OH bonds within tobacco leaves, directly related to changes in the internal chemical composition of tobacco leaves; while tobacco leaf image data contains visual feature information such as color, texture, and morphology, reflecting the appearance of the tobacco leaves. Through principal component analysis, the two types of data are reduced in dimension and then fused to construct a new feature set that can comprehensively characterize the internal and external characteristics of tobacco leaves. This provides a more comprehensive and reliable information basis for distinguishing tobacco leaf maturity, and has the technical advantages of richer information dimensions and more comprehensive feature expression compared to traditional single-modal methods.

[0082] The technical concept of the dual-modal feature attention model adopts a multi-layer perceptron structure. Through the synergy of interactive attention modules, feature extraction modules, and feature fusion modules, it achieves adaptive fusion and weight optimization of different modal features. The model extracts internal feature associations of a single modality through a multi-head self-attention mechanism, learns cross-modal feature mapping relationships through a mutual attention mechanism, and finally achieves deep fusion of multimodal features through a feature fusion module. Compared with traditional fusion methods based on simple feature splicing or fixed-weight weighted averaging, this attention-based fusion method can adaptively learn the importance of different modal features, achieve deep interaction and selective fusion at the feature level, and is more flexible and effective. Especially under conditions of diverse tobacco varieties and complex growing environments, the model can intelligently adjust the attention paid to different modal features, improving the generalization ability and adaptability of the discriminant model.

[0083] The synergistic effect of the above two core technical ideas has enabled the present invention to achieve a technological breakthrough in the field of tobacco leaf maturity discrimination. Bimodal feature fusion provides a rich source of information, while the bimodal feature attention model can intelligently mine and integrate this information. This synergistic effect enables the system to adaptively adjust the fusion strategy under different data quality conditions to achieve accurate discrimination of tobacco leaf maturity. Especially in the actual baking environment, due to unstable acquisition conditions, the quality of different modal data fluctuates. The present invention dynamically optimizes the model parameters through the fusion weight adjustment function, so that the discrimination system maintains high precision and high stability. Compared with the traditional fixed-mode discrimination method, this intelligent and adaptive discrimination method has stronger environmental adaptability and discrimination reliability, and provides reliable technical support for the precise baking of tobacco leaves.

[0084] Specifically, the core principle of this invention lies in accurately distinguishing tobacco leaf maturity through bimodal data fusion and dynamic weight adjustment. This method leverages the complementary strengths of near-infrared spectroscopy and image analysis technologies to construct an adaptive fusion framework that dynamically optimizes feature fusion strategies based on data quality, thereby improving the accuracy and robustness of tobacco leaf maturity determination.

[0085] At the feature extraction level, this method simultaneously collects tobacco leaf near-infrared spectral data and tobacco leaf image data to characterize the internal chemical composition and external color and shape characteristics of the tobacco leaves, respectively. Near-infrared spectroscopy can detect the absorption characteristics of groups such as NH, CH, and OH bonds in tobacco leaves, reflecting changes in the leaves' internal composition; tobacco leaf images contain visual features such as color, texture, and morphology, reflecting the appearance of the leaves. By preprocessing these two types of data and performing principal component analysis to reduce dimensionality, key features of tobacco leaf characteristics are extracted, reducing data redundancy and improving feature expression efficiency.

[0086] At the feature fusion level, the present invention designs a dual-modal feature attention model, which adopts a multi-layer perceptron structure and includes an interactive attention module, a feature extraction module, and a feature fusion module. This model extracts the internal feature associations of a single modality through a multi-head self-attention mechanism, learns the cross-modal feature mapping relationship through a mutual attention mechanism, and finally realizes the deep fusion and weight optimization of multimodal features through a feature fusion module. This fusion method based on the attention mechanism can adaptively learn the importance of different modal features, realize deep interaction and selective fusion at the feature level, and is more flexible and effective than traditional fusion methods such as simple feature splicing or weighted averaging.

[0087] At the dynamic adjustment level, the present invention innovatively proposes a fusion weight adjustment function, which dynamically adjusts the weight distribution of the interactive attention module in the dual-modal feature attention model based on three parameters: the signal-to-noise ratio of the tobacco leaf near-infrared spectral data, the clarity score of the tobacco leaf image data, and the data consistency evaluation value. When the balance value is in different intervals, the logarithmic, linear, or exponential weight adjustment function is used to dynamically balance the weights of the tobacco leaf near-infrared spectral features and the tobacco leaf image features. This mechanism can effectively deal with the problem of quality fluctuations in different modal data in actual baking environments. When the quality of a certain modal data decreases, its weight in the fusion process is automatically reduced, thereby maintaining the stability and accuracy of the discriminant model.

[0088] Through the synergistic effect of these three levels, the present invention achieves accurate identification of tobacco leaf maturity, effectively resolving the technical issue of insufficient accuracy in identifying tobacco leaf maturity using multimodal information fusion. This technical solution is theoretically sound and practically feasible, providing technical support for precise tobacco leaf curing.

[0089] A specific embodiment 1 of the present invention is provided below. The specific implementation of each step in this embodiment 1 is described in detail as follows.

[0090] The specific implementation of step S01 is to collect data from tobacco leaf samples using a near-infrared spectrometer and a high-resolution digital camera, and establish a tobacco leaf maturity grading standard. The specific implementation process is to first use a near-infrared spectrometer with a wavelength range of 800 to 2500 nm and a sampling interval of 2 nm to scan the tobacco leaf samples. 32 spectra are collected for each sample and the average is taken to eliminate random errors to obtain tobacco leaf near-infrared spectral data. The average value during the spectral data collection process is calculated using the formula:

[0091]

[0092] Where R λ is the average spectral reflectance at wavelength λ; R λ,i is the spectral reflectance at wavelength λ obtained by the i-th measurement; N is the total number of measurements, which is 32 here.

[0093] At the same time, a high-resolution digital camera with a resolution of no less than 3000×2000 pixels was used to capture color images of tobacco leaves at a vertical distance of 30 cm under standard D65 light conditions with an illumination intensity of 800 Lux and a color temperature of 5500K. Images of both the front and back sides of each sample were captured. By hiring more than five experts with more than 10 years of tobacco leaf grading experience, a three-level grading standard for tobacco leaf maturity was established based on indicators such as tobacco leaf color, aroma, oil content, tissue structure, and elasticity, namely immature, moderately mature, and overmature. The expert score was calculated using the weighted average method:

[0094]

[0095] Where S mat is the comprehensive score of tobacco leaf maturity; j is the weight of the j-th evaluation index; s j,k is the score of the kth expert on the jth indicator; M is the number of evaluation indicators, which is 5 here; K is the number of experts, which is not less than 5. mat The values ​​are divided into three levels: immature (S mat <0.4), moderately mature (0.4≤S mat ≤0.7), overmature (S mat The purpose of this step is to obtain multimodal data that can comprehensively characterize tobacco leaf characteristics and provide basic data support for subsequent analysis.

[0096] The specific implementation of step S02 is to pre-process the tobacco leaf near-infrared spectral data to obtain pre-processed tobacco leaf near-infrared spectral data and extract the characteristic information of the internal chemical components of the tobacco leaf. The specific implementation process is to first perform a standard normal variable transformation on the original spectral data to eliminate the influence of baseline drift. The transformation formula is:

[0097]

[0098] Where, X SNV,i is the spectrum value of the i-th wavelength point after standard normal variable transformation; X i is the value of the i-th wavelength point in the original spectrum data; is the average value of the spectrum values ​​at all wavelength points; σ X is the standard deviation of the spectral values ​​at all wavelengths.

[0099] The Savitzky-Golay smoothing algorithm is then applied to smooth the spectral data. The window width is set to 15 wavelength points and the polynomial order is 3 to reduce spectral noise. The calculation formula of the Savitzky-Golay smoothing algorithm is:

[0100]

[0101] Where, X SG,i is the spectrum value of the i-th wavelength point after smoothing; X i+j is the value of the i+jth wavelength point in the original spectrum; c j is the convolution coefficient, which is determined by the window width and the polynomial order; m is the half window width, which is 7 here, corresponding to a window width of 15.

[0102] Then, a multivariate scattering correction algorithm is used to eliminate the scattering effect caused by the uneven size of the sample particles. The correction parameters are selected to obtain the optimal value obtained by iterative optimization. The calculation formula of the multivariate scattering correction algorithm is:

[0103]

[0104] Where, X MSC,i is the spectrum value of the i-th wavelength point after multivariate scattering correction; X SG,i is the spectrum value of the i-th wavelength point after smoothing; a and b are the intercept and slope of the linear regression, respectively, which are obtained by regressing the sample spectrum against the standard spectrum.

[0105] Then, the first-order derivative transformation is applied to enhance the spectral features, and the first-order derivative of the spectrum is calculated using the difference method with a differential interval of 5 nm. The calculation formula is:

[0106]

[0107] Where, X D1,i is the first derivative value of the i-th wavelength point; XMSC,i+Δ and X MSC,i-Δ are the multivariate scattering corrected spectral values ​​at Δ wavelength points in the backward and forward directions, respectively; Δλ is the wavelength interval, which is 5 nm here.

[0108] Finally, a range selection algorithm was used to identify wavelength ranges with high correlation with tobacco leaf chemical composition. These ranges primarily included the 1100-1200 nm, 1400-1600 nm, and 1900-2100 nm ranges, corresponding to the absorption characteristics of groups such as NH, CH, and OH bonds in tobacco leaves. Range selection was performed using a stepwise regression analysis method, with a correlation threshold of 0.7. Wavelength ranges above this threshold were retained. This step aims to eliminate non-sample information interference, enhance the characteristic information of the tobacco leaf's internal chemical composition, and improve the accuracy of subsequent analysis.

[0109] The specific implementation of step S03 is to perform image processing on the tobacco leaf image data to extract tobacco leaf color parameters and tobacco leaf morphological characteristic parameters, and quantify the external characteristics of the tobacco leaves. The specific implementation process is to first pre-process the original image using a Gaussian filter algorithm with a filter kernel size of 5×5 and a standard deviation of 1.5 to remove image noise. The calculation formula for the two-dimensional convolution kernel of the Gaussian filter is:

[0110]

[0111] Where G(x, y) is the value of the Gaussian kernel at position (x, y); σ is the standard deviation of the Gaussian function, which is 1.5 here; and x and y are the pixel coordinates relative to the center of the kernel.

[0112] Filtered image I f Obtained through convolution operation:

[0113]

[0114] Where, I f (x, y) is the pixel value of the filtered image at position (x, y); I(xs, yt) is the pixel value of the original image at position (xs, yt); k is the radius of the convolution kernel, which is 2 here, corresponding to a 5×5 convolution kernel.

[0115] Then, the Otsu threshold segmentation algorithm is applied to extract the outline of the tobacco leaf. The segmentation threshold is automatically calculated, generally between grayscale values ​​45 and 65, to separate the tobacco leaf from the background. The Otsu threshold method determines the optimal threshold by maximizing the inter-class variance. The calculation formula is:

[0116]

[0117] Where, is the inter-class variance under threshold t; ω0(t) and ω1(t) are the probabilities of foreground and background pixels respectively;

[0118] μ0(t) and μ1(t) are the average grayscale values ​​of foreground and background pixels respectively. The optimal threshold t * That is to make Maximum threshold:

[0119]

[0120] Then, the RGB color space is converted to the HSV color space, and the values ​​of the three channels of hue, saturation, and lightness are extracted. The mean, standard deviation, skewness, and kurtosis of each channel in the tobacco leaf area are calculated to form a 12-dimensional chromaticity feature vector. The conversion formula from RGB to HSV is:

[0121] V = max(R, G, B);

[0122]

[0123] Where R, G, and B are the values ​​of the red, green, and blue channels, respectively, ranging from 0 to 1; H is the hue value, ranging from 0 to 360 degrees; S is the saturation value, ranging from 0 to 1; and V is the lightness value, ranging from 0 to 1.

[0124] The calculation formula of the chromaticity eigenvector C is:

[0125] C=[μ H ,σ H , s H , k H , μ S ,σ S , s S , k S , μ V ,σ V , s V , k V ];

[0126] Where μ, σ, s, and k represent the mean, standard deviation, skewness, and kurtosis, respectively; the subscripts H, S, and V represent the hue, saturation, and value channels, respectively.

[0127] Then, based on the segmented tobacco leaf contour, the geometric morphological features of the tobacco leaf, such as area, perimeter, aspect ratio, roundness, and rectangularity, are calculated. The roundness threshold is set to 0.5. A value lower than this indicates that the tobacco leaf shape is irregular. The calculation formula for tobacco leaf morphological features is:

[0128]

[0129] Where A is the tobacco leaf area, and the unit is the number of pixels; p i is the i-th pixel belonging to the tobacco leaf area; n is the total number of pixels in the tobacco leaf area; P is the perimeter of the tobacco leaf, in pixels; li is the length of the i-th boundary pixel segment; m is the total number of boundary pixel segments; AR is the aspect ratio; L and W are the length and width of the tobacco leaf respectively; C is the roundness; and R is the rectangularity.

[0130] Finally, the Gobota algorithm is used to extract the vein distribution characteristics of tobacco leaves. Parameters such as the number of main veins, the number of lateral veins, the vein distribution density, and the vein direction angle are extracted to form an 8-dimensional morphological feature vector. The key steps in vein extraction are image enhancement and binarization. The enhancement formula is:

[0131] I e =I g -I gm ;

[0132] Where, I e is the enhanced image; I g is a grayscale image; I gm The grayscale image after applying the mean filter. The purpose of this step is to quantitatively characterize the external color and shape characteristics of the tobacco leaves and obtain characteristic parameters that can reflect the appearance quality of the tobacco leaves.

[0133] The specific implementation of step S04 is to use the principal component analysis method to perform dimensionality reduction processing on the pre-processed tobacco leaf near-infrared spectrum data, tobacco leaf color parameters, and tobacco leaf morphological characteristic parameters, respectively, to obtain the principal component characteristics of the tobacco leaf near-infrared spectrum and the principal component characteristics of the tobacco leaf image. The specific implementation process is to first perform principal component analysis on the pre-processed tobacco leaf near-infrared spectrum data, calculate the eigenvalues ​​and eigenvectors, sort them in descending order of eigenvalues, and select the top K principal components with a cumulative contribution rate of 95%, usually with a K value of 8 to 12, as the principal component characteristics of the tobacco leaf near-infrared spectrum. The calculation steps of the principal component analysis are as follows:

[0134] Calculate the covariance matrix ∑:

[0135]

[0136] Where ∑ is the covariance matrix; X i is the feature vector of the i-th sample; is the average value of all sample feature vectors; N is the total number of samples.

[0137] Solve the characteristic equation to obtain the eigenvalue λ j and the eigenvector v j :

[0138] ∑v j =λ j v j ;

[0139] Where λ j is the jth eigenvalue; v j is the corresponding eigenvector.

[0140] Arrange the eigenvalues ​​and eigenvectors in descending order of eigenvalue size, select the first K eigenvectors to form the projection matrix P, and the K value is determined by the cumulative contribution rate:

[0141]

[0142] Where d is the original feature dimension.

[0143] Calculate principal component scores:

[0144]

[0145] Where Y NIR is the principal component feature of tobacco leaf near-infrared spectrum; P is the projection matrix; X is the original eigenvector; is the average value of all sample feature vectors.

[0146] Then, the tobacco leaf color parameters and tobacco leaf morphological characteristic parameters are combined to form a 20-dimensional feature vector. The principal component analysis method is also used for dimensionality reduction. The top M principal components with a cumulative contribution rate of 90% are selected. Usually, the M value is 5 to 8, which is used as the principal component features of the tobacco leaf image. Finally, the obtained principal component features are standardized so that the data range of each principal component is unified to between 0 and 1, which facilitates subsequent data fusion. The standardization formula is:

[0147]

[0148] Where Y norm is the standardized principal component feature; Y is the original principal component feature; Y min and Y max are the minimum and maximum values ​​of the principal component features, respectively. The purpose of this step is to reduce data redundancy through dimensionality reduction, extract the most representative internal and external feature information of tobacco leaves, improve computational efficiency, and avoid the dimensionality curse problem.

[0149] The specific implementation of step S05 is to fuse the principal component features of the tobacco leaf near-infrared spectrum with the principal component features of the tobacco leaf image to obtain fused data, calculate data quality assessment indicators, and apply a bimodal feature attention model to weight the feature importance of the fused data. The specific implementation process is to first use a feature cascade method to connect the principal component features of the tobacco leaf near-infrared spectrum with the principal component features of the tobacco leaf image in the feature dimension to form a K+M dimensional fused feature vector. The calculation formula is:

[0150] F=[Y NIR ; Y IMG ];

[0151] Where F is the fusion feature vector; Y NIRis the principal component feature of tobacco leaf near infrared spectrum, with dimension K; Y IMG is the principal component feature of the tobacco leaf image, with a dimension of M; [;] represents the vector concatenation operation.

[0152] Then, the signal-to-noise ratio of tobacco leaf near-infrared spectral data was calculated by dividing the average signal intensity of the main spectral band by the standard deviation of the background noise. The signal-to-noise ratio threshold was set to 15. A value lower than this indicates poor data quality. The signal-to-noise ratio calculation formula is:

[0153]

[0154] Where, SNR NIR is the signal-to-noise ratio of tobacco leaf near-infrared spectroscopy data; μ signal is the average signal intensity of the main band of the spectrum; σ noise is the standard deviation of background noise.

[0155] Next, we calculate the clarity score of the tobacco leaf image data. We use the Laplace operator to calculate the image clarity. The clarity score threshold is set to 0.6. A value below this value indicates poor image quality. The clarity score calculation formula is:

[0156]

[0157] Where, CF IMG is the clarity score of tobacco leaf image data; is the Laplace operator value of the image at position (x, y); N pix is the total number of pixels in the image.

[0158] Then, the data consistency evaluation value is calculated. The normalized mutual information method is used to calculate the correlation between the two modal features. The consistency threshold is set to 0.4. A value lower than this indicates that the representations of the two data sources are inconsistent. The formula for calculating the data consistency evaluation value is:

[0159]

[0160] Where, DC is the data consistency evaluation value; I(Y NIR ; Y IMG ) is the mutual information between the principal component features of tobacco leaf near-infrared spectrum and the principal component features of tobacco leaf image; H(Y NIR ) and H(Y IMG ) are the entropies of the two features respectively.

[0161] Finally, the fused data and three evaluation metrics are fed into a pre-trained bimodal feature attention model. This model learns the correlation and importance between features from different modalities and adaptively adjusts feature weights to generate weighted fused data. This step aims to effectively fuse multi-source tobacco leaf data, dynamically adjusting the contribution ratio of different modal data based on data quality, and improving the expressiveness and discriminative performance of the fused features.

[0162] The specific implementation method of step S06 is to construct a tobacco leaf state discrimination model based on weighted fusion data, and use a fusion weight adjustment function to optimize the parameters of the tobacco leaf state discrimination model. The specific implementation process is to first use the weighted fusion data as input to construct a tobacco leaf state discrimination model based on a deep convolutional neural network. The network structure includes an input layer, four convolution layers, two fully connected layers and an output layer, wherein the convolution kernel size is 3×3, the number of convolution kernels in each layer is 32, 64, 128 and 256 respectively, the number of neurons in the fully connected layer is 1024 and 512 respectively, and the number of neurons in the output layer is 3, corresponding to three maturity levels. Then, the balance value based on the signal-to-noise ratio of the tobacco leaf near-infrared spectral data, the clarity score of the tobacco leaf image data and the data consistency evaluation value is calculated. The balance value calculation formula is:

[0163]

[0164] Where B is the balance value; SNR NIR ′、CF IMG ′ and DC′ are the normalized signal-to-noise ratio of tobacco leaf near-infrared spectral data, the clarity score of tobacco leaf image data, and the data consistency evaluation value, respectively; w1, w2, and w3 are the weights of each indicator, which are set to 0.4, 0.4, and 0.2, respectively.

[0165] Then select the corresponding weight adjustment function according to the balance value:

[0166] When B∈[0, 0.3], the logarithmic weight adjustment function is used:

[0167] α NIR =a1-b1log(1-B);

[0168] α IMG =1-α NIR ;

[0169] When B∈(0.3, 0.7], a linear weight adjustment function is used:

[0170] α NIR =a2-b2B;

[0171] α IMG =1-α NIR ;

[0172] When B∈(0.7, 1], the exponential weight adjustment function is used:

[0173]

[0174] α IMG =1-α NIR ;

[0175] Where, α NIR and α IMG are the weights of the principal component features of tobacco leaf near-infrared spectra and tobacco leaf images, respectively; a1, b1, a2, b2, a3 and b3 are parameters determined by optimization of the validation set.

[0176] Finally, the weight adjustment function is used to adjust the parameters of the interactive attention module in the bimodal feature attention model to optimize the tobacco leaf state discrimination model's dependence on data from different modalities. The goal of this step is to build a high-performance tobacco leaf state discrimination model and improve its adaptability and robustness by dynamically adjusting weights.

[0177] The specific implementation method of step S07 is to input the tobacco leaf data to be tested into the tobacco leaf state discrimination model, realize the color and shape state discrimination of the tobacco leaf, and output the tobacco leaf maturity grade discrimination result. The specific implementation process is to first collect near-infrared spectral data and image data of the tobacco leaf sample to be tested, using the same equipment and parameters as step S01. Then, according to the processing flow of steps S02 to S05, the collected data is preprocessed, feature extracted, dimension reduced, data fused and feature weighted to obtain the weighted fusion data of the tobacco leaf to be tested. The weighted fusion data is then input into the optimized tobacco leaf state discrimination model, and the probability distribution of the three maturity grades is obtained by forward propagation calculation. The probability calculation formula is:

[0178]

[0179] Where, P(y=j|F w ) is the probability that the tobacco leaf to be tested belongs to the jth maturity level; F w is the weighted fusion data; z j is the output value of the jth neuron in the output layer; j∈{1, 2, 3}, corresponding to the three levels of immaturity, moderate maturity and over-maturity respectively.

[0180] Finally, the level with the highest probability is selected as the discrimination result. The discrimination result determination formula is:

[0181] y * =argmax j P(y=j|F w );

[0182] Where y *This is the final judgment result. When the highest probability exceeds 85%, the judgment result is highly reliable. When the highest probability is less than 65%, re-data collection or manual intervention is required. The purpose of this step is to apply the constructed tobacco leaf state discrimination model to determine the maturity level of the tested tobacco leaves, providing accurate guidance for the tobacco leaf curing process.

[0183] The bimodal feature attention model adopts a multi-layer perceptron structure, which mainly consists of an input layer, multiple hidden layers, and an output layer. The hidden layer contains an interactive attention module, a feature extraction module, and a feature fusion module. The detailed structure of the model is as follows:

[0184] The input layer receives two modal feature data, namely the principal component eigenvectors of tobacco leaf near-infrared spectrum and and the principal component eigenvector of tobacco leaf image K is typically 8 to 12, and M is typically 5 to 8. These two parameters are determined based on the cumulative contribution rate of principal component analysis. The input layer normalizes the eigenvectors, using the min-max normalization method to map the eigenvalues ​​to a range between 0 and 1, ensuring that the different modal features are within the same numerical range.

[0185] The interactive attention module includes a multi-head self-attention mechanism and a mutual attention mechanism. The multi-head self-attention mechanism processes the feature correlation within a single modality separately, and its calculation formula is:

[0186] Q=XW Q , K=XW K , V=XW V ;

[0187]

[0188] Z = AV;

[0189] Where X is the input feature; W Q 、W K and W V is the linear transformation matrix; Q, K and V are the query matrix, key matrix and value matrix respectively; d k is the feature dimension, here 64; A is the attention weight matrix; Z is the weighted feature output. The multi-head attention mechanism divides the input features into h heads, each head calculates attention independently, and then splices the results:

[0190] Z=Concat(Z1,Z2,...,Z h )W O ;

[0191] Where Z i is the output of the i-th attention head; W Ois the output linear transformation matrix; h is the number of attention heads, which is 8 here. The mutual attention mechanism uses the features of one modality as the query and the features of the other modality as the key and value. The calculation formula is the same as that of the self-attention.

[0192] The feature extraction module uses a fully connected layer to extract high-level feature representations, and its calculation formula is:

[0193] H1=ReLU(XW1+b1);

[0194] H2=ReLU(H1W2+b2);

[0195] Where X is the input feature; W1 and W2 are weight matrices; b1 and b2 are bias vectors; H1 and H2 are hidden layer outputs; ReLU is the activation function, defined as ReLU(x) = max(0, x).

[0196] The feature fusion module combines spatial attention and channel attention to achieve effective fusion of multimodal features. The spatial attention calculation formula is:

[0197]

[0198] F max =max i∈{1,2,...,C} F i ;

[0199] F spa =σ(f 7×7 ([F avg ; F max ]));

[0200] Where, F i is the i-th channel of the feature map; C is the number of channels of the feature map; F avg and F max are the average pooling and maximum pooling results in the channel dimension respectively; [;] represents the splicing operation in the channel dimension; f 7×7 represents a convolution operation with a convolution kernel size of 7×7; σ is the Sigmoid activation function, defined as F spa is the spatial attention feature map.

[0201] The channel attention calculation formula is:

[0202]

[0203] z max =max i,j F(i, j);

[0204] F cha =σ(W2ReLU(W1z avg)+W2ReLU(W1z max ));

[0205] Where F(i, j) is the value of the feature map at position (i, j); H and W are the height and width of the feature map respectively; z avg and z max are the average pooling and maximum pooling results in the spatial dimension respectively; W1 and W2 are the weight matrices of the fully connected layer; F cha is the channel attention vector. The final fusion feature is calculated as:

[0206] F fused =F⊙F cha ⊙F spa +F;

[0207] Where ⊙ represents element-by-element multiplication; F fused is the fused feature.

[0208] The output layer generates weighted fusion data and feature weight distribution, which is calculated as follows:

[0209] F w =W3F fused +b3;

[0210] α=Softmax(W4F fused +b4);

[0211] Where W3 and W4 are weight matrices; b3 and b4 are bias vectors; F w is the weighted fusion data; α is the feature weight distribution, including α NIR and α IMG Two components, representing the weights of the two modal features.

[0212] The loss function of the bimodal feature attention model consists of three parts: reconstruction loss, contrast loss, and classification loss. The calculation formula is:

[0213] L=λ1L rec +λ2L con +λ3L cls ;

[0214] Where, L is the total loss; L rec is the reconstruction loss, using mean square error; L con For the contrast loss, InfoNCE loss is used; L cls is the classification loss, and the cross entropy loss is adopted; λ1, λ2 and λ3 are weight coefficients, which are set to 0.4, 0.3 and 0.3 respectively.

[0215] The reconstruction loss is calculated as:

[0216]

[0217] Where, is the weighted fusion data of the i-th sample; F (i) is the original fusion feature of the i-th sample; N is the number of samples; ||·||2 represents the Euclidean norm.

[0218] The contrast loss calculation formula is:

[0219]

[0220] Where, and are the near-infrared spectral characteristics and image characteristics of tobacco leaves of the i-th sample respectively; τ is the temperature parameter, which is set to 0.5; · represents the inner product operation.

[0221] The classification loss calculation formula is:

[0222]

[0223] Where, is the one-hot encoding of the true label of the i-th sample; The probability that the i-th sample belongs to the j-th category is predicted by the model.

[0224] The fusion weight adjustment function is calculated based on three parameters: the signal-to-noise ratio of tobacco leaf near-infrared spectral data, the clarity score of tobacco leaf image data, and the data consistency evaluation value, to obtain a balance value between 0 and 1. The calculation formula is as described above. When the balance value is in different ranges, different weight adjustment functions are used, where:

[0225] Logarithmic weight adjustment function (balanced value in the range of 0 to 0.3):

[0226] α NIR =0.8-0.2log(1-B);

[0227] α IMG =1-α NIR ;

[0228] In the formula, when B is close to 0, α NIR Close to 0.8, it means that the near-infrared spectral feature weight is dominant; when B is close to 0.3, α NIR It is about 0.65, which still maintains a high weight.

[0229] Linear weight adjustment function (balanced value in the range of 0.3 to 0.7):

[0230] α NIR =0.85-0.7B;

[0231] αIMG =1-α NIR ;

[0232] In the formula, when B increases from 0.3 to 0.7, α NIR Linearly decreasing, α IMG Linear increase to achieve a smooth transition of the weights of the two modal features.

[0233] Exponential weight adjustment function (balanced value in the range of 0.7 to 1):

[0234] α NIR =0.6e -2B ;

[0235] α IMG =1-α NIR ;

[0236] In the formula, when B is close to 0.7, α NIR is about 0.25; when B is close to 1, α NIR It decreases rapidly to close to 0, indicating that the image feature weight is dominant.

[0237] The fusion weight adjustment function adjusts the parameters of the interactive attention module in the dual-modal feature attention model. Specifically, it adjusts the weight distribution of each attention head in the multi-head attention mechanism. By dynamically adjusting the attention score matrix, feature selection and information fusion of different modal data are achieved. The adjustment method is to modify the scaling factor in the attention score calculation formula:

[0238]

[0239] Where A NIR and A IMG are the attention weight matrices corresponding to the tobacco leaf near-infrared spectral features and tobacco leaf image features respectively; α NIR and α IMG is the weight coefficient of the two modal features.

[0240] The steps for establishing a training dataset for the bimodal feature attention model include sample collection, data acquisition, expert rating, data preprocessing, feature extraction, and dataset partitioning. The detailed steps are as follows:

[0241] The sample collection phase first selected representative tobacco varieties, including the three mainstream varieties of Honghua Dajinyuan, K326 and Yunyan 87, covering tobacco leaf samples from different climatic regions and cultivation conditions. A total of 3,000 tobacco leaf samples were collected, including 1,000 samples of Honghua Dajinyuan, 1,000 samples of K326 varieties, and 1,000 samples of Yunyan 87 varieties. The sample collection time spanned a full growing season, including tobacco leaves at different maturity stages, ensuring that the data set covered various maturity states. Sample collection followed the principle of random stratified sampling. Tobacco plants were randomly selected from different plots, and the 8th to 16th tobacco leaves were taken from each plant to ensure consistency in sample location. The collected tobacco leaf samples were immediately sealed and stored in a constant temperature and humidity environment, with the temperature controlled at 22±2°C and the relative humidity controlled at 65±5% to prevent changes in sample characteristics.

[0242] During the data collection phase, near-infrared spectral data and image data were collected simultaneously for each tobacco leaf sample. Near-infrared spectral data were collected using a near-infrared spectrometer with a wavelength range of 800 to 2500 nm and a resolution better than 2 nm. The sampling interval was set to 2 nm. Each sample was scanned from the center, covering a 20 mm diameter circular area. To reduce random error, each sample was scanned 32 times, and the average value was used as the final spectral data. Image data were collected using a digital camera with a resolution of 4000 × 3000 pixels, equipped with a standard D65 light source with a color temperature of 5500K and an illumination intensity of 800 Lux. The camera was held 30 cm vertically from the tobacco leaf surface to ensure consistent image size and resolution. Images of both the front and back sides of each sample were collected, and a standard color chart and scale were placed at the edge of the image for subsequent color correction and size calibration. During the data collection process, the ambient temperature was maintained at 22 ± 2°C and the relative humidity at 65 ± 5% to prevent environmental factors from affecting data collection.

[0243] During the expert rating phase, five experts with more than 10 years of experience in tobacco leaf grading were invited to grade the maturity of the collected tobacco leaf samples. The rating adopted the Delphi method, in which the experts first independently rated the samples, then summarized their opinions and discussed the divergent samples until a consensus was reached. The rating criteria were based on five indicators: tobacco leaf color, aroma, oil content, tissue structure, and elasticity. Each indicator was scored on a scale of 1 to 5, with a total score between 5 and 25 points. Based on the total score, the maturity of the tobacco leaves was divided into three levels: immature (5 to 12 points), moderately mature (13 to 20 points), and overmature (21 to 25 points). The rating process was carried out under standard light conditions to ensure the consistency of visual assessment. In order to test the reliability of the rating, 10% of the samples were randomly selected for repeated rating, and the rating consistency was calculated. The consistency requirement was to reach more than 90%. After expert rating, the final label distribution was: 1,000 immature samples, 1,200 moderately mature samples, and 800 overmature samples.

[0244] In the data preprocessing phase, outlier detection is first performed on the near-infrared spectral data. The Mahalanobis distance method is used to identify multivariate outliers. The threshold is set to 3.0, and samples exceeding this threshold are marked as abnormal samples. The Mahalanobis distance calculation formula is:

[0245]

[0246] Where D M (x) is the Mahalanobis distance of sample x; μ is the mean vector of all samples; ∑ is the covariance matrix of the samples.

[0247] The spectral signal-to-noise ratio (SNR) was then calculated by dividing the signal intensity of the main absorption peak by the standard deviation of the non-absorbing region. The SNR threshold was set at 10, and samples below this threshold were eliminated. The spectral data were then baseline corrected, smoothed, and normalized. Baseline correction was performed using a polynomial fitting method, smoothing was performed using the Savitzky-Golay algorithm, and normalization was performed using a standard normal variate transformation. Image data were first color corrected, with white balance adjusted using a standard color chart within the image to ensure color consistency. Image segmentation was then performed, using the Otsu thresholding method to separate the tobacco leaves from the background and extract the tobacco leaf regions. Image sharpness was then calculated using a Laplace transform to assess image sharpness. The sharpness score threshold was set at 0.5, and samples below this threshold were eliminated. Finally, image standardization was performed, with all images resized to a uniform resolution of 2000 × 1500 pixels and brightness normalized. Data preprocessing eliminated approximately 150 outliers, leaving 2850 valid samples.

[0248] The feature extraction phase first performs principal component analysis on the preprocessed near-infrared spectral data to extract the principal component features of the spectrum. The covariance matrix is ​​calculated, and the eigenvalues ​​and eigenvectors are solved. The top 10 principal components with a cumulative contribution rate of 95% are selected as the principal component features of the tobacco leaf near-infrared spectrum. HSV color space features are then extracted from the image data. The mean, standard deviation, skewness, and kurtosis of the hue, saturation, and lightness channels are calculated to form a 12-dimensional chromaticity feature vector. Morphological parameters, including geometric parameters such as leaf area, perimeter, aspect ratio, roundness, and rectangularity, as well as vein distribution, are then extracted to form an 8-dimensional morphological feature vector. The chromaticity and morphological feature vectors are merged into a 20-dimensional feature vector, which is then reduced in dimension using principal component analysis. The top six principal components with a cumulative contribution rate of 90% are selected as the principal component features of the tobacco leaf image. Finally, the two principal component features are normalized to ensure consistent numerical ranges for subsequent fusion processing. After feature extraction, the 10-dimensional principal component features of tobacco leaf near-infrared spectrum and the 6-dimensional principal component features of tobacco leaf image are obtained to form the feature vector of each sample.

[0249] Stratified random sampling was used to partition the dataset into training and validation sets in a ratio of 7 to 3. To ensure a balanced distribution of samples across varieties and maturity levels, the samples were first stratified by tobacco variety and maturity level, ensuring that samples within each subgroup were equally distributed. The Honghua Dajinyuan variety was divided into a training set of 665 samples and a validation set of 285; the K326 variety was divided into a training set of 672 samples and a validation set of 288; and the Yunyan 87 variety was divided into a training set of 658 samples and a validation set of 282. Furthermore, immature samples were divided into a training set of 686 samples and a validation set of 294; moderately mature samples were divided into a training set of 826 samples and a validation set of 354; and over-mature samples were divided into a training set of 483 samples and a validation set of 207. A fixed random seed was used during the partitioning process to ensure reproducible results. The training set was used for model parameter learning, and the validation set was used for hyperparameter tuning and model performance evaluation. Finally, a paired training sample set containing the principal component features of tobacco leaf near-infrared spectra, principal component features of tobacco leaf images and label information is constructed. The data format is (X NIR , X IMG , y), where X NIR is the principal component eigenvector of the 10-dimensional tobacco leaf near-infrared spectrum, X IMG is the principal component eigenvector of the 6-dimensional tobacco leaf image, and y is the maturity grade label.

[0250] The detailed structure of the tobacco leaf state discrimination model adopts an improved deep convolutional neural network design, consisting of three parts: feature extraction module, feature fusion module, and classification module. The feature extraction module uses a two-stream structure to process the principal component features of the tobacco leaf near-infrared spectrum and the principal component features of the tobacco leaf image respectively. Each stream contains two one-dimensional convolution layers. The convolution operation formula is:

[0251] H l =f(W l *H l-1 +b l );

[0252] Where H l is the output feature map of the lth layer; W l and b l are the convolution kernel and bias of the lth layer respectively; * represents the convolution operation; f is the activation function, which uses the ReLU function.

[0253] The feature fusion module uses the attention mechanism to achieve adaptive fusion of two features. It includes two submodules: channel attention and spatial attention. The calculation formula is as described above. The classification module includes two fully connected layers and a Softmax output layer. The calculation formula is:

[0254] z=W5(dropout(W6(dropout(H fused ))+b6))+b5;

[0255]

[0256] Where H fused is the fused feature; W5, W6, b5 and b6 are the weight matrix and bias vector; dropout represents the Dropout operation with a dropout rate of 0.5; z is the linear output of the output layer; is the predicted probability distribution.

[0257] The steps for training the bimodal feature attention model include setting the number of epochs to 200, the learning rate to 0.001, the batch size to 32, using the Adam optimizer for parameter optimization, and the loss function to adopt a weighted combination of cross-entropy loss and contrastive learning loss. An early stopping mechanism is introduced to prevent overfitting, and the verification frequency is set to once every 5 epochs. Training is stopped when the accuracy does not improve after 10 consecutive verifications. The optimal model parameters are saved based on the performance of the verification set. A learning rate decay strategy is adopted during training, and the learning rate is reduced to 0.8 times the original every 50 epochs. Finally, the model with the best performance on the verification set is selected as the pre-training model.

[0258] In summary, this method for intelligently discriminating tobacco leaf color and shape during the tobacco baking process achieves precise identification of tobacco leaf maturity through the fusion of near-infrared spectroscopy and image data, combined with deep learning technology. This method's innovation lies in the adaptive fusion of data from different sources using a bimodal feature attention model. This method dynamically adjusts the model's reliance on data from different modalities through a fusion weight adjustment function, improving the adaptability and robustness of the discrimination model and addressing the inaccuracy of traditional single-modality discrimination methods.

[0259] In order to better understand and implement the present invention, Example 2 of a specific application scenario of the present invention is provided below: Researchers conducted intelligent and precise discrimination research on tobacco leaves of three tobacco varieties: Honghua Dajinyuan, K326 and Yunyan 87. 300 tobacco leaf samples from different producing areas were selected, 100 samples of each variety. The tobacco leaves were collected under standard harvesting conditions, and the maturity level was discriminated using the method of the present invention. First, a near-infrared spectrometer with a wavelength range of 800 to 2500nm and a sampling interval of 2nm was used to scan and collect the samples. 32 spectral acquisitions were performed for each sample and the average was taken to obtain stable near-infrared spectral data. At the same time, a digital camera with a resolution of 4000×3000 pixels was used to capture images of the front and back of the tobacco leaves under standard light conditions (light intensity 800Lux, color temperature 5500K). The expert group rated the samples and scored them according to five indicators: color, aroma, oil content, tissue structure and elasticity. They were divided into three levels: immature, moderately mature and over-mature. The grade distribution of each variety is shown in Table 1:

[0260] Table 1 Maturity grade distribution of three tobacco varieties

[0261] Tobacco varieties Unripe (portions) Moderately mature (servings) Overripe (portions) Total (copies) Red Flower Gold Yuan 28 48 24 100 K326 32 45 23 100 Clouds and Smoke 87 35 42 23 100 total 95 135 70 300

[0262] The collected near-infrared spectral data were preprocessed, including standard normal variate transformation, Savitzky-Golay smoothing (window width 15, polynomial order 3), multivariate scattering correction, and first-order derivative transformation. Three key wavelength ranges were selected using an interval selection algorithm: 1100 to 1200 nm, 1400 to 1600 nm, and 1900 to 2100 nm, corresponding to the absorption characteristics of groups such as NH bonds, CH bonds, and OH bonds in tobacco leaves. Table 2 shows the average absorbance values ​​of the three tobacco varieties in these key wavelength ranges:

[0263] Table 2 Average absorbance values ​​of three tobacco varieties in key wavelength ranges

[0264] Tobacco varieties 1100-1200nm 1400-1600nm 1900-2100nm Red Flower Gold Yuan 0.427 0.683 0.576 K326 0.415 0.652 0.581 Clouds and Smoke 87 0.432 0.671 0.564

[0265] Tobacco leaf images were processed by first using a Gaussian filter (kernel size 5×5, standard deviation 1.5) to remove noise. Then, the Otsu thresholding method (threshold range 45-65) was used for image segmentation. The RGB images were converted to the HSV color space to extract chromaticity parameters and calculate morphological feature parameters, including area, perimeter, aspect ratio, roundness, and rectangularity. Table 3 shows the average HSV chromaticity parameters of tobacco leaves at different maturity levels:

[0266] Table 3 Average HSV chromaticity parameters of tobacco leaves at different maturity levels

[0267] Maturity Level Hue mean (H) Saturation mean (S) Mean brightness (V) Immature 92.4 0.537 0.613 Moderately mature 73.6 0.642 0.578 Overmaturity 49.2 0.598 0.542

[0268] Principal component analysis (PCA) was used to reduce the dimensionality of the preprocessed near-infrared spectral data. The top 10 principal components with a cumulative contribution rate of 95% were selected as the principal component features of the tobacco leaf near-infrared spectra. PCA was also performed on the tobacco leaf image features, and the top 6 principal components with a cumulative contribution rate of 90% were selected as the principal component features of the tobacco leaf image. Data quality assessment indicators were calculated, including the tobacco leaf near-infrared spectral data signal-to-noise ratio (average value of 18.6), the tobacco leaf image data clarity score (average value of 0.73), and the data consistency assessment value (average value of 0.62).

[0269] A bimodal feature attention model was constructed, employing multi-head self-attention and mutual-attention mechanisms to achieve multimodal feature fusion. Feature importance was weighted using a corresponding weight adjustment function based on balance value calculation. The dataset was divided into a training set (210 samples) and a validation set (90 samples) in a 7:3 ratio to ensure a balanced distribution of samples across varieties and maturity levels. Model training parameters were set as follows: 200 iterations, a learning rate of 0.001, a batch size of 32, and the Adam optimizer was used. After model training and optimization, classification performance on the validation set is shown in Table 4:

[0270] Table 4 Comparison of classification performance of different discrimination methods on the validation set

[0271] Identification method Accuracy (%) Accuracy (%) Recall rate (%) F1 score Using only near-infrared spectroscopy 83.2 84.1 82.6 0.833 Using only image features 81.5 82.3 80.9 0.816 Simple feature fusion 86.4 87.2 85.8 0.865 Method of the present invention 93.7 94.2 93.1 0.936

[0272] Finally, 30 new tobacco leaf samples were selected for actual discrimination testing. Near-infrared spectroscopy and image data were collected, and the maturity level was discriminated using the established tobacco leaf state discrimination model, and compared with the expert rating results. The discrimination accuracy of the method of the present invention reached 93.3%, of which 28 sample discrimination results were consistent with the expert ratings, and only 2 samples were inconsistent. The inconsistent samples were all slight deviations from adjacent grades (immature was misjudged as moderately mature, or moderately mature was misjudged as overmature), and there were no serious misjudgments that spanned two grades.

[0273] Traditional tobacco leaf maturity discrimination mainly relies on expert experience for manual rating, which has problems such as strong subjectivity, poor consistency, and low efficiency. Or only single modal data (simple image or spectral data) is used for discrimination, and the accuracy rate is generally around 80%. The method of the present invention achieves accurate discrimination of tobacco leaf maturity through the dual-modal data fusion of near-infrared spectrum and image, combined with deep learning technology, and improves the discrimination accuracy by about 13%. By dynamically adjusting the model's dependence on different modal data through the fusion weight adjustment function, it effectively solves the problem that single modal data may fail under specific conditions and enhances the robustness of the model. At the same time, the method of the present invention can realize automated discrimination, significantly improve work efficiency, provide intelligent guidance for the tobacco leaf baking process, and has good practical value.

[0274] It should be noted that the variables involved in the present invention are explained in detail as shown in Tables 5, 6, 7 and 8 below.

[0275] Table 5 Variable Explanation Table (Part 1)

[0276]

[0277]

[0278] Table 6 Variable Explanation Table (Part 2)

[0279]

[0280] Table 7 Variable Explanation Table (Part 3)

[0281]

[0282] Table 8 Variable Explanation Table (Part 4)

[0283]

[0284]

[0285] The above description is only a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with this technical field can easily think of changes or replacements within the technical scope disclosed by the present invention, which should be covered by the scope of protection of the present invention.

Claims

1. A method for intelligently distinguishing the color and shape of tobacco leaves during an intelligent and precise baking process, characterized in that: include: Collect tobacco leaf near-infrared spectral data and tobacco leaf image data to establish tobacco leaf maturity grading standards; Preprocessing tobacco leaf near-infrared spectral data to obtain preprocessed tobacco leaf near-infrared spectral data, and extracting characteristic information of internal chemical components of tobacco leaves from the preprocessed tobacco leaf near-infrared spectral data; Image processing is performed on tobacco leaf image data to extract tobacco leaf color parameters and tobacco leaf morphological characteristic parameters, and to quantify the external characteristics of tobacco leaves; the principal component analysis method is used to perform dimensionality reduction processing on the pre-processed tobacco leaf near-infrared spectral data, tobacco leaf color parameters, and tobacco leaf morphological characteristic parameters, respectively, to obtain the principal component characteristics of tobacco leaf near-infrared spectra and tobacco leaf images; the principal component characteristics of tobacco leaf near-infrared spectra and tobacco leaf images are fused to obtain fused data, and the bimodal feature attention model is used to weight the feature importance of the fused data to obtain weighted fused data; a tobacco leaf state discrimination model is constructed based on the weighted fused data, and the fusion weight adjustment function is used to optimize the tobacco leaf state discrimination model parameters; the tobacco leaf data to be tested is input into the tobacco leaf state discrimination model to realize tobacco leaf color and shape state discrimination, and the tobacco leaf maturity grade discrimination result is output.

2. The method according to claim 1, characterized in that Tobacco leaf near-infrared spectral data refers to the spectral reflectance data obtained by scanning tobacco leaf samples through a near-infrared spectrometer within the near-infrared band, which mainly reflects the absorption characteristics of NH bonds, CH bonds and OH bonds inside the tobacco leaves.

3. The method according to claim 2, characterized in that Tobacco leaf image data refers to the RGB color images of tobacco leaves collected under standard light source conditions by a high-resolution digital camera, which contains visual feature information of tobacco leaf color, tobacco leaf texture and tobacco leaf morphology.

4. The method according to claim 3, characterized in that Tobacco leaf maturity grades include immature, moderately mature and over-mature.

5. The method according to claim 4, characterized in that Tobacco leaf chromaticity parameters refer to the hue, saturation, and lightness values ​​in the HSV color space extracted from the tobacco leaf RGB color image, which are used to quantitatively characterize the color characteristics of tobacco leaves.

6. The method according to claim 5, characterized in that Tobacco leaf morphological characteristic parameters refer to the geometric morphological characteristics of tobacco leaf area, tobacco leaf circumference, tobacco leaf aspect ratio, tobacco leaf roundness, tobacco leaf rectangularity and tobacco leaf vein distribution, which are used to quantitatively characterize the shape characteristics of tobacco leaves.

7. The method according to claim 6, characterized in that Principal component analysis refers to a statistical analysis method that reduces high-dimensional data to low-dimensional data. It converts the original features into new linearly independent features through orthogonal transformation, called principal components, which are used to reduce data redundancy and retain key information.

8. The method according to claim 7, characterized in that Data fusion refers to the combination of the principal component features of tobacco leaf near-infrared spectra and the principal component features of tobacco leaf images at the feature level.

9. The method according to claim 8, characterized in that The bimodal feature attention model refers to a deep learning model based on the attention mechanism, which realizes the adaptive fusion and weight distribution of multi-source data by learning the correlation and importance between different modal features.

10. The method according to claim 9, characterized in that The tobacco leaf state discrimination model refers to a classification model based on a deep convolutional neural network structure, which is used to receive weighted fusion data and output the tobacco leaf maturity grade discrimination results.

Citation Information

Patent Citations

  • Tobacco leaf classification method based on spectrum and machine vision coupling

    CN110705655A

  • Tobacco leaf data fusion curing characteristic parameter monitoring method based on curing barn observation window

    CN117668487A

  • Structural damage identification method and system in combination with multi-modal information and artificial intelligence

    CN118552795A

  • Tobacco replacement determination method based on multi-modal hybrid fusion

    WO2024259785A1

Cited By

  • Deep learning-based digital SERS kynurenine ultramicro detection method

    CN121049232A

  • Digital sers kynurenine ultra trace detection method based on deep learning

    CN121049232B

  • Intelligent polygonatum sibiricum grading system based on multi-mode visual spectrum

    CN121198624A

  • Double-branch lightweight tobacco leaf grading method based on mixed framework

    CN121459064A