Single tobacco leaf automatic grading method based on multilayer feature fusion network

Through the combination of a multi-layer feature fusion network and attention discrimination module, the problem of poor fusion of global and local features in single-piece tobacco leaf grading is solved, the grading accuracy is improved, labor costs are reduced, and the intelligence of tobacco leaf grading is promoted.

CN120355982APending Publication Date: 2025-07-22SOUTHWEST JIAOTONG UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510384247.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-22

AI Technical Summary

Technical Problem

The existing single-piece tobacco leaf grading method cannot effectively integrate global and local characteristics, resulting in low grading accuracy and cannot meet the actual needs of the tobacco industry.

Method used

Using a multi-layer feature fusion network method, the Swin Transformer and SMobileNetV2 feature extraction networks are used to extract global and local features, and the feature fusion is layer by layer through feature fusion branches, and hierarchical results are generated in combination with the attention discrimination module.

Benefits of technology

It significantly improves the accuracy of tobacco leaf grading, reduces labor costs, and promotes the intelligent development of tobacco leaf grading, which can effectively capture subtle differences between individual tobacco leaves of adjacent grades and reduces grading errors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120355982A_ABST
    Figure CN120355982A_ABST
Patent Text Reader

Abstract

The invention discloses a single tobacco leaf automatic grading method based on a multilayer feature fusion network, and belongs to the technical field of tobacco leaf intelligent grading. The method comprises the following steps: collecting a single tobacco leaf image sample set; training a single tobacco leaf grading model by using the single tobacco leaf image sample set to obtain a trained single tobacco leaf grading model; the single tobacco leaf grading model comprises a global branch used for acquiring global features of a single tobacco leaf image sample, a local branch used for acquiring local features of the single tobacco leaf image sample, and an image fusion module used for fusing the global features and the local features, the feature fusion branch is used for obtaining high-level fusion features; and the attention discrimination module is used for outputting a prediction result according to the high-level fusion features. And obtaining a to-be-tested single tobacco leaf image, and outputting a grading result of the to-be-tested single tobacco leaf image by using the trained single tobacco leaf grading model. According to the method, the problem of low grading accuracy caused by poor global and local feature fusion of the single tobacco leaf is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of intelligent tobacco grading, and particularly relates to an automatic grading method for single tobacco leaves based on a multi-layer feature fusion network. Background Art

[0002] Tobacco leaf grading is a time-consuming and laborious task. Manual grading is prone to problems such as fatigue-induced misgrading and low grading efficiency. To promote the intelligence of tobacco leaf grading and improve the grading accuracy, algorithms for single tobacco leaf grading have been studied.

[0003] Early tobacco leaf grading algorithms mainly used image processing technology to extract the appearance features of tobacco leaves, because a large part of the quality of tobacco leaves is presented on the appearance. Based on the extracted features, a classifier was used to predict the grade of a single tobacco leaf. The obvious drawback of such a method is that only the explicit features of tobacco leaves are extracted, and deeper features (such as oil content) cannot be mined. Therefore, existing technical solutions collect the spectral information of tobacco leaves, which can extract more internal features of tobacco leaves, but this method has complex feature processing and high costs for collection equipment, making it impossible to be widely used.

[0004] In recent years, with the booming development of deep learning, the field of image classification has developed rapidly. Applying deep learning algorithms to the task of single tobacco leaf grading has greatly improved the grading accuracy. Most methods based on deep learning use convolutional neural networks to directly extract the global features of single tobacco leaves for grading, and this method does not utilize the rich detailed features of single tobacco leaves. Therefore, the current mainstream research is to separately extract the global and local features of single tobacco leaves, and then use these two types of features to achieve accurate grading. Since the features of single tobacco leaves of different grades are similar, existing methods cannot effectively fuse the global and local features, resulting in the grading accuracy of single tobacco leaves not meeting the actual requirements of the tobacco industry. Summary of the Invention

[0005] Aiming at the above deficiencies in the prior art, the automatic grading method for single tobacco leaves based on a multi-layer feature fusion network provided by the present invention solves the problem of low grading accuracy caused by poor fusion of global and local features of single tobacco leaves.

[0006] To achieve the above invention purpose, the technical solution adopted by the present invention is: an automatic grading method for single tobacco leaves based on a multi-layer feature fusion network, including: Collecting images of single tobacco leaves, and performing data preprocessing and enhancement on the images of single tobacco leaves to obtain a sample set of images of single tobacco leaves; Train a single tobacco leaf grading model using a single tobacco leaf image sample set to obtain a trained single tobacco leaf grading model; the single tobacco leaf grading model includes a global branch for obtaining the global features of the single tobacco leaf image sample, a local branch for obtaining the local features of the single tobacco leaf image sample, a feature fusion branch for fusing the global features and the local features to obtain high-level fusion features, and an attention discrimination module for outputting a prediction result based on the high-level fusion features. Obtain an image of the single tobacco leaf to be measured, and use the trained single tobacco leaf grading model to output the grading result of the image of the single tobacco leaf to be measured.

[0007] Further, the global branch uses a Swin Transformer feature extraction network; based on the four groups of stacked Transformer Block structures of the Swin Transformer feature extraction network, the global features of the single tobacco leaf image sample in the first stage with a size of the global features of the single tobacco leaf image sample in the second stage with a size of the global features of the single tobacco leaf image sample in the third stage with a size of and the global features of the single tobacco leaf image sample in the fourth stage with a size of are obtained in sequence; is the height of the feature map; is the width of the feature map.

[0008] Furthermore, the local branch adopts the SMobileNetV2 feature extraction network; the SMobileNetV2 feature extraction network includes a convolutional layer with a 3×3 convolutional kernel, a first SMobileNetV2 Block feature extraction block, a second SMobileNetV2 Block feature extraction block, a third SMobileNetV2 Block feature extraction block, a fourth SMobileNetV2 Block feature extraction block, a fifth SMobileNetV2 Block feature extraction block, a sixth SMobileNetV2 Block feature extraction block, a seventh SMobileNetV2 Block feature extraction block, an eighth SMobileNetV2 Block feature extraction block, a ninth SMobileNetV2 Block feature extraction block, a tenth SMobileNetV2 Block feature extraction block, a first MobileNetV2 Block feature extraction block, a second MobileNetV2 Block feature extraction block, a third MobileNetV2 Block feature extraction block, an eleventh SMobileNetV2 Block feature extraction block, a twelfth SMobileNetV2 Block feature extraction block, a thirteenth SMobileNetV2 Block feature extraction block, and a fourteenth SMobileNetV2 Block feature extraction block, which are connected in sequence; the outputs of the third SMobileNetV2 Block feature extraction block, the sixth SMobileNetV2 Block feature extraction block, the third MobileNetV2 Block feature extraction block, and the fourteenth SMobileNetV2 Block feature extraction block are obtained in sequence and used as the local features of the single tobacco leaf image sample in the first stage with a size of , the local features of the single tobacco leaf image sample in the second stage with a size of , the local features of the single tobacco leaf image sample in the third stage with a size of , and the local features of the single tobacco leaf image sample in the fourth stage with a size of ; is the height of the feature map; is the width of the feature map.

[0009] Furthermore, the feature fusion branch is used to perform layer-by-layer fusion on the global features of the single tobacco leaf image samples in each stage and the local features of the single tobacco leaf image samples in the corresponding stage to obtain high-level fusion features.

[0010] Furthermore, the layer-by-layer fusion of the global features of the single tobacco leaf image samples in each stage and the local features of the single tobacco leaf image samples in the corresponding stage to obtain high-level fusion features is specifically as follows: Obtain the multi-channel attention of the global features of the single tobacco leaf image sample in the current stage, and multiply the multi-channel attention of the global features of the single tobacco leaf image sample in the current stage by the global features of the single tobacco leaf image sample in the current stage to obtain the channel-weighted global features in the current stage; Obtain the spatial attention of the local features of the single tobacco leaf image sample in the current stage, and multiply the spatial attention of the local features of the single tobacco leaf image sample in the current stage by the local features of the single tobacco leaf image sample in the current stage to obtain the spatially-information-weighted local features in the current stage; Downsample the fused features of the previous stage through 1×1 convolution and average pooling to obtain a feature map with a matching size ; concatenate with the global features of the single tobacco leaf image sample in the current stage and the local features of the single tobacco leaf image sample in the current stage in the channel dimension to obtain the result of fused feature channel concatenation; perform a 1×1 convolution operation on the result of fused feature channel concatenation to obtain the result of the convolution operation of the fused features; perform layer normalization on the result of the convolution operation of the fused features to obtain the first fused feature ; Concatenate the channel-weighted global features in the current stage, the spatially-information-weighted local features in the current stage, and the first fused feature activated by the GeLU activation function in the channel dimension to obtain the preliminary fused features in the current stage; Input the preliminary fused features in the current stage into an inverted residual multi-layer perceptron, and add the output result of the inverted residual multi-layer perceptron and the fused features of the previous stage through a skip connection to obtain the fused features in the current stage; Perform layer-by-layer fusion on the global features and local features of the single tobacco leaf image sample in each stage, and use the fused features of the fourth stage as the high-level fused features.

[0011] Further, the obtaining of the multi-channel attention of the global features of the single tobacco leaf image sample in the current stage is specifically as follows: Perform max pooling, average pooling, and soft pooling operations on the global features of the single tobacco leaf image sample in the current stage respectively to obtain the preliminary max pooling result of the global features, the preliminary average pooling result of the global features, and the preliminary soft pooling result of the global features, and perform linear operations and ReLU activation function activation on the preliminary max pooling result of the global features, the preliminary average pooling result of the global features, and the preliminary soft pooling result of the global features respectively to obtain the max pooling result of the global features, the average pooling result of the global features, and the soft pooling result of the global features; Add the max pooling result of the global features and the average pooling result of the global features element-wise to obtain the result of global pooling addition, and multiply the result of global pooling addition and the soft pooling result of the global features element-wise to obtain the result of global pooling multiplication; The result of the global pooling multiplication is linearly operated through the first linear layer to obtain the result of the global pooling linear operation, and the result of the global pooling linear operation is activated through the sigmoid activation function to obtain the multi-channel attention of the global features of the single tobacco leaf image sample in the current stage.

[0012] Furthermore, the specific method for obtaining the spatial attention of the local features of the single tobacco leaf image sample in the current stage is as follows: Perform max pooling and average pooling on the local features of the single tobacco leaf image sample in the current stage in the channel dimension to obtain the max pooling result of the local features and the average pooling result of the local features; Perform a channel concatenation operation on the max pooling result of the local features and the average pooling result of the local features to obtain the locally aggregated features with channel information; Convolve the locally aggregated features with channel information through a 7×7 convolutional layer to obtain the result of the local pooling convolution, and activate the result of the local pooling convolution through the sigmoid activation function to obtain the spatial attention of the local features of the single tobacco leaf image sample in the current stage.

[0013] Furthermore, the inverted residual multi-layer perceptron includes a depth convolutional layer with residual, a first convolutional layer, and a second convolutional layer; The depth convolutional layer with residual is used to perform a depth convolution operation on the preliminary fusion features in the current stage, and perform a skip connection between the result of the depth convolution operation activated by the GeLU activation function and the preliminary fusion features in the current stage to obtain the result of the skip connection; The first convolutional layer is used to perform a convolution operation on the result of the skip connection after layer normalization to obtain the result of the first convolution operation; The second convolutional layer is used to perform a convolution operation on the result of the first convolution operation to obtain the result of the second convolution operation, and perform layer normalization on the result of the second convolution operation to obtain the output result of the inverted residual multi-layer perceptron.

[0014] Furthermore, the specific method for outputting the prediction result according to the high-level fusion features is as follows: Input the high-level fusion features into the attention discrimination module to generate a discriminative feature representation ; Input into adaptive average pooling, and input the result of the adaptive average pooling into the second linear layer to output the prediction result.

[0015] Furthermore, the attention discrimination module includes five parallel branches; The first parallel branch is used to extract the feature map of the high-level fusion features by using the dilated convolution with the first dilation rate ; A second parallel branch for extracting a feature map of high-level fusion features using dilated convolution with a second dilation rate ; Add and element-wise, and perform global average pooling on the result of adding and element-wise to obtain the corresponding average pooling results of and ; Pass the corresponding average pooling results of and through two layers of pointwise convolutional layers for convolution operation to obtain the first global channel information; and pass the result of adding and element-wise through two layers of pointwise convolutional layers for convolution operation to obtain the first local channel information; Add the first global channel information and the first local channel information element-wise to obtain the result of adding the first channel information, and activate the result of adding the first channel information through a sigmoid activation function to obtain the attention weight ; Multiply and element-wise to obtain the result of multiplying and ; A third parallel branch for extracting a feature map of high-level fusion features using dilated convolution with a third dilation rate ; Obtain the residual attention weight through 1 - ; Multiply and element-wise to obtain the result of multiplying and ; Add the result of multiplying and and the result of multiplying and element-wise to obtain and fused features ; Add and element-wise, and perform global average pooling on the result of adding and element-wise to obtain the corresponding average pooling results of and ; Pass the corresponding average pooling results of and through two layers of pointwise convolutional layers for convolution operation to obtain the second global channel information; and pass the result of adding and ​The result of the element addition is convolved through two layers of pointwise convolutional layers to obtain the second local channel information; the second global channel information and the second local channel information are added element - by - element to obtain the result of the second channel information addition, and the result of the second channel information addition is activated through the sigmoid activation function to obtain the attention weights ; Multiply with element - by - element to obtain multiplied by the result of the multiplication; The fourth parallel branch is used to extract the feature map of the high - level fusion features by using the dilated convolution with the fourth dilation rate ; Obtain the residual attention weights through 1 - ; Multiply with element - by - element to obtain multiplied by the result of the multiplication; Multiply the result of with and the result of the multiplication, and add them element - by - element with the result of the multiplication of and to obtain and the fused features ; The fifth parallel branch is used to extract the feature map of the high - level fusion features by using the global pooling operation ; And perform channel concatenation processing on , , and to obtain the result of the branch channel concatenation, and reduce the channel dimension of the result of the branch channel concatenation through the third convolutional layer to obtain a discriminative feature representation ; The first dilation rate < the second dilation rate < the third dilation rate < the fourth dilation rate.

[0016] The beneficial effects of the present invention are as follows: Since the global and local features of single - piece tobacco leaves at different grades are similar, a single - piece tobacco leaf grading method with a multi - layer feature fusion network is designed, which effectively fuses the global (leaf shape) and local (vein, texture) features of single - piece tobacco leaves, significantly improves the accuracy of tobacco leaf grading, reduces the labor cost and promotes the intelligent development of tobacco leaf grading; An attention discrimination module is designed and applied to high - level fusion features to further make full use of global and local features, capture the features with subtle differences between adjacent - grade single - piece tobacco leaves, and reduce the grading errors of adjacent - grade single - piece tobacco leaves. Brief Description of the Drawings

[0017] Figure 1 is the flowchart of the method of the present invention.

[0018] Figure 2 Schematic diagram of single-piece tobacco leaf images at different grades collected in the embodiments of the present invention.

[0019] Figure 3 Flow chart of data preprocessing for single-piece tobacco leaf images in the embodiments of the present invention Figure 4 Schematic diagram of the single-piece tobacco leaf image after data augmentation in the embodiments of the present invention.

[0020] Figure 5 Network model diagram in the embodiments of the present invention

[0021] Figure 6 Specific module diagram of the network model in the embodiments of the present invention

[0022] Figure 7 Confusion matrix diagram of predictions obtained by using the proposed trained network model to test the test set in the embodiments of the present invention Specific implementation manners

[0023] The following describes the specific implementation manners of the present invention to facilitate those skilled in the art of the present technology to understand the present invention. However, it should be clear that the present invention is not limited to the scope of the specific implementation manners. For those of ordinary skill in the art of the present technology, as long as various changes are within the spirit and scope of the present invention defined and determined by the appended claims, these changes are obvious, and all inventions and creations using the concept of the present invention are within the scope of protection.

[0024] As Figure 1 shown, in an embodiment of the present invention, an automatic grading method for single-piece tobacco leaves based on a multi-layer feature fusion network includes: Collecting single-piece tobacco leaf images, and performing data preprocessing and augmentation on the single-piece tobacco leaf images to obtain a single-piece tobacco leaf image sample set; Training a single-piece tobacco leaf grading model using the single-piece tobacco leaf image sample set to obtain a trained single-piece tobacco leaf grading model; the single-piece tobacco leaf grading model includes a global branch for obtaining global features of the single-piece tobacco leaf image sample, a local branch for obtaining local features of the single-piece tobacco leaf image sample, a feature fusion branch for fusing the global features and the local features to obtain high-level fusion features, and an attention discrimination module for outputting a prediction result according to the high-level fusion features; Obtaining a single-piece tobacco leaf image to be measured, and using the trained single-piece tobacco leaf grading model to output a grading result of the single-piece tobacco leaf image to be measured.

[0025] In this embodiment, an industrial camera is used to collect single-piece tobacco leaf images and perform data preprocessing and augmentation, and the implementation method is as follows: Use an industrial RGB camera to take pictures of single tobacco leaves; In this embodiment, the industrial RGB camera MV-CS050-10GC of Hikvision is used, and the FA lens of Hikvision is adopted for the lens.

[0026] Make labels for the obtained single tobacco leaf images and divide them into a training set and a test set; In this embodiment, the classification gives image labels according to the GB2635-92 standard. As Figure 2 shown, a total of 5 grades of single tobacco leaves are collected. The resolution of the original collected images is 2448×2048, the image format is jpg, and a total of 1578 images are collected. According to each category, they are randomly divided into a training set and a test set according to the ratio of 7:3.

[0027] Perform a series of image preprocessing operations on the training set and the test set, and perform data augmentation on the preprocessed images; In this embodiment, first uniformly crop all the original pictures to a size of 2448×900, and then utilize the significant difference between the foreground and background of the tobacco leaf image in the HSV color space in the S component to completely separate the tobacco leaf target from the background; further obtain the minimum circumscribed rectangle of the tobacco leaf in the background-segmented image and expand it into a square. The entire preprocessing process is as Figure 3 . Perform data augmentation on the preprocessed training set and test set respectively. After augmentation, there are 3238 images in the training set and 440 images in the test set; the data augmentation includes horizontal flipping, vertical flipping, random rotation ([-45°, 45°]), and translation. The augmented pictures are as Figure 4 shown. These are all situations that may occur in the actual scenario, so as to balance the samples and enhance the generalization of the model.

[0028] The global branch adopts a Swin Transformer feature extraction network; based on the four groups of stacked Transformer Block structures of the Swin Transformer feature extraction network, the global features of the single tobacco leaf image samples in the first stage with a size of , the global features of the single tobacco leaf image samples in the second stage with a size of , the global features of the single tobacco leaf image samples in the third stage with a size of , and the global features of the single tobacco leaf image samples in the fourth stage with a size of are obtained in sequence; is the height of the feature map; is the width of the feature map.

[0029] In this embodiment, the global branch uses Swin Transformer to extract global features, and the local branch uses SMobileNetV2 to extract local features, as Figure 5 shown. Specifically, the input image is convolved with a convolution kernel of size 4×4 and a stride of 4 to obtain image patches of dimension C ×4×4. Then these image patches are flattened and projected to C dimensions, and a feature map of size is output. The feature map is then input into a global feature extractor composed of 12 Swin TransformerBlock feature extraction blocks and 3 patch merging operations. Each time a patch merging operation is performed, the feature map is downsampled by a factor of 2. Therefore, the sizes of the feature maps in Stage2, Stage3, and Stage4 are , and .

[0030] The local branch adopts the SMobileNetV2 feature extraction network; the SMobileNetV2 feature extraction network includes a convolutional layer with a convolution kernel of 3×3, a first SMobileNetV2 Block feature extraction block, a second SMobileNetV2 Block feature extraction block, a third SMobileNetV2 Block feature extraction block, a fourth SMobileNetV2 Block feature extraction block, a fifth SMobileNetV2 Block feature extraction block, a sixth SMobileNetV2 Block feature extraction block, a seventh SMobileNetV2 Block feature extraction block, an eighth SMobileNetV2 Block feature extraction block, a ninth SMobileNetV2 Block feature extraction block, a tenth SMobileNetV2 Block feature extraction block, a first MobileNetV2 Block feature extraction block, a second MobileNetV2 Block feature extraction block, a third MobileNetV2 Block feature extraction block, an eleventh SMobileNetV2 Block feature extraction block, a twelfth SMobileNetV2 Block feature extraction block, a thirteenth SMobileNetV2 Block feature extraction block, and a fourteenth SMobileNetV2 Block feature extraction block, which are connected in sequence; the outputs of the third SMobileNetV2 Block feature extraction block, the sixth SMobileNetV2 Block feature extraction block, the third MobileNetV2 Block feature extraction block, and the fourteenth SMobileNetV2 Block feature extraction block are obtained in sequence and used as the local features of the single tobacco leaf image sample in the first stage with a size of , the local features of the single tobacco leaf image sample in the second stage with a size of , the local features of the single tobacco leaf image sample in the third stage with a size of , and the local features of the single tobacco leaf image sample in the fourth stage with a size of ; is the height of the feature map; is the width of the feature map.

[0031] In this embodiment, in the local branch, First, through a convolution with a convolution kernel of 3×3, a dimension of The feature map is then input into a local feature extractor composed of 14 SMobileNetV2 Blocks and 3 MobileNetV2 Blocks. SMobileNetV2 has 17 inverted bottleneck layers, among which 4 inverted bottleneck layers use convolutional layers with a stride of 2 to compress the size of the feature map, so the change in the size of the feature map is consistent with that of the global branch. The SMobileNetV2 Block was proposed in this literature (Chen D, Zhang Y, He Z, et al. Feature-reinforced dual-encoder aggregation network for flue-cured tobacco grading[J]. Computers and Electronics in Agriculture, 2023, 210: 107887.).

[0032] The feature fusion branch is used to layer-by-layer fuse the global features of the single-leaf tobacco image samples at each stage and the local features of the single-leaf tobacco image samples at the corresponding stage to obtain high-level fusion features.

[0033] The layer-by-layer fusion of the global features of the single-leaf tobacco image samples at each stage and the local features of the single-leaf tobacco image samples at the corresponding stage to obtain high-level fusion features is specifically as follows: Obtain the multi-channel attention of the global features of the single-leaf tobacco image samples at the current stage, and multiply the multi-channel attention of the global features of the single-leaf tobacco image samples at the current stage by the global features of the single-leaf tobacco image samples at the current stage to obtain the channel-weighted global features at the current stage; Obtain the spatial attention of the local features of the single-leaf tobacco image samples at the current stage, and multiply the spatial attention of the local features of the single-leaf tobacco image samples at the current stage by the local features of the single-leaf tobacco image samples at the current stage to obtain the spatially-information-weighted local features at the current stage; Downsample the fusion features of the previous stage through 1×1 convolution and average pooling to obtain a feature map with a matching size ; Perform channel concatenation with the global features of the single-leaf tobacco image samples at the current stage and the local features of the single-leaf tobacco image samples at the current stage to obtain the result of channel concatenation of the fusion features; perform 1×1 convolution operation on the result of channel concatenation of the fusion features to obtain the result of the convolution operation of the fusion features; perform layer normalization on the result of the convolution operation of the fusion features to obtain the first fusion feature ; The channel-weighted global features at the current stage, the spatially-information-weighted local features at the current stage, and the first fusion feature activated by the GeLU activation function Perform channel splicing to obtain the preliminary fusion features at the current stage; Input the preliminary fusion features at the current stage into an inverted residual multi-layer perceptron, and add the output result of the inverted residual multi-layer perceptron to the fusion features of the previous stage through skip connection to obtain the fusion features at the current stage; Perform layer-by-layer fusion on the global features and local features of the single tobacco leaf image samples at each stage, and use the fusion features of the fourth stage as the high-level fusion features.

[0034] In this embodiment, the feature maps output by the local branch and the global branch in four stages are all input into the feature fusion branch. In the feature fusion branch, the output of the previous layer is used as the input to the next layer, as Figure 5 shown. In the feature fusion branch, a global and local feature fusion module is designed using channel and spatial attention mechanisms to adaptively fuse global and local features layer by layer and output high-level fusion features.

[0035] In this embodiment, as Figure 6 (a), the fusion features of the previous layer First, perform downsampling through 1×1 convolution and average pooling to obtain a feature map with a matching size , then concatenate with the global feature , the local feature perform channel splicing, 1×1 convolution operation and Layer Normal layer to obtain . and , Then, fuse the global, local and previous layer features through the GeLU activation function and channel splicing operation to obtain the preliminary fusion features.

[0036] The multi-channel attention for obtaining the global features of the single tobacco leaf image sample at the current stage is specifically: Perform max pooling, average pooling and soft pooling operations on the global features of the single tobacco leaf image sample at the current stage respectively to obtain the preliminary max pooling result of the global feature, the preliminary average pooling result of the global feature and the preliminary soft pooling result of the global feature, and perform linear operations and ReLU activation function activation on the preliminary max pooling result of the global feature, the preliminary average pooling result of the global feature and the preliminary soft pooling result of the global feature respectively to obtain the max pooling result of the global feature, the average pooling result of the global feature and the soft pooling result of the global feature; Add the max pooling result of the global feature and the average pooling result of the global feature element by element to obtain the result of adding global pooling, and multiply the result of adding global pooling with the soft pooling result of the global feature element by element to obtain the result of multiplying global pooling; The result of global pooling multiplication is linearly operated through the first linear layer to obtain the result of global pooling linear operation, and the result of global pooling linear operation is activated through the sigmoid activation function to obtain the multi-channel attention of the global features of the single tobacco leaf image sample at the current stage.

[0037] In this embodiment, the global features Output three feature maps through three pooling methods, and then input them into the linear shared layer to obtain three channel attention maps. The channel attention maps of average pooling and max pooling are added and then multiplied by the channel attention map of soft pooling. Finally, the channel attention weights are obtained through a linear layer and the Sigmoid function. This weight is then multiplied by to obtain the channel-weighted global features .

[0038] The specific method for obtaining the spatial attention of the local features of the single tobacco leaf image sample at the current stage is as follows: Perform max pooling and average pooling on the local features of the single tobacco leaf image sample at the current stage in the channel dimension to obtain the max pooling result of the local features and the average pooling result of the local features; Perform channel splicing operation on the max pooling result of the local features and the average pooling result of the local features to obtain the local features with aggregated channel information; Perform convolution on the local features with aggregated channel information through a 7×7 convolutional layer to obtain the result of local pooling convolution, and activate the result of local pooling convolution through the sigmoid activation function to obtain the spatial attention of the local features of the single tobacco leaf image sample at the current stage.

[0039] In this embodiment, as Figure 6 (a), the local features Perform average pooling and max pooling respectively in the channel dimension, and then obtain a feature map with a dimension of through channel splicing operation (H and W are the height and width). Input this feature map into a 7×7 convolutional layer to obtain a feature map with a dimension of . Obtain the spatial attention map through the Sigmoid function, and multiply this attention map by to obtain the locally feature with weighted spatial information .

[0040] The inverted residual multi-layer perceptron includes a depth convolutional layer with residual, a first convolutional layer, and a second convolutional layer; The depth convolutional layer with residual is used to perform depth convolution operation on the preliminary fusion features at the current stage, and perform skip connection on the result of the depth convolution operation activated by the GeLU activation function and the preliminary fusion features at the current stage to obtain the result of the skip connection; The first convolutional layer is used to perform a convolutional operation on the result of the layer-normalized skip connection to obtain a first convolutional operation result; The second convolutional layer is used to perform a convolutional operation on the first convolutional operation result to obtain a second convolutional operation result, and perform layer normalization on the second convolutional operation result to obtain the output result of the inverted residual multi-layer perceptron.

[0041] In this embodiment, the inverted residual multi-layer perceptron is composed of a depth convolutional layer with a residual and two ordinary convolutional layers. Specifically, the preliminarily fused features first pass through a 3×3 depth convolution and a GeLU activation function, and then the output is added to the preliminarily fused features through a skip connection. The added result then passes through two 1×1 convolutional layers and a Layer Normal layer to obtain an output. This output is added to the fused features of the previous layer after downsampling through a skip connection to obtain fused features . Through such a fusion process in four stages, high-level fused features are finally output.

[0042] Outputting a prediction result according to the high-level fused features specifically includes: Inputting the high-level fused features into an attention discrimination module to generate a discriminative feature representation ; Input into adaptive average pooling, and input the result of the adaptive average pooling into a second linear layer to output a prediction result.

[0043] The attention discrimination module includes five parallel branches; The first parallel branch is used to extract the feature map of the high-level fused features by using dilated convolution with a first dilation rate ; The second parallel branch is used to extract the feature map of the high-level fused features by using dilated convolution with a second dilation rate ; Add and element-wise, and perform global average pooling on the result of the element-wise addition of and to obtain and corresponding average pooling results; Pass the average pooling results corresponding to and through two layers of pointwise convolutional layers to perform a convolutional operation to obtain first global channel information; and and The result of adding elements is convolved through two layers of pointwise convolutional layers to obtain the first local channel information; the first global channel information and the first local channel information are added element-wise to get the result of adding the first channel information, and the result of adding the first channel information is activated through the sigmoid activation function to obtain the attention weights ; Multiply by element-wise to get the result of multiplying by ; The third parallel branch is used to extract the feature map of the high-level fusion features by using dilated convolution with the third dilation rate ; Obtain the residual attention weights through 1 - ; Multiply by element-wise to get the result of multiplying by ; Multiply the result of multiplying by and element-wise, and add the result to the result of multiplying by element-wise to obtain the fused features of and ; Add to element-wise, perform global average pooling on the result of adding and element-wise to get the corresponding average pooling result of and ; Convolve the corresponding average pooling results of and through two layers of pointwise convolutional layers to obtain the second global channel information; and convolve the result of adding and element-wise through two layers of pointwise convolutional layers to obtain the second local channel information; add the second global channel information and the second local channel information element-wise to get the result of adding the second channel information, and activate the result of adding the second channel information through the sigmoid activation function to obtain the attention weights ; Multiply by element-wise to get the result of multiplying by ; The fourth parallel branch is used to extract the feature map of the high-level fusion features by using dilated convolution with the fourth dilation rate ; Obtain the residual attention weights through 1 - ​​ ; Multiply the elements of and to obtain the result of multiplying and ; Add the result of multiplying and and the result of multiplying and element - by - element to obtain the fused feature of and ; ; The fifth parallel branch is used to extract the feature map of the high - level fused feature by using the global pooling operation ; And splice the channels of , , and to obtain the result of the branch channel splicing, and reduce the channel dimension of the result of the branch channel splicing through the third convolutional layer to obtain a discriminative feature representation ; The first dilation rate < the second dilation rate < the third dilation rate < the fourth dilation rate.

[0044] In this embodiment, as shown in Figure 6 (b), the high - level fused feature first passes through 5 parallel branches. Among them, 4 branches are dilated convolutions with dilation rates of 1, 6, 12, and 18 respectively, and the last branch is the global average pooling branch. Through these 5 branches, multi - scale context information of the high - level fused feature is obtained, and a feature map with diverse context - awareness ability is generated , , , , ; Compared with , it has rich local details. Compared with , it has rich global context information. Add and element - by - element, and successively pass the added result through global average pooling, point - wise convolution, ReLU activation function, and point - wise convolution to extract global channel information. Then pass the added result through point - wise convolution, ReLU activation function, and point - wise convolution to extract local channel information. Add the global channel information and the local channel information element - by - element and obtain the attention weight through the sigmoid function , and obtain the residual attention weight through 1 - ; Multiply and element - by - element, multiply and and Perform element-wise multiplication and then perform element-wise addition on the multiplication results to obtain and the fused features ; Compared with it has rich local details, compared with it has rich global context information. Combine and according to the fusion steps of and to obtain and the fused features ; Combine and by channel concatenation, and then concatenate the output result with and by channel concatenation. The concatenated result is output through a 1×1 convolution to obtain a discriminative feature representation F .

[0045] In this embodiment, under the supervision of the cross-entropy loss, a trained single-leaf tobacco automatic grading model is obtained, and the trained single-leaf tobacco automatic grading model is used to detect the single-leaf tobacco image to be detected, and the grading result of the single-leaf tobacco is obtained.

[0046] Sum over all samples and all classes to calculate the overall loss, and use this to supervise the training of the single-leaf tobacco grading model. As Figure 7 shown, 440 test sets after data augmentation are passed through steps S2 to S4 to obtain prediction results, which are then compared with the true labels to calculate the prediction accuracy of each grade and generate a confusion matrix.

[0047] To verify the effectiveness of the method of the present invention, data is collected at an actual tobacco purchasing station and made into a dataset for algorithm research. The dataset of the present invention includes 5 single-leaf tobacco grades, namely B2F, B3F, C2F, C3F, and C4F. To verify the model performance, the dataset is divided into a training set and a test set, and data preprocessing and augmentation are performed on the training set and the test set. After data augmentation, the total number of pictures is 3678, including 3238 in the training set and 440 in the test set. Regarding the optimization algorithm of the model, the SGD algorithm is used, the initial learning rate is set to 0.001, and the learning rate update strategy is StepLR.

[0048] The hardware devices are as follows: Processor Intel(R) Xeon(R) Silver 4110 CPU @ 2.10GHz; Memory (RAM) 32.0GB; Dedicated graphics card, NVIDIA GeForce GTX 2080Ti; System type, Windows10; Development tools, Python and the pytorch deep learning framework.

[0049] The evaluation metrics used are accuracy (Acc), macro F1 (macro-F1), average inference time (Time), and model parameter size (Weight).

[0050] The implementation plan verification consists of three parts.

[0051] ① Compare the proposed method with other classification algorithms using a popular single-leaf tobacco grading algorithm to verify the effectiveness and advancement of the proposed method. ② Conduct ablation experiments on the global and local feature fusion module and the attention discrimination module proposed in this solution to verify the contribution of each part to the method. ③ Design a comparative experiment on the global and local feature fusion method to verify the effectiveness of the fusion method proposed in this solution.

[0052] To verify the advancement of the method proposed in the present invention, the method of the present invention is compared with several state-of-the-art classification algorithms and single-leaf tobacco grading methods. Among them, FDANet removes the manually extracted feature part and completely uses the features extracted by deep learning for classification. The method of the present invention is significantly superior to other comparison methods, with the accuracy and macro-F1 increasing by 2.73% and 2.71% respectively compared to the optimal FDANet. Moreover, the accuracy of the method of the present invention is the best for all three tobacco grades, and the accuracy rankings for the other two grades are also good. Although the method of the present invention has 187.4MB more in parameter size and sacrifices 12.9ms in inference time compared to MobileNetV2, it still meets the memory and computing speed requirements of industrial embedded devices. In summary, the method of the present invention has better performance for the single-leaf tobacco grading task.

[0053] Table 1

[0054] This group of experiments is designed to verify the contributions of the global and local feature fusion module (MLFF) and the attention discrimination module (ADM) proposed in the present invention to the method. The experimental results are shown in Table 2. The experimental results show that by adding MLFF to the basic model, the accuracy rate is increased by 0.91%, and the macro-F1 score is increased by 0.94%. Based on this, by adding ADM, the accuracy rate is increased by 2.73% compared with the basic model, and the macro-F1 is increased by 2.72%. Therefore, the two proposed modules significantly improve the classification accuracy of single tobacco leaves. Since the network model adopts a three-branch parallel feature fusion method, the parameter size increases by 58.9 MB, but the final model inference time only sacrifices 1.5 ms, still meeting the actual industrial requirements.

[0055] Table 2

[0056] The present invention proposes to use the multi-layer feature global and local fusion method (MLFF) to fuse the global-local features of tobacco leaves. In order to verify the effectiveness of MLFF as a model feature fusion method, another 3 schemes are selected for comparison. Scheme 1: The global feature and the local feature are only fused using skip connections; Scheme 2: On the basis of Scheme 1, multi-channel attention is added to the global feature; Scheme 3: On the basis of Scheme 2, spatial attention is added to the local feature; Scheme 4 is MLFF proposed by the present invention, which uses a three-branch parallel method on the basis of the previous ones to fuse the global-local features layer by layer. The experimental results are shown in Table 3, where 1, 2, 3, and 4 in the table respectively refer to Scheme 1, 2, 3, and 4. Scheme 1 has the worst performance among these four schemes, with the lowest recognition accuracy for tobacco leaves, the accuracy rate is only 88.64%, and the Macro-F1 is only 88.69%. Since Scheme 1 only uses skip connections, the parameter size and inference time are the smallest, which are 123.MB and 19.4 ms respectively. For Scheme 2 and Scheme 3, with the addition of channel attention and spatial attention, the accuracy rate gradually increases, indicating that the attention mechanism promotes feature fusion. Scheme 4, that is, the method MLFF proposed by the present invention, has the highest recognition accuracy for single tobacco leaves, with an accuracy rate of 91.59%, which is 2.95% higher than Scheme 1. However, due to the three-branch layer-by-layer feature fusion, the model parameters increase by 59.3 MB compared with Scheme 1, but the inference time only sacrifices 1.4 ms.

[0057] Table 3

Claims

1. An automatic grading method for single tobacco leaves based on a multi-layer feature fusion network, characterized in that, Including: Collecting single tobacco leaf images, performing data preprocessing and enhancement on the single tobacco leaf images to obtain a single tobacco leaf image sample set; Training a single tobacco leaf grading model using the single tobacco leaf image sample set to obtain a trained single tobacco leaf grading model; the single tobacco leaf grading model includes a global branch for obtaining global features of single tobacco leaf image samples, a local branch for obtaining local features of single tobacco leaf image samples, a feature fusion branch for fusing the global features and local features to obtain high-level fusion features, and an attention discrimination module for outputting a prediction result based on the high-level fusion features; Obtaining a single tobacco leaf image to be measured, and using the trained single tobacco leaf grading model to output the grading result of the single tobacco leaf image to be measured.

2. The automatic grading method for single tobacco leaves based on a multi-layer feature fusion network according to claim 1, wherein, The global branch adopts a Swin Transformer feature extraction network; The global branch is based on four groups of stacked Transformer Block structures of the Swin Transformer feature extraction network, and sequentially obtains the global features of the single tobacco leaf image sample in the first stage with a size of , the global features of the single tobacco leaf image sample in the second stage with a size of , the global features of the single tobacco leaf image sample in the third stage with a size of , and the global features of the single tobacco leaf image sample in the fourth stage with a size of ; is the height of the feature map; is the width of the feature map.

3. The automatic grading method for single tobacco leaves based on a multi-layer feature fusion network according to claim 1, wherein The local branch adopts an SMobileNetV2 feature extraction network; The described SMobileNetV2 feature extraction network includes a convolutional layer with a 3×3 convolutional kernel, a first SMobileNetV2 Block feature extraction block, a second SMobileNetV2 Block feature extraction block, a third SMobileNetV2 Block feature extraction block, a fourth SMobileNetV2 Block feature extraction block, a fifth SMobileNetV2 Block feature extraction block, a sixth SMobileNetV2 Block feature extraction block, a seventh SMobileNetV2 Block feature extraction block, an eighth SMobileNetV2 Block feature extraction block, a ninth SMobileNetV2 Block feature extraction block, a tenth SMobileNetV2 Block feature extraction block, a first MobileNetV2 Block feature extraction block, a second MobileNetV2 Block feature extraction block, a third MobileNetV2 Block feature extraction block, an eleventh SMobileNetV2 Block feature extraction block, a twelfth SMobileNetV2 Block feature extraction block, a thirteenth SMobileNetV2 Block feature extraction block, and a fourteenth SMobileNetV2 Block feature extraction block, which are connected in sequence; the outputs of the third SMobileNetV2 Block feature extraction block, the sixth SMobileNetV2 Block feature extraction block, the third MobileNetV2 Block feature extraction block, and the fourteenth SMobileNetV2 Block feature extraction block are obtained in sequence and used as the local features of a single tobacco leaf image sample in the first stage with a size of , the local features of a single tobacco leaf image sample in the second stage with a size of , the local features of a single tobacco leaf image sample in the third stage with a size of , and the local features of a single tobacco leaf image sample in the fourth stage with a size of ; is the height of the feature map; is the width of the feature map.

4. The automatic grading method for single tobacco leaves based on a multi-layer feature fusion network according to claim 1, characterized in that, The feature fusion branch is used to perform layer-by-layer fusion on the global features of single tobacco leaf image samples at each stage and the local features of single tobacco leaf image samples at the corresponding stage to obtain high-level fusion features.

5. The automatic grading method for single tobacco leaves based on a multi-layer feature fusion network according to claim 4, wherein, The layer-by-layer fusion of the global features of single tobacco leaf image samples at each stage and the local features of single tobacco leaf image samples at the corresponding stage to obtain high-level fusion features is specifically as follows: Obtaining the multi-channel attention of the global features of single tobacco leaf image samples at the current stage, and multiplying the multi-channel attention of the global features of single tobacco leaf image samples at the current stage by the global features of single tobacco leaf image samples at the current stage to obtain the channel-weighted global features at the current stage; Obtaining the spatial attention of the local features of single tobacco leaf image samples at the current stage, and multiplying the spatial attention of the local features of single tobacco leaf image samples at the current stage by the local features of single tobacco leaf image samples at the current stage to obtain the spatially-information-weighted local features at the current stage; The fused features of the previous stage are downsampled through 1×1 convolution and average pooling to obtain a feature map with a matching size ; The is concatenated with the global features of the single tobacco leaf image sample in the current stage and the local features of the single tobacco leaf image sample in the current stage in the channel dimension to obtain the result of fused feature channel concatenation; Performing a 1×1 convolution operation on the fused feature channel splicing result to obtain the convolution operation result of the fused features; Perform layer normalization on the convolution operation result of the fused features to obtain the first fused feature ; Perform channel concatenation on the channel-weighted global features of the current stage, the spatially-information-weighted local features of the current stage, and the first fused features activated by the GeLU activation function to obtain the preliminary fused features of the current stage; Inputting the preliminary fused features at the current stage into an inverted residual multi-layer perceptron, and adding the output result of the inverted residual multi-layer perceptron and the fused features at the previous stage through a skip connection to obtain the fused features at the current stage; Performing layer-by-layer fusion on the global features of single tobacco leaf image samples at each stage and the local features of single tobacco leaf image samples, and using the fused features at the fourth stage as the high-level fusion features.

6. The automatic grading method for single tobacco leaves based on a multi-layer feature fusion network according to claim 5, wherein, The obtaining of the multi-channel attention of the global features of single tobacco leaf image samples at the current stage is specifically as follows: Performing max pooling, average pooling, and soft pooling operations on the global features of single tobacco leaf image samples at the current stage respectively to obtain the preliminary max pooling result of the global features, the preliminary average pooling result of the global features, and the preliminary soft pooling result of the global features, and performing linear operations and ReLU activation function activation on the preliminary max pooling result of the global features, the preliminary average pooling result of the global features, and the preliminary soft pooling result of the global features respectively to obtain the max pooling result of the global features, the average pooling result of the global features, and the soft pooling result of the global features; Add the global feature max pooling result and the global feature average pooling result element-wise to obtain the global pooling addition result, and multiply the global pooling addition result with the global feature soft pooling result element-wise to obtain the global pooling multiplication result; Perform a linear operation on the global pooling multiplication result through the first linear layer to obtain the global pooling linear operation result, and activate the global pooling linear operation result through the sigmoid activation function to obtain the multi-channel attention of the global feature of the single tobacco leaf image sample at the current stage.

7. The automatic grading method for single tobacco leaves based on a multi-layer feature fusion network according to claim 5, characterized in that, The obtaining of the spatial attention of the local feature of the single tobacco leaf image sample at the current stage is specifically as follows: Perform max pooling and average pooling on the local feature of the single tobacco leaf image sample at the current stage in the channel dimension respectively to obtain the local feature max pooling result and the local feature average pooling result; Perform a channel concatenation operation on the local feature max pooling result and the local feature average pooling result to obtain the locally aggregated feature with channel information; Perform convolution on the locally aggregated feature with channel information through a 7×7 convolutional layer to obtain the local pooling convolution result, and activate the local pooling convolution result through the sigmoid activation function to obtain the spatial attention of the local feature of the single tobacco leaf image sample at the current stage.

8. The automatic grading method for single tobacco leaves based on a multi-layer feature fusion network according to claim 5, characterized in that, The inverted residual multi-layer perceptron includes a depth convolutional layer with residual, a first convolutional layer, and a second convolutional layer; The depth convolutional layer with residual is used to perform a depth convolution operation on the preliminary fusion feature at the current stage, and perform a skip connection on the result of the depth convolution operation activated by the GeLU activation function and the preliminary fusion feature at the current stage to obtain the skip connection result; The first convolutional layer is used to perform a convolution operation on the result of the skip connection after layer normalization to obtain the first convolution operation result; The second convolutional layer is used to perform a convolution operation on the first convolution operation result to obtain the second convolution operation result, and perform layer normalization on the second convolution operation result to obtain the output result of the inverted residual multi-layer perceptron.

9. The automatic grading method for single tobacco leaves based on a multi-layer feature fusion network according to claim 1, wherein The outputting of the prediction result according to the high-level fusion feature is specifically as follows: Input the high-level fusion features into the attention discrimination module to generate discriminative feature representations ; Input into the adaptive average pooling, and input the result of the adaptive average pooling into the second linear layer to output the prediction result. Input into the adaptive average pooling, and input the result of the adaptive average pooling into the second linear layer to output the prediction result.

10. The automatic grading method for single tobacco leaves based on a multi-layer feature fusion network according to claim 9, wherein The attention discrimination module includes five parallel branches; The first parallel branch is used to extract the feature map of the high-level fusion features by using the dilated convolution with the first dilation rate ; A second parallel branch for extracting a feature map of high-level fusion features using dilated convolutions with a second dilation rate ; Add element-wise with , and perform global average pooling on the result of the element-wise addition of and to obtain the corresponding average pooling results of and ; Pass the corresponding average pooling results of and through two layers of pointwise convolutional layers for convolution operations to obtain the first global channel information; and pass the result of the element-wise addition of and through two layers of pointwise convolutional layers for convolution operations to obtain the first local channel information; Add the first global channel information and the first local channel information element-wise to get the result of adding the first channel information, and activate the result of adding the first channel information through the sigmoid activation function to obtain the attention weight ; Multiply element-wise with to obtain the result of the multiplication of and ; The third parallel branch is used to extract the feature map of the high-level fusion features by using dilated convolution with the third dilation rate ; Obtain the residual attention weight through 1- ; Multiply element-wise with to get the result of multiplying with ; Add the result of multiplying element-wise with and the result of multiplying to obtain the fused feature of and ; ; ; ; Add and element - by - element. Perform global average pooling on the result of adding the elements of and to obtain the corresponding average pooling results of and ; Add and 's corresponding average pooling results through two layers of point - wise convolutional layers to obtain the second global channel information; And add the results of adding the elements of and through two layers of point - wise convolutional layers to obtain the second local channel information; Add the second global channel information and the second local channel information element - by - element to get the result of adding the second channel information, and activate the result of adding the second channel information through the sigmoid activation function to obtain the attention weight ; Multiply and element - by - element to get the result of multiplying and ; The fourth parallel branch is used to extract the feature map of high-level fusion features by using dilated convolution with a fourth dilation rate ; through 1- Obtain the residual attention weight ; Multiply with element-wise to get and The result of the multiplication; Add the result of multiplying with element-wise, and add the result of multiplying with element-wise to obtain and The fused feature ; The fifth parallel branch is used to extract the feature map of high-level fusion features by means of global pooling operation ; and , , and are subjected to channel concatenation processing to obtain the result of branch channel concatenation, and the result of branch channel concatenation is passed through the third convolutional layer to reduce the channel dimension and obtain a discriminative feature representation ; The first dilation rate < the second dilation rate < the third dilation rate < the fourth dilation rate.