Lightweight wheat seed classification method fusing mixed attention mechanism

By adopting a lightweight classification method with a fusion and mixed attention mechanism in wheat seed classification, the problems of insufficient identification accuracy and high computing resource consumption in the existing technology are solved, efficient and low-cost wheat seed classification are achieved, and the development of agricultural intelligence is promoted.

CN119919718APending Publication Date: 2025-05-02HENAN INST OF SCI & TECH
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411975174.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-02

AI Technical Summary

Technical Problem

The prior art has problems such as insufficient identification accuracy and high computing resource consumption in crop seed classification, which is difficult to meet the needs of agricultural experts and farmers.

Method used

A lightweight wheat seed classification method with a fusion mixed attention mechanism is adopted. By constructing a lightweight classification model, combining a mixed attention module and a stacked inverted residual convolution layer, local features of different scales are extracted to reduce the amount of model parameters and calculations.

Benefits of technology

While maintaining high recognition accuracy, the storage space and computing costs of the model are significantly reduced, and are suitable for equipment with limited resources and promote the development of agricultural intelligence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119919718A_ABST
    Figure CN119919718A_ABST
Patent Text Reader

Abstract

The invention relates to a lightweight wheat seed classification method fusing a mixed attention mechanism, which comprises the following steps: respectively shooting collected wheat and each seed from two angles of an upward ventral ditch and a downward ventral ditch, taking black light-absorbing flannelette as a background during shooting, and obtaining and constructing a wheat grain image set; constructing a lightweight classification model, and inputting the wheat grain image set into the model for training and verification; obtaining a trained and verified lightweight classification model, and classifying the to-be-classified seeds through the trained and verified lightweight classification model; according to the lightweight classification model, a mixed attention mechanism and a stacked inverted residual convolutional layer are combined, so that the model parameter quantity and the calculation quantity are reduced under the condition that the network performance is not reduced; and more local features are learned by using convolution of different scales, so that a relatively small storage space is occupied while relatively high identification precision is maintained, and real-time classification and identification of the wheat grain image on a future low-performance terminal are facilitated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of computer vision image processing and crop seed classification, and in particular to a lightweight wheat seed classification method integrating a hybrid attention mechanism. Background Art

[0002] In the early research on the classification of crop seed varieties, traditional machine learning methods such as support vector machines (SVM), K-nearest neighbor (KNN) and decision trees [3-5] were widely used. These methods usually rely on manually extracted features such as color, shape, texture, etc., and use these features as input data for classification. For example, by extracting the color and shape features of the seeds, it is possible to classify different varieties of seeds. However, manual feature extraction requires a lot of domain knowledge and experience, and shows certain limitations when faced with complex and diverse seed image data;

[0003] Deep learning recognition methods are important methods in the field of image recognition in recent years and have been widely used in the study of crop species. Image recognition methods based on deep learning mainly include convolutional neural networks (CNN), recurrent neural networks (RNN) and deep belief networks (DBN). The characteristics of deep learning are that it can automatically learn features from data and can perform high-level abstract representation and learning layer by layer. Its advantage is that it can automatically learn features from data, avoiding the process of manual feature extraction, and can improve the generalization ability and robustness of the model. By using a large amount of data and high-performance equipment to train neural networks, the performance and effect of the model are further improved. However, the performance of crop seed image classification needs to be improved, and the computing resource requirements of each deep learning model are still high. Therefore, in order to meet the needs of agricultural experts and farmers, lightweight image classification models are imperative. Lightweight models can not only run on resource-constrained devices, but also significantly reduce computing costs and time consumption while ensuring classification accuracy. This will help to more widely apply image classification technology in actual agricultural production, improve work efficiency and promote the development of intelligent agriculture. Summary of the invention

[0004] In response to the needs in the prior art, the present invention provides a lightweight wheat seed classification method that integrates a hybrid attention mechanism, aiming to enable the model to occupy less storage space while maintaining high recognition accuracy.

[0005] A lightweight wheat seed classification method integrating hybrid attention mechanism includes the following steps:

[0006] Step 1: For each variety of wheat collected, 500 seeds with full grains were selected, and each seed was photographed at two angles, with the ventral groove upward and the ventral groove downward, with black light-absorbing flannel as the background, to obtain and construct a wheat grain image set;

[0007] Step 2: Build a lightweight classification model and input the wheat grain image set into the model for training and validation;

[0008] The lightweight classification model includes an input layer, a first convolutional layer, a maximum pooling layer, a hybrid attention module, a stacked inverted residual convolutional layer, a second convolutional layer, a global average pooling layer, a fully connected layer, and an output layer, which are connected in sequence. The hybrid attention module extracts local features of different scales by connecting channel attention and spatial attention in parallel, and the stacked inverted residual convolutional layer extracts deep features of the image by stacking basic units and downsampling units.

[0009] Step 3: Obtain a trained and verified lightweight classification model, and classify the seeds to be classified using the trained and verified lightweight classification model.

[0010] Further, a training set is divided from the wheat grain image set, and the training process of the training set in the lightweight classification model is as follows:

[0011] Step 2.1: Crop the training set to 224×224 through the input layer and normalize it;

[0012] Step 2.2: Preliminarily extract low-level features from the normalized training set through the first convolutional layer to obtain a low-level feature map;

[0013] Step 2.3: Reduce the size of the low-level feature map through the maximum pooling layer to obtain a small-size feature map;

[0014] Step 2.4: Use the channel attention in the hybrid attention module to extract feature information in the small-size feature map based on the ECA mechanism, and obtain feature map A; at the same time, use the spatial attention module in the hybrid attention module to extract feature information in the small-size feature map based on the TPSA mechanism, and obtain feature map B; add feature map A and feature map B to obtain the output of the hybrid attention module, which is recorded as feature map C;

[0015] Step 2.5: The feature map C obtained by the mixed attention enters the first core layer, the second core layer and the third core layer in the stacked inverted residual convolution layer in turn. The first core layer contains 1 downsampling unit and a basic unit repeated 4 times, and its input feature map dimension is transformed from 24×56×56 to 116×28×28; the second core layer contains 1 downsampling unit and a basic unit repeated 8 times, and its input feature map dimension is transformed from 116×28×28 to 232×14×14; the third core layer contains 1 downsampling unit and a basic unit repeated 4 times, and its input dimension is transformed from 232×14×14 to 464×7×7, and the final output feature map can be obtained, which is recorded as feature map D;

[0016] Step 2.6: After the feature map D passes through the second convolutional layer, the global average layer, the fully connected layer, and the output layer in sequence, the final output of the lightweight classification model is obtained, which is the predicted category;

[0017] Step 2.7: Pass the output of the fully connected layer to the Softmax layer, and measure the difference between the predicted category and the actual category through the cross entropy loss function in the Softmax layer;

[0018] Step 2.8: Optimize the model parameters of the lightweight classification model through the momentum stochastic gradient descent optimizer;

[0019] Step 2.9: Complete the training of the lightweight classification model.

[0020] Further:

[0021] In step 2.4, the specific steps to obtain feature map A are:

[0022]

[0023] w=σ(C1D k (y))

[0024]

[0025] Among them, X i,j (c) is a small-size feature map; y is the feature map output after global average pooling, and the transformation dimension is y∈R C×1 , H, W, and C represent the height, width, and number of channels of the input image, respectively; Represents y i The set of k adjacent channels, C1D represents one-dimensional convolution, k is the convolution kernel, σ is the sigmoid function, || Odd represents the nearest odd number; b,γ are constants with values ​​of 1 and 2 respectively; w is the channel attention weight;

[0026] Finally, the channel attention weight w obtained in the previous step is multiplied by the original input feature map to obtain the feature map A.

[0027] Further:

[0028] The specific steps to obtain the feature map B are:

[0029] First, global average pooling is performed on each channel of the small-size feature map to obtain the global feature vector G. Then, a 1×1 convolution is used to reduce the dimension of the global feature vector G to obtain the feature map G′.

[0030] The feature map G′ is convolved with two convolution kernels of different sizes to obtain the feature maps P1 and P2, which are then summed and fused.

[0031] Use 1×1 convolution to increase the dimension of the summed and fused feature map to obtain the feature map P′; perform Sigmoid activation on the feature map P′ to obtain the attention weight A;

[0032] Multiply the small-size feature map by the attention weight A element-wise to obtain the feature map B.

[0033] Further: In step 2.5,

[0034] The input of the basic unit is first processed by the channel segmentation layer. After the channel segmentation, one branch passes through the first convolution layer, BN layer, RELU function, depth-wise separable convolution layer, BN layer, second convolution layer, BN layer, RELU function in turn, and is concatenated with the original residual of the other branch, and finally the channel shuffled output is performed.

[0035] Further, the input of the downsampling unit is spliced ​​and connected after passing through two branches respectively, and finally output after channel shuffling; among them, one branch passes through the first convolution layer, BN layer, RELU function, depthwise separable convolution layer, BN layer, second convolution layer, BN layer, RELU function in sequence; the other branch passes through the depthwise separable convolution layer, BN layer, second convolution layer, BN layer, RELU function in sequence.

[0036] Further: The cross entropy loss function is:

[0037]

[0038] Among them, N 1 represents the number of categories in the data set, y i Represents the encoding of the true label category, is the distribution probability of the label category; is the predicted probability The natural logarithm of

[0039] R Lossis the average value of the loss function for each sample, which is:

[0040]

[0041] Where N 2 represents the number of samples in the data set; Represents the loss function for each sample.

[0042] Further: The momentum stochastic gradient descent optimizer is:

[0043]

[0044] V t =βV t-1 +(1-β)g

[0045] θ k =θ k-1 -ηw k

[0046] θ k-1 is the model parameter vector at step k-1, is the loss function L with respect to the parameter θ k-1 The gradient of , η is the learning rate, β is the momentum parameter, which is taken as 0.9 in this paper, and w k is the momentum, which represents the weighted accumulation of historical gradients.

[0047] Further: When the loss value of the cross entropy loss function slowly approaches the global optimal solution, the objective function is adjusted by cosine annealing to make it closer to the global optimal solution.

[0048] The beneficial effects of the present invention are as follows: the lightweight classification model combines the hybrid attention mechanism with the stacked inverted residual convolution layer, so that the model reduces the number of model parameters and the amount of calculation without reducing the network performance; and uses convolutions of different scales to learn more local features, while maintaining high recognition accuracy and occupying less storage space, which is helpful to realize real-time classification and recognition of wheat grain images on low-performance terminals in the future. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a flow chart of the present invention;

[0050] Figure 2 It is a schematic diagram of the structure of the lightweight classification model in the present invention;

[0051] Figure 3 Schematic diagram of the structure of the hybrid attention module in the present invention;

[0052] Figure 4 Schematic diagram of the structure of the basic unit in the stacked inverted residual convolution layer;

[0053] Figure 5 Schematic diagram of the structure of the downsampling unit in the stacked inverted residual convolution layer. DETAILED DESCRIPTION

[0054] The present invention is described in detail below in conjunction with the accompanying drawings. The embodiments of the present invention are described in detail below, and examples of the embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals throughout represent the same or similar elements or elements with the same or similar functions. The embodiments described below with reference to the accompanying drawings are exemplary and are only used to explain the present invention, and cannot be interpreted as limiting the present invention. The directional terms such as left, middle, right, top, and bottom in the embodiments of the present invention are only relative concepts or are based on the normal use state of the product, and should not be considered as restrictive.

[0055] A lightweight wheat seed classification method integrating hybrid attention mechanism includes the following steps:

[0056] Step 1: For each variety of wheat collected, 500 seeds with full grains were selected, and each seed was photographed at two angles, with the ventral groove upward and the ventral groove downward, with black light-absorbing flannel as the background, to obtain and construct a wheat grain image set;

[0057] Step 2: Build a lightweight classification model and input the wheat grain image set into the model for training and validation;

[0058] The lightweight classification model includes an input layer, a first convolutional layer, a maximum pooling layer, a hybrid attention module, a stacked inverted residual convolutional layer, a second convolutional layer, a global average pooling layer, a fully connected layer, and an output layer, which are connected in sequence. The hybrid attention module extracts local features of different scales by connecting channel attention and spatial attention in parallel, and the stacked inverted residual convolutional layer extracts deep features of the image by stacking basic units and downsampling units.

[0059] The training set is divided from the wheat grain image set. The training process of the training set in the lightweight classification model is as follows:

[0060] Step 2.1: Crop the training set to 224×224 through the input layer and perform normalization to standardize the input image for subsequent processing;

[0061] Step 2.2: Preliminarily extract low-level features, such as edge features and texture features, from the normalized training set through the first convolutional layer to reduce the spatial resolution and obtain a low-level feature map;

[0062] Step 2.3: Reduce the size of the low-level feature map through the maximum pooling layer to obtain a small-size feature map;

[0063] Step 2.4: Use the channel attention in the hybrid attention module to extract feature information in the small-size feature map based on the ECA mechanism, and obtain feature map A; at the same time, use the spatial attention module in the hybrid attention module to extract feature information in the small-size feature map based on the TPSA mechanism, and obtain feature map B; add feature map A and feature map B to obtain the output of the hybrid attention module, which is recorded as feature map C;

[0064] In step 2.4, the specific steps to obtain feature map A are:

[0065] Small size feature map X∈R H×W×C , converted to X∈R through a global average pooling 1×1×C , the resulting feature map y is then reshaped into X∈R C×1 to satisfy the requirements of the subsequent convolution operation; this is achieved by applying a one-dimensional convolution and utilizing the associated one-dimensional convolution weights; that is:

[0066]

[0067] w=σ(C1D k (y))

[0068]

[0069] Among them, X i,j (c) is a small-size feature map; y is the feature map output after global average pooling. To meet the requirements of subsequent convolution operations, the transformation dimension is y∈R C×1 , H, W, and C represent the height, width, and number of channels of the input image, respectively; Represents y i The set of k adjacent channels, C1D represents one-dimensional convolution, k is the convolution kernel, σ is the sigmoid function, || Odd represents the nearest odd number; b,γ are constants with values ​​of 1 and 2 respectively; w is the channel attention weight, which transforms the dimension and restores its shape, w∈R 1×1×C ;

[0070] Finally, multiply the channel attention weight w obtained in the previous step by the original input feature map to obtain the feature map A;

[0071] Step 2.5: The feature map C obtained by mixed attention enters the first core layer, the second core layer and the third core layer in the stacked inverted residual convolution layer. The first core layer contains 1 downsampling unit and a basic unit repeated 4 times, and its input feature map dimension is transformed from 24×56×56 to 116×28×28; the second core layer contains 1 downsampling unit and a basic unit repeated 8 times, and its input feature map dimension is transformed from 116×28×28 to 232×14×14; the third core layer It consists of a downsampling unit and a basic unit repeated 4 times. Its input dimension is transformed from 232×14×14 to 464×7×7, and the final output feature map is obtained, which is recorded as feature map D. Among them, the input of the basic unit is first processed by the channel segmentation layer. After the channel segmentation, one branch passes through the first convolution layer, BN layer, RELU function, depth-separable convolution layer, BN layer, second convolution layer, BN layer, RELU function in turn, and then is spliced ​​and connected with the original residual of the other branch, and finally the channel shuffle is output. The input of the downsampling unit passes through two branches and is spliced ​​and connected, and finally the channel shuffle is output; among them, one branch passes through the first convolution layer, BN layer, RELU function, depth-separable convolution layer, BN layer, second convolution layer, BN layer, RELU function in turn; the other branch passes through the depth-separable convolution layer, BN layer, second convolution layer, BN layer, RELU function in turn.

[0072] Both branches use depthwise separable convolution and channel shuffling. Depthwise separable convolution decomposes the standard convolution into two smaller operations: depthwise convolution and pointwise convolution. Depthwise convolution applies convolution kernels independently on each input channel, while pointwise convolution uses 1×1 convolution to combine the outputs of these channels. This method can greatly reduce the amount of calculation and parameters. The channel shuffling operation rearranges the channels of the feature map in the network to achieve information exchange between channels. This operation is usually performed after channel segmentation to ensure that features from different branches can be effectively fused together to improve the network's expressiveness.

[0073] This basic unit has fewer convolution operations on the right branch than the Shuffle Net V1 basic unit, and the degree of structural fragmentation is reduced. In addition, point-by-point grouped convolution is replaced by point-by-point convolution to reduce memory access costs. In addition, the three convolution operations have the same number of input and output channels, which is more conducive to improving network speed. After the convolution operation, the connection and channel rearrangement operations performed by the two branches can be combined with the Channel Split of the next unit to form an element-level operation, which reduces the number of element-level operations and helps improve computational efficiency;

[0074] The specific steps to obtain the feature map B are:

[0075] First, global average pooling is performed on each channel of the small-size feature map to obtain the global feature vector G. The global average pooling layer can effectively capture global information and reduce the spatial dimension of the feature map. Then, 1×1 convolution is used to reduce the dimension of the global feature vector G to obtain the feature map G′, thereby reducing the number of channels. In this paper, the number of channels is reduced to 1 / 16 of the original number.

[0076] The feature map G′ is convolved with two convolution kernels of different sizes. This embodiment uses 3×3 and 5×5 convolution kernels. The two convolution layers can capture features of different scales and enhance the diversity of feature extraction, thereby improving the expressiveness of the model. The feature map P1 and the feature map P2 are obtained, and then the feature map P1 and the feature map P2 are summed and fused.

[0077] Use 1×1 convolution to increase the dimension of the summed and fused feature map to obtain the feature map P′; perform Sigmoid activation on the feature map P′ to obtain the attention weight A; Sigmoid mainly normalizes the value of the feature map, and the purpose of increasing the dimension is to restore the original number of channels to keep it consistent with the input feature map;

[0078] Multiply the small-size feature map by the attention weight A element by element to obtain the feature map B;

[0079] Step 2.6: After the feature map D passes through the second convolutional layer, the global average layer, the fully connected layer, and the output layer in sequence, the final output of the lightweight classification model is obtained, which is the predicted category;

[0080] Step 2.7: Pass the output of the fully connected layer to the Softmax layer, and measure the difference between the predicted category and the actual category through the cross entropy loss function in the Softmax layer;

[0081] The cross entropy loss function is:

[0082] This function calculates the negative log-likelihood of the model's predictions, guiding the optimization of model parameters. The cross-entropy loss function provides a measure of model performance by evaluating the deviation between the predicted probability and the true label. Lightweight classification models use this loss function to quantify the difference between their output and the true label.

[0083]

[0084] Among them, N 1 represents the number of categories in the data set, y i Represents the encoding of the true label category, is the distribution probability of the label category; is the predicted probability The natural logarithm of

[0085] R Loss is the average value of the loss function for each sample, which is:

[0086]

[0087] Where N 2 represents the number of samples in the data set; represents the loss function of each sample; the cross entropy loss function quantifies the difference between the probability distribution of the true category and the probability distribution predicted by the model. It uses the characteristics of the logarithmic function to emphasize the penalty for wrong predictions, thereby guiding the model to continuously optimize during the training process and improve classification performance; the average loss function provides an overall performance indicator to help us understand the performance of the model on the entire data set; by combining the two, we can more comprehensively evaluate and optimize machine learning models;

[0088] Step 2.8: Optimize the model parameters of the lightweight classification model through the momentum stochastic gradient descent optimizer;

[0089] The momentum stochastic gradient descent optimizer is:

[0090]

[0091] V t =βV t-1 +(1-β)g

[0092] θ k =θ k-1 -ηw k

[0093] θ k-1 is the model parameter vector at step k-1, is the loss function L with respect to the parameter θ k-1 The gradient of , η is the learning rate, β is the momentum parameter, which is taken as 0.9 in this paper, and w k is the momentum, which represents the weighted accumulation of historical gradients;

[0094] In the process of using the gradient descent algorithm to optimize the objective function, when the loss value of the cross entropy loss function slowly approaches the global optimal solution, the update step size is adjusted through cosine annealing to make it smaller, so that the objective function is closer to the global optimal solution. The learning rate decay strategies include equal interval adjustment, exponential decay adjustment, cosine annealing adjustment, adaptive adjustment, custom adjustment, etc. This paper uses cosine annealing adjustment as the learning rate decay strategy; the cosine annealing algorithm reduces the learning rate through the cosine function, and the learning rate decays in a cosine function type. First, it slowly decreases with the cosine function to determine the correct optimization direction, and then it decreases rapidly to speed up the convergence of the function. If the gradient descent algorithm falls into a local optimal solution during training, the learning rate can be suddenly increased in the second half of the cosine function to jump out of the local optimal solution and find the path to the global optimal solution. It is a stochastic gradient descent method with restart; the cosine annealing algorithm is:

[0095]

[0096] Among them, i represents the index, represents the range of learning rate, which we set to 1 and 0.01 respectively, T cur represents the current cycle, and T i refers to the number of the i-th cycle;

[0097] Step 2.9: Complete the training of the lightweight classification model;

[0098] Step 3: Obtain a trained and verified lightweight classification model, and classify the seeds to be classified using the trained and verified lightweight classification model.

[0099] Experimental results of the comparison model on the wheat test dataset

[0100]

[0101] The table above gives a comprehensive comparison of the six comparison models in the wheat seed classification task. The experimental results show that LWheatNet (lightweight classification model) is considered the most effective model, with an accuracy of 98.99%, a minimum loss of 0.0413, and 1.33 million parameters. This performance highlights LWheatNet's ability to balance accuracy and computational efficiency. MobileNet V2 and MobileNet V3 also performed well, surpassing AlexNet and VGG16. It is worth noting that MobileNet V3's accuracy is 1.41 percentage points higher than AlexNet and 7.06 percentage points higher than VGG16, indicating that it uses fewer parameters and lower loss values ​​while ensuring accuracy in the classification task. ShuffleNet V2 further demonstrates its advantage through its efficiency, achieving an accuracy of 98.79% with only 1.26 million parameters. Although its loss value is slightly higher than LWheatNet, at 0.0545, this makes it one of the most resource-efficient models in the comparison. In contrast, VGG16 and AlexNet, despite their larger parameter sizes, did not lead to better performance in this case. The results show that increased model complexity is not associated with improved accuracy, indicating that model architecture and design are key factors in determining performance.

[0102] The above shows and describes the basic principles, main features and advantages of the present invention. It should be understood by those skilled in the art that the present invention is not limited to the above embodiments. The above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, which fall within the scope of the present invention. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A lightweight wheat seed classification method integrating hybrid attention mechanism, characterized by: The following steps are involved: Step 1: For each variety of wheat collected, 500 seeds with full grains were selected, and each seed was photographed at two angles, with the ventral groove upward and the ventral groove downward, with black light-absorbing flannel as the background, to obtain and construct a wheat grain image set; Step 2: Build a lightweight classification model and input the wheat grain image set into the model for training and validation; The lightweight classification model includes an input layer, a first convolutional layer, a maximum pooling layer, a hybrid attention module, a stacked inverted residual convolutional layer, a second convolutional layer, a global average pooling layer, a fully connected layer, and an output layer, which are connected in sequence. The hybrid attention module extracts local features of different scales by connecting channel attention and spatial attention in parallel, and the stacked inverted residual convolutional layer extracts deep features of the image by stacking basic units and downsampling units. Step 3: Obtain a trained and verified lightweight classification model, and classify the seeds to be classified using the trained and verified lightweight classification model.

2. The lightweight wheat seed classification method integrating hybrid attention mechanism according to claim 1 is characterized in that: The training set is divided from the wheat grain image set. The training process of the training set in the lightweight classification model is as follows: Step 2.1: Crop the training set to 224×224 through the input layer and normalize it; Step 2.2: Preliminarily extract low-level features from the normalized training set through the first convolutional layer to obtain a low-level feature map; Step 2.3: Reduce the size of the low-level feature map through the maximum pooling layer to obtain a small-size feature map; Step 2.4: Use the channel attention in the hybrid attention module to extract feature information in the small-size feature map based on the ECA mechanism, and obtain feature map A; at the same time, use the spatial attention module in the hybrid attention module to extract feature information in the small-size feature map based on the TPSA mechanism, and obtain feature map B; add feature map A and feature map B to obtain the output of the hybrid attention module, which is recorded as feature map C; Step 2.5: The feature map C obtained by the mixed attention enters the first core layer, the second core layer and the third core layer in the stacked inverted residual convolution layer in turn. The first core layer contains 1 downsampling unit and a basic unit repeated 4 times, and its input feature map dimension is transformed from 24×56×56 to 116×28×28; the second core layer contains 1 downsampling unit and a basic unit repeated 8 times, and its input feature map dimension is transformed from 116×28×28 to 232×14×14; the third core layer contains 1 downsampling unit and a basic unit repeated 4 times, and its input dimension is transformed from 232×14×14 to 464×7×7, and the final output feature map can be obtained, which is recorded as feature map D; Step 2.6: After the feature map D passes through the second convolutional layer, the global average layer, the fully connected layer, and the output layer in sequence, the final output of the lightweight classification model is obtained, which is the predicted category; Step 2.7: Pass the output of the fully connected layer to the Softmax layer, and measure the difference between the predicted category and the actual category through the cross entropy loss function in the Softmax layer; Step 2.8: Optimize the model parameters of the lightweight classification model through the momentum stochastic gradient descent optimizer; Step 2.9: Complete the training of the lightweight classification model.

3. The lightweight wheat seed classification method integrating hybrid attention mechanism according to claim 2 is characterized in that: In step 2.4, the specific steps to obtain feature map A are: in=σ(C1D k (y)) Among them, X i,j (c) is a small-size feature map; y is the feature map output after global average pooling, and the transformation dimension is y∈R C×1 , H, W, and C represent the height, width, and number of channels of the input image, respectively; represents the set of k adjacent channels of yi, C1D represents one-dimensional convolution, k is the convolution kernel, σ is the sigmoid function, || Odd represents the nearest odd number; b, γ are constants with values ​​of 1 and 2 respectively; w is the channel attention weight; Finally, the channel attention weight w obtained in the previous step is multiplied by the original input feature map to obtain the feature map A.

4. The lightweight wheat seed classification method integrating hybrid attention mechanism according to claim 3 is characterized in that: The specific steps to obtain the feature map B are: First, global average pooling is performed on each channel of the small-size feature map to obtain the global feature vector G. Then, a 1×1 convolution is used to reduce the dimension of the global feature vector G to obtain the feature map G′. The feature map G′ is convolved with two convolution kernels of different sizes to obtain the feature maps P1 and P2, which are then summed and fused. Use 1×1 convolution to increase the dimension of the summed and fused feature map to obtain the feature map p′; perform Sigmoid activation on the feature map p′ to obtain the attention weight A; Multiply the small-size feature map by the attention weight A element-wise to obtain the feature map B.

5. The lightweight wheat seed classification method integrating hybrid attention mechanism according to claim 2, characterized in that: In step 2.5, The input of the basic unit is first processed by the channel segmentation layer. After the channel segmentation, one branch passes through the first convolution layer, BN layer, RELU function, depth-wise separable convolution layer, BN layer, second convolution layer, BN layer, RELU function in turn, and is concatenated with the original residual of the other branch, and finally the channel shuffled output is performed.

6. The lightweight wheat seed classification method integrating hybrid attention mechanism according to claim 2, characterized in that: The input of the downsampling unit is spliced ​​and connected after passing through two branches, and finally output after channel shuffling; one branch passes through the first convolution layer, BN layer, RELU function, depthwise separable convolution layer, BN layer, second convolution layer, BN layer, RELU function in sequence; the other branch passes through the depthwise separable convolution layer, BN layer, second convolution layer, BN layer, RELU function in sequence.

7. The lightweight wheat seed classification method integrating hybrid attention mechanism according to claim 2, characterized in that: The cross entropy loss function is: Among them, N1 represents the number of categories in the data set, y i Represents the encoding of the true label category, is the distribution probability of the label category; is the predicted probability The natural logarithm of R Loss is the average value of the loss function for each sample, which is: Among them, N2 represents the number of samples in the data set; Represents the loss function for each sample.

8. The lightweight wheat seed classification method integrating hybrid attention mechanism according to claim 7 is characterized in that: The momentum stochastic gradient descent optimizer is: V t =βV t-1 +(1-β)g i k =θ k-1 -ηw k θ k-1 is the model parameter vector at step k-1, is the loss function L with respect to parameter θ k-1 The gradient of , η is the learning rate, β is the momentum parameter, which is taken as 0.9 in this paper, and w k is the momentum, which represents the weighted accumulation of historical gradients.

9. The lightweight wheat seed classification method integrating hybrid attention mechanism according to claim 8, characterized in that: When the loss value of the cross entropy loss function slowly approaches the global optimal solution, the objective function is adjusted by cosine annealing to make it closer to the global optimal solution.

Citation Information

Cited By

  • Feature extraction and weight reduction method for cattle face recognition scene

    CN120783369A