Plant disease detection method based on hybrid convolution attention mechanism
The LabInt-Net model addresses computational and training challenges in plant disease detection by integrating Lab color mode, convolutional neural networks, and multi-scale attention mechanisms, achieving over 93% accuracy in precision and recall for plant disease detection.
Patent Information
- Application Number
- CN202510485642.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-17
- Publication Date
- 2025-07-15
AI Technical Summary
The existing plant disease detection methods have problems such as excessive calculation volume, difficulty in training and insufficient detection accuracy in complex contexts.
The LabInt-Net detection model based on the hybrid convolutional attention mechanism is adopted. Through Lab color mode, convolutional neural network and multi-scale hybrid attention mechanism, combined with spatial reconstruction and multi-scale pooling, the attention of the model on disease characteristics is dynamically adjusted to reduce the impact of light interference and background noise.
The detection accuracy and adaptability of the model in complex backgrounds is improved, and efficient and accurate plant disease identification is achieved, which is especially suitable for disease detection tasks in complex backgrounds.
Smart Images

Figure CN120318691A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of deep learning image processing, and particularly to a plant disease detection method based on a hybrid convolutional attention mechanism. Background Art
[0002] Plant diseases have become one of the most serious problems in the global agricultural production structure. The severe spread of diseases has caused global food production losses of hundreds of billions of US dollars, and such phenomena are particularly obvious in less developed countries. Plant diseases not only cause a sharp reduction in crop yields, but also affect crop quality, thereby hindering the development of the global agricultural economy. With the continuous expansion of global climate change and agricultural production scale, the occurrence frequency and damage degree of pests and diseases are increasing day by day, and these factors bring more severe challenges to plant health management.
[0003] The detection of plant diseases initially relied on manual methods. Workers needed to identify diseases from a large number of plants and judge the types of lesions. However, manual detection has obvious limitations: long-term work leads to fatigue of personnel and reduced efficiency; due to differences in personnel experience, the judgment of the same disease may be different, resulting in inconsistent detection results.
[0004] Traditional machine learning methods rely on manual feature extraction. This requires operators to have certain professional knowledge to ensure the effectiveness and comprehensiveness of features. However, the features extracted manually are usually low-dimensional and easy to interpret, and it is difficult to significantly improve the detection accuracy.
[0005] Deep learning technology has become increasingly popular in recent years due to its powerful learning ability and intelligent feature extraction ability. After being fully trained, a deep learning model can also show good generalization ability when facing unknown data. As a rising star, deep learning is often used in fields such as computer vision and language processing. Especially in visual recognition, compared with traditional machine learning, it has great advantages in terms of convenience and accuracy.
[0006] Currently, popular network models include Make R-CNN proposed by He, which is based on Faster R-CNN. By adding a branch to generate the segmentation mask of the target, it can perform pixel-level segmentation. However, when performing instance segmentation, the computational cost is too high and the memory usage is particularly large, resulting in a slow detection speed; SSD proposed by Liu is a single-stage object detection algorithm that can directly detect on feature maps of different sizes. Therefore, it has a fast detection speed and can be used for real-time tasks. However, because it uses pre-set anchor boxes, SSD performs worse than some two-stage models when detecting small objects; SwinTransformer-B proposed by Liu is based on the Transformer architecture and proposes a mobile window mechanism, which is good at capturing global dependencies. However, compared with convolutional neural networks, Transformer is more difficult to train and requires a larger dataset and a longer training time; YOLOv7, as a relatively popular plant disease detection algorithm recently, is popular because of its fast speed, high precision, and lightweight design. Although YOLOv7 performs well in detecting large objects, it still has deficiencies in detecting small objects, and in complex backgrounds, YOLOv7 even has problems with a decrease in detection accuracy. Summary of the Invention
[0007] The purpose of the present invention is to provide a plant disease detection method based on a hybrid convolutional attention mechanism to solve the defects such as excessive computational cost and difficult training in the prior art.
[0008] A plant disease detection method based on a hybrid convolutional attention mechanism includes the following steps:
[0009] S1, Prepare the farm crops to be detected;
[0010] S2, Use the LabInt-Net detection model to detect the disease types of the crops;
[0011] Among them, the construction process of the LabInt-Net detection model includes the following steps:
[0012] S21, Obtain the picture data of the damaged crops in the actual farmland and construct a dataset according to the type classification;
[0013] S22, Based on the dataset, use the Lab color model, convolutional neural network, attention mechanism, and spatial reconstruction to establish an initial LabInt-Net detection model;
[0014] S23, Based on the dataset, iteratively optimize the current LabInt-Net detection model. When the detection accuracy and loss value reach a predetermined threshold, obtain the final LabInt-Net detection model.
[0015] Preferably, the specific steps in S21 include: preprocessing the collected data, mainly performing image cleaning and denoising operations, and dividing the dataset into a training set, a validation set, and a test set according to the ratio of 80%, 10%, and 10%.
[0016] Preferably, the specific steps in S22 include:
[0017] Using the Lab color model, divide the picture into a color channel and a lightness channel to reduce the interference of light on image details;
[0018] Use the combination of a convolutional neural network and spatial reconstruction to capture multi-scale spatial information;
[0019] Use a multi-scale hybrid attention mechanism to dynamically adjust the model's attention to disease characteristics and reduce the interference of background noise.
[0020] Preferably, the multi-scale hybrid attention mechanism specifically includes:
[0021] Use channel attention to enable direct interaction between channels to achieve complementary information;
[0022] Use a shuffle convolutional module to enable information flow within channels;
[0023] Use multi-scale spatial attention to extract multi-scale features and remove redundant features at the same time.
[0024] Preferably, the specific training steps of the LabInt-Net detection model in S23 are as follows:
[0025] S231, in the Lab color model, divide the input image into a color stream and a lightness stream;
[0026] S232, based on the color stream and the lightness stream, perform dual-branch independent feature mixing convolution, and multiply and fuse the convolution results to obtain a new feature map;
[0027] S233, based on the fused feature map, perform channel shuffling and a multi-scale hybrid attention mechanism to increase the contribution of key features to the model accuracy and remove redundant features to obtain an attention feature map;
[0028] S234, based on the attention feature map, generate a detection result through batch normalization and a fully connected layer.
[0029] Preferably, the feature mixing convolution specifically includes:
[0030] Adopt a spatial reconstruction unit and a pyramid pooling module to improve the model's spatial and structural sensitivity to different scale features;
[0031] Adopt deep convolution and multi-scale spatial attention mechanism to enhance the non-linear expression ability and convergence speed of the model.
[0032] Preferably, the loss function of the LabInt-Net detection model in the training stage is defined as:
[0033]
[0034] where y i is the true label (0 or 1), p i is the probability that the model predicts belonging to this category, and n is the number of samples.
[0035] Compared with the prior art, the present invention has the following advantages:
[0036] 1. LabInt realizes double-branch independent convolution in the Lab color mode, which can effectively reduce the influence of illumination. By combining multiple modules such as spatial reconstruction and dense feature extraction, the accuracy of the model is improved. And the superimposed use of multi-scale pooling and multi-scale convolution enables the model to maintain good detection accuracy even in complex background environments.
[0037] 2. Integrate spatial reconstruction and multi-scale pooling, that is, by capturing multi-scale spatial and channel information, enhance the expression ability of the model for shallow features, improve the adaptability and detection accuracy of the model. And use deep convolution to strengthen the information flow between channels and increase the fine-grainedness of the model.
[0038] 3. The multi-scale hybrid attention module dynamically adjusts the attention of the model to disease features, reduces the dependence of the model on the background, and reduces the interference of background noise. This is very important for disease detection in complex backgrounds in actual agricultural production. BRIEF DESCRIPTION OF THE DRAWINGS
[0039] Figure 1 is the overall flowchart of the plant disease detection method in the present invention.
[0040] Figure 2 is the partial flowchart of the plant disease detection method in the present invention.
[0041] Figure 3 is the specific implementation flowchart of the plant disease detection method in the present invention.
[0042] Figure 4 is the schematic diagram of the multi-scale hybrid attention model in the present invention.
[0043] Figure 5 is the schematic diagram of the reconstruction pooling convolution in the feature hybrid convolution.
[0044] Figure 6Schematic diagram of the depth attention convolution in the feature mixing convolution.
[0045] Figure 7 Schematic diagram of embedding a shuffle convolution module between channel and spatial attention. Detailed implementation manners
[0046] In order to make the technical means, creative features, achieved purposes and effects of the present invention easy to understand, the present invention will be further described below in conjunction with specific implementation manners.
[0047] As Figures 1 to 3 shown, a plant disease detection method based on a hybrid convolution attention mechanism is provided, including the following steps:
[0048] S1, Prepare the farm crops to be detected;
[0049] S2, Use the LabInt-Net detection model to detect the disease types of the crops;
[0050] Among them, the construction process of the LabInt-Net detection model includes the following steps:
[0051] S21, Obtain the picture data of the damaged crops in the actual farmland, and construct a data set according to the type classification;
[0052] Specifically, preprocess the collected data, mainly perform image cleaning and denoising work to ensure the image quality. Considering that the data background in the actual farmland is often complex and redundant, the size of the input image is uniformly modified to 32×32 pixels. This preserves the key features while greatly reducing the computational cost. After establishing the data set, divide the data set into a training set, a validation set, and a test set according to the ratio of 80%, 10%, and 10%. Among them, the training set data is used to train the prediction model, and the validation set data is used to select the hyperparameters of the model until the model effect is optimal.
[0053] S22, Based on the data set, use the Lab color model, convolutional neural network, attention mechanism, and spatial reconstruction to establish an initial LabInt-Net detection model;
[0054] LabInt-Net uses double-branch independent convolutions to extract color stream and lightness stream features respectively, retains key information through multi-scale pooling and spatial reconstruction. The dense feature extraction module enhances the disease area localization ability, realizes feature-level correlation learning through multiplication fusion, and effectively distinguishes diseases with similar textures. Introduce channel group shuffling to promote cross-channel information interaction, and combine the attention mechanism to dynamically screen key features. Finally, the detection result is output through the fully connected layer. This structure improves the disease detection accuracy while maintaining the lightweight characteristics through multi-modal feature fusion and channel recombination strategies, and is particularly suitable for plant disease recognition tasks under complex backgrounds.
[0055] Reasonably set the error threshold of the model, with a value range between 0.001 and 0.01, set the learning rate at 0.0003, and the maximum number of iterations is 50 times.
[0056] For the entire model, define its loss function during the training phase as follows:
[0057]
[0058] Among them, y i is the true label (0 or 1), p i is the probability that the model predicts to belong to this category, and n is the number of samples.
[0059] S23. Based on the dataset, iteratively optimize the current LabInt-Net detection model. When the detection accuracy and loss value reach a predetermined threshold, obtain the final LabInt-Net detection model, and then be able to accurately detect crop diseases.
[0060] In this embodiment, the specific training steps of the LabInt-Net detection model in S23 are as follows:
[0061] S231. In the Lab color mode, divide the input image into a color stream and a lightness stream;
[0062] S232. Based on the color stream and the lightness stream, perform dual-branch independent feature mixing convolution, and multiply and fuse the convolution results to obtain a new feature map; the feature mixing convolution includes reconstruction pooling convolution (such as Figure 5 ) and depth attention convolution (such as Figure 6 ).
[0063] The reconstruction pooling convolution covers two core parts: the spatial reconstruction unit (SRC) and the pyramid pooling convolution module (PPM), aiming to improve the spatial and structural sensitivity of the model to features of different scales. In the traditional PPM, information redundancy problems often occur due to simple splicing. Therefore, the reconstruction pooling convolution adds SRC before the PPM. SRC greatly reduces redundant information through operations such as convolution and batch normalization. The specific formula of SRC is as follows:
[0064]
[0065] F3 = ELU(BN3(Conv3(F2)))
[0066] In the formula, F represents the input feature map, ELU and GeLU represent activation functions, and BN represents batch normalization. To improve the non-linear expression ability, combine the effective non-linear transformation of ELU and the smooth gradient retention characteristic of GeLU to enhance the model expression ability.
[0067] The depth attention convolution adopts a hierarchical architecture. First, depthwise convolution (DW) is used to extract channel-independent spatial features, and the output F′ is used to reduce the computational complexity. Then, multi-scale hybrid attention is introduced to generate F′ a , focusing on lesions and suppressing redundancy. Element-wise fusion of F′ and F′ a After that, 3×3 convolution is used to extract deep features. At the same time, 1×1 convolution is used to adjust the channel dimension of the original F and fuse it multiplicatively, reducing the training oscillation. Finally, ELU activation is used to improve the non-linearity and convergence speed. This structure dynamically allocates resources, maintains high efficiency, and enhances complex feature learning.
[0068] S233, based on the fused feature map, performs channel shuffle and multi-scale hybrid attention mechanism to increase the contribution of key features to the model accuracy, remove redundant features, and obtain the attention feature map;
[0069] The multi-scale hybrid attention model (such as Figure 4 ) combines channel attention and multi-scale spatial attention to comprehensively extract features, which can effectively improve the detection accuracy.
[0070] The channel attention sub-module adopts a lightweight design: First, average pooling is performed on the feature map, then the combination of 1×1 convolution and ELU is used to enhance the non-linear expression ability, and then 1×1 convolution is used to restore the dimension. Finally, the attention map is generated by Sigmoid and fused with the original map. To improve the generalization of the model, a shuffle convolution module (CSC) (such as Figure 7 ) is embedded between channel and spatial attention, and its formulas are shown in (2), (3), and (4):
[0071]
[0072] y′ g,c′,h,w =y c′,g,h,w (4)
[0073] y″ = ELU(y′) (5)
[0074] Equation (2) is the convolution operation, and K c′,c,i,j is the weight of the convolution kernel on the input channel c′ and the output channel c; x c′,h+i,w+j is the value of the input feature map; y c,h,w is the value of the output feature map.
[0075] Equation (3) is the channel shuffle, g is the group index, and c′ is the shuffled channel index.
[0076] Equation (4) is the ELU activation function.
[0077] The spatial attention sub-module adopts a parallel multi-scale convolution structure: the feature map is respectively subjected to feature extraction through 3×3 and 5×5 convolution branches. After enhancing the non-linearity through an activation function, the number of channels is compressed to 1 using convolution, and then the spatial attention weight is generated through Sigmoid. The attention results of different scales are concatenated by channel. After restoring the channel dimension through 1×1 convolution, it is multiplied by the original feature to dynamically adjust the spatial feature distribution and focus on the lesion area.
[0078] S234, based on the attention feature map, generates detection results through batch normalization and a fully connected layer. When the detection results meet the preset detection conditions, the final LabInt detection model is obtained, and then it can accurately detect crop diseases.
[0079] To prevent the overfitting problem that occurs during the training process, an L2 regularization constraint is added on the basis of the cross-entropy loss function.
[0080] The L2 regularization constraint is achieved by adding an additional regularization term, which is the sum of the squares of the model parameters. In this way, it penalizes those overly large parameter values, preventing the model's parameters from becoming too large, thereby reducing the complexity of the model. A model with lower complexity usually has better generalization ability and can perform more stably on new data. The specific calculation formula is as follows:
[0081]
[0082] The hardware facilities of this implementation use an 18vCPU AMD EPYC 9754 128-Core Processor, equipped with RTX 4090D(24GB)*1 video memory and 60GB of memory. The model configuration information is as follows: the image input size (pixels) is 32×32, the batch size is 256, the number of channels is 3, the number of epochs is 50, the learning rate is 0.0003, the maximum norm threshold is 0.8, the weight decay is 1e-4, the number of groups is 16, and the cross-entropy loss function is selected.
[0083] Common model evaluation metrics mainly include: Accuracy, Precision, Recall, F1 Score, and mean Average Precision (mAP). The higher the values of the above metrics, the better. Their formulas are shown in (7), (8), (9), (10), (11):
[0084]
[0085] Among them, m is the number of correctly predicted samples, n is the total number of samples, TP is the true positive, FP is the false positive, FN is the false negative, and AP i is the average precision for the i-th category.
[0086] The following table shows the performance indicators of the embodiments.
[0087] Table 1 Performance Table
[0088]
[0089] It can be seen from the above data that the proposed model LabInt in this paper performs well in terms of indicators such as precision and recall, and both the F1 and mAP values reach 93%, indicating that the model can efficiently identify diseases and is suitable for complex tasks with high precision requirements.
[0090] In summary, the LabInt model constructed in this study is the result of innovative integration based on in-depth research on convolutional neural networks and attention mechanisms. Combining the powerful feature extraction ability of convolutional neural networks with the focusing advantage of attention mechanisms for key information, an efficient model suitable for plant disease detection is created. The evaluation indicators used in the model, such as accuracy, precision, recall, F1 value, and average precision mAP, are widely used in related fields and have been verified to effectively measure the performance of the model.
[0091] Aiming at the problems existing in traditional plant disease detection methods and existing deep learning detection models, this study makes full use of the multi-dimensional information of images and proposes a detection model that integrates a variety of advanced technologies. During the feature extraction process, considering the characteristics of plant disease images affected by factors such as illumination and complex backgrounds, the double-branch independent convolution in the Lab color mode is used to reduce illumination interference, spatial reconstruction and multi-scale pooling are used to mine shallow disease features, and deep convolution is used to optimize the calculation cost and model fine-grainedness. At the same time, a multi-scale hybrid attention mechanism is introduced to effectively reduce the interference of complex backgrounds and enhance the attention to disease features.
[0092] The model training results show that the LabInt model can accurately extract the key features of plant disease images. Even when facing images with similar textures or affected by complex backgrounds, it can accurately identify diseases. The experimental results also show that this model performs well in key indicators such as precision and recall, can well capture the change trend of plant disease features, make full use of image information to model feature associations, and thus significantly improve the accuracy of plant disease detection, having great application value and broad development prospects in actual agricultural production scenarios.
[0093] Therefore, the above-disclosed embodiments are illustrative in all respects and not exclusive. All changes within the scope of the present invention or within the scope equivalent to the present invention are encompassed by the present invention.
Claims
1. A plant disease detection method based on a hybrid convolutional attention mechanism, characterized in that It includes the following steps: S1. Prepare the farm crops to be detected; S2. Use the LabInt-Net detection model to detect the disease types of the crops; Among them, the construction process of the LabInt-Net detection model includes the following steps: S21. Obtain the picture data of the damaged crops in the actual farmland, and construct a data set according to the type classification; S22. Based on the data set, use the Lab color mode, convolutional neural network, attention mechanism and spatial reconstruction to establish an initial LabInt-Net detection model; S23. Based on the data set, iteratively optimize the current LabInt-Net detection model. When the detection accuracy and loss value reach the predetermined threshold, obtain the final LabInt-Net detection model.
2. The plant disease detection method based on the hybrid convolutional attention mechanism according to claim 1, wherein Specifically included in S21: Preprocess the collected data, mainly perform image cleaning and denoising work, and divide the data set into a training set, a validation set, and a test set according to the ratio of 80%, 10%, and 10%.
3. The plant disease detection method based on the hybrid convolutional attention mechanism according to claim 1, wherein Specifically included in S22: Use the Lab color mode to divide the picture into a color channel and a brightness channel, so as to reduce the interference of light on the image details; Use the combination of convolutional neural network and spatial reconstruction to capture multi-scale spatial information; Use the multi-scale hybrid attention mechanism to dynamically adjust the attention of the model to the disease characteristics and reduce the interference of background noise.
4. The plant disease detection method based on the hybrid convolutional attention mechanism according to claim 3, wherein The multi-scale hybrid attention mechanism specifically includes: Use channel attention to enable direct interaction between channels to achieve complementary information; Use the shuffle convolutional module to enable information flow within the channels; Use multi-scale spatial attention to extract multi-scale features and remove redundant features at the same time.
5. The plant disease detection method based on the hybrid convolutional attention mechanism according to claim 1, characterized in that The specific training steps of the LabInt-Net detection model in S23 are as follows: S231. In the Lab color mode, divide the input image into a color stream and a brightness stream; S232. Based on the color stream and the brightness stream, perform double-branch independent feature mixing convolution, and multiply and fuse the convolution results to obtain a new feature map; S233. Based on the fused feature map, perform channel shuffle and multi-scale hybrid attention mechanism to increase the contribution of key features to the model accuracy, remove redundant features, and obtain an attention feature map; S234. Based on the attention feature map, generate a detection result through batch normalization and a fully connected layer.
6. The plant disease detection method based on the hybrid convolutional attention mechanism according to claim 5, characterized in that, The feature mixing convolution specifically includes: Adopt a spatial reconstruction unit and a pyramid pooling module to improve the spatial and structural sensitivity of the model to different scale features; Adopt depth convolution and multi-scale spatial attention mechanism to enhance the nonlinear expression ability and convergence speed of the model.
7. The plant disease detection method based on the hybrid convolutional attention mechanism according to claim 1, characterized in that The loss function of the LabInt-Net detection model in the training stage is defined as: Among them, y i is the true label (0 or 1), p i is the probability that the model predicts belonging to this category, and n is the number of samples.