Apple leaf disease multi-classification method based on complex environment background

By constructing the LCAMNet model, combining the MSDM and MFEN modules and an improved triple attention mechanism, the accuracy and robustness of apple leaf disease identification in complex environments is solved, and efficient disease classification is achieved.

CN120411618AActive Publication Date: 2025-08-01INNER MONGOLIA AGRICULTURAL UNIVERSITY

Patent Information

Application Number
CN202510477090.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-01
Estimated Expiration
2045-04-16

AI Technical Summary

Technical Problem

The existing apple leaf disease recognition method based on convolutional neural network is insufficient in the context of complex environments, making it difficult to effectively capture diverse disease characteristics and environmental interference, resulting in low classification accuracy.

Method used

A lightweight fusion attention multi-branch network (LCAMNet) is constructed, and the multi-scale downsampling module (MSDM) and multi-scale feature extraction module (MFEN) are combined with an improved triple attention mechanism to capture the diverse feature information of apple leaf diseases, enhancing the robustness and accuracy of the model.

Benefits of technology

It improves the accuracy and robustness of apple leaf disease classification, is suitable for environments with limited resources and complex backgrounds, and has good versatility and efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120411618A_ABST
    Figure CN120411618A_ABST
Patent Text Reader

Abstract

The invention discloses an apple leaf disease multi-classification method based on a complex environment background, and the method comprises the following steps: constructing a common data set FGVC8 and a self-built complex environment background data set SCEBD, and carrying out the preprocessing of collected image data; constructing a multi-scale down-sampling module MSDM by combining a plurality of down-sampling strategies; constructing a multi-scale feature extraction module MFEN, and capturing diversified feature information of apple leaf diseases through a multi-branch structure; an improved triple attention mechanism is introduced into the MFEN, and key feature information is further extracted; and establishing a lightweight fusion attention multi-branch network LCAMNet model. According to the apple leaf disease multi-classification method based on the complex environment background, the accuracy and robustness of disease recognition are improved, and the method is suitable for apple leaf disease classification tasks with limited resources and complex backgrounds and has good universality and high efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the cross - technical field of agricultural information and computer vision, and particularly to a multi - classification method for apple leaf diseases under complex environmental backgrounds. Background Art

[0002] China is the largest apple producer in the world, with its annual output accounting for more than 50% of the global total output, ranking first in the world. The growth process of apples is vulnerable to the invasion of pathogens such as fungi and viruses, which can trigger a series of diseases. If early diseases are not detected and measures are not taken in time, the diseases will spread rapidly, affecting the quality and yield of apples and causing huge agricultural losses. Manual identification of apple leaf diseases is labor - intensive and time - consuming, and is easily interfered by subjective factors such as fatigue and mood. Therefore, it is very necessary to use information technology to achieve rapid detection of apple leaf disease types.

[0003] In the early stage, machine - learning - based methods were widely used for crop leaf disease identification. Mainly, global features (such as color, texture, shape) designed by hand were used to describe disease spots, and traditional image - processing techniques (such as edge detection, gray - level co - occurrence matrix, etc.) were used to extract features, and then the features were input into classifiers (such as SVM, KNN, etc.) for identification. However, these methods have significant drawbacks: one is that the global feature description ability is limited, and it is difficult to accurately capture the local features of complex disease spots; the other is that they are sensitive to noise, light changes and environmental factors in the image, resulting in unstable features and affecting the classification accuracy.

[0004] In recent years, with the rapid development of computer technology and the continuous improvement of image recognition accuracy, convolutional neural networks (CNNs) have shown significant advantages in automatic feature extraction and data recognition. By introducing operations such as local connection and weight sharing, CNNs can automatically learn and extract multi - level features from the original image without the need for manually designed features. This makes CNNs perform well in various image recognition tasks, especially in the field of crop disease recognition.

[0005] However, although CNNs have made significant progress in the task of crop disease recognition, when applying this deep - learning method to apple leaf disease classification, some major challenges still remain. The diversity of apple leaf disease spots and the complexity of their morphology make disease recognition extremely difficult. In addition, environmental factors such as light, shadow, and leaf reflection will also affect the image quality, and further affect the training and accuracy of the CNN model. Due to the diversity of apple leaf disease spots and environmental interference, how to improve the robustness and accuracy of the model for these complex scenarios is still an important problem in current research. Summary of the Invention

[0006] The object of the present invention is to provide a multi-classification method for apple leaf diseases under a complex environmental background, which improves the accuracy and robustness of disease recognition, is applicable to the classification task of apple leaf diseases with limited resources and complex backgrounds, and has good versatility and high efficiency.

[0007] To achieve the above object, the present invention provides a multi-classification method for apple leaf diseases under a complex environmental background, including the following steps:

[0008] Step S1, constructing a public dataset FGVC8 and a self-built complex environmental background dataset SCEBD, and preprocessing the collected image data;

[0009] Step S2, proving that two cascaded 3×3 convolutions are equal to a 5×5 convolution, and further comparing their computational amounts;

[0010] Step S3, constructing a multi-scale downsampling module MSDM by combining multiple downsampling strategies;

[0011] Step S4, constructing a multi-scale feature extraction module MFEN to capture diverse feature information of apple leaf diseases through a multi-branch structure;

[0012] Step S5, introducing an improved triplet attention mechanism into MFEN to further extract key feature information;

[0013] Step S6, based on the above constructed MSDM module, MFEN module and improved triplet attention mechanism, establishing a lightweight fusion attention multi-branch network LCAMNet model.

[0014] Preferably, in step S3, a multi-scale downsampling module MSDM is constructed by combining multiple downsampling strategies. The MSDM module includes two branches. The first branch includes two paths. By applying max pooling (i.e., path 1) and average pooling (path 2) in parallel, the global features and significant features in the image are respectively extracted;

[0015] First, average pooling extracts the global features of the image by averaging the pixel values within the pooling region, reducing the influence of local noise and enabling the model to capture the overall information of the image;

[0016] Second, max pooling extracts local significant features by selecting the maximum value in the pooling window, highlighting the texture and morphological changes in the disease area;

[0017] Then, by concatenating the results of max pooling and average pooling, the complementarity between significant features and global features is achieved;

[0018] Finally, 1×1 convolution, batch normalization BN and ReLU activation function are used to aggregate the channels.

[0019] Preferably, in step S3, the second branch includes path 2 and path 3, which are parallel convolutional downsampling modules;

[0020] In path 3, the input feature map first passes through a depth convolution of 3×3 with a stride of 2 + BN to retain and downsample the overall features; then passes through a 1×1 conv + BN + Relu to adjust the number of channels to ensure that the channels of the output feature map match those of path 1 and path 2 of the first branch;

[0021] In path 4, the input feature map first passes through a 1×1 conv + BN + Relu, then passes through a depth convolution of 3×3 with a stride of 2 + BN, and finally passes through a 1×1 conv + BN + Relu to extract local features; by concatenating the results of path 3 and path 4, the overall features and local features are fused;

[0022] The sampled outputs of the first branch and the second branch are concatenated in the channel dimension to integrate the global and local features extracted by the two branches. The concatenated feature map is further optimized through a channel shuffle operation to enhance the interaction between channels and promote information fusion.

[0023] Preferably, in step S4, a multi-scale feature extraction module MFEN is constructed to capture diverse feature information of apple leaf diseases by evenly distributing the input number of channels into four independent branches; among them, the four independent branches are: a feature retention branch, a local detail branch, a deep feature branch, and a significant feature branch.

[0024] Preferably, in the feature retention branch, the input features obtain output features through skip connections to retain the low-level information of apple leaf diseases.

[0025] Preferably, the local detail branch is used to extract the lesion edges, local texture changes, and disease regions of apple leaf disease pictures; in the local detail branch, the feature map input by the second branch first passes through a 1×1 convolution to adjust the number of channels of the feature map; then the feature map passes through a BN layer and a Relu layer; subsequently, the feature map passes through a 3×3 depth convolution, BN; the depth convolution performs convolution independently on each input channel, extracts local features on each channel, and is suitable for extracting fine-grained and independent feature information; finally, the feature map passes through a 1×1 point convolution to fuse the local features extracted from each channel in the depth convolution; by weighted combination of features of different channels, the model captures cross-channel global information.

[0026] Preferably, the deep feature branch extracts the multi-level representation of apple leaf disease images. The feature map input to the third branch first goes through a 1×1 convolution, BN, and Relu; then through two 3×3 depth convolutions, BN; and finally through a 1×1 convolution, BN, and Relu. Two cascaded 3×3 depth convolutions are used in this branch to reduce the network parameters and computational amount.

[0027] Preferably, in the significant feature branch, the feature map input to the fourth branch first extracts features through a max pooling layer; then, the feature map is further processed through a 1×1 convolution to ensure the fusion of features in different channels and improve the overall feature expression ability.

[0028] Perform a concatenation operation on the output feature maps of the above four branches, and introduce channel shuffle to optimize the information flow and fusion between different channels; the channel shuffle operation realizes the information exchange of features from different convolutional channels by rearranging the features of each channel.

[0029] Preferably, in step S5, the triple attention mechanism includes three branches. The first two branches are used to capture the interaction information between channels and between spaces respectively, and the third branch is used to establish a separate spatial attention. The outputs of these three branches are fused to generate the final attention feature.

[0030] Initial tensor X C×H×W , where C is the number of image channels, H is the image height, and W is the image width; for the first branch, the input shape remains C×H×W, which is responsible for aggregating the interaction features of H and W, i.e., the spatial dimensions; for the second branch, the shape of the tensor obtained through rotation is H×C×W, which is responsible for aggregating the interaction features of C and W, i.e., the channel and W spatial dimensions; for the third branch, the shape of the tensor obtained through rotation is W×H×C, which is responsible for aggregating H and C, i.e., the H spatial dimension and the channel.

[0031] After that, the triple tensor is passed to Z-pool. The Z-pool layer is responsible for reducing the zeroth dimension of the tensor to 2 by concatenating the average pooling and max pooling features in this dimension.

[0032] Then, two cascaded 3×3 convolutional layers, Relu, are used to reduce the zeroth dimension of the tensor to 1; the attention weights are generated by the sigmoid activation layer, then applied to the input tensor, and the rotated tensor is restored to the original input shape; the final attention mechanism is obtained by averaging and summing the results of the three branches.

[0033] Preferably, in step S6, a lightweight fusion attention multi-branch network LCAMNet model is established. The LCAMNet model includes a basic convolution Conv1 with a size of 3×3, a max-pooling layer with a 3×3 pooling kernel, and stages 2, 3, and 4, as well as a convolution Conv5 with a size of 1×1, an average pooling layer, and a fully connected layer;

[0034] Among them, stages 2, 3, and 4 all contain a multi-scale downsampling module MSDM and a multi-scale feature extraction module MFEN; an improved triple attention mechanism is stacked after the MFEN in stage 4;

[0035] For multi-channel input images, the LCAMNet model first preliminarily processes the input images through a Conv1 layer with a size of 3×3 and a stride of 2 to extract low-level features; then, a max-pooling layer with a size of 3×3 and a stride of 2 is used to downsample the images to reduce the dimension of the output features and obtain the output feature map;

[0036] Then, the feature map passes through stages 2, 3, and 4 to compress redundant information and gradually extract the features of apple leaf disease images; among them, each stage includes a cascaded multi-scale downsampling module MSDM and a multi-scale feature extraction module MFEN; the MSDM module is used to downsample the extracted features to improve the calculation efficiency while retaining key features; on this basis, the MFEN module is stacked to further capture the multi-scale features of the disease;

[0037] The MSDM and MFEN modules are stacked three times in total, and by introducing an improved triple attention mechanism at the end of stage 4, key feature information is further extracted to enhance the expression ability of the model;

[0038] Finally, a 1×1 convolutional layer Conv5 is connected to adjust the number of channels for feature fusion, and then it is processed by a 7×7 average pooling layer and a fully connected layer classifier for classification to output the final category result;

[0039] The training process of the LCAMNet model is as follows:

[0040] First, data preprocessing is performed, including data cleaning, image size adjustment, dataset division, data augmentation, and data normalization;

[0041] Next, the hyperparameters in the training process are set, specifically including the number of training epochs Epoch, batch size Batch size, optimizer, and loss function, and the model weights are initialized;

[0042] Then, the data is fed into the model in batches for training;

[0043] During the training process, each time a batch of data is sent to the training set, the Batch size is incremented by 1; if the Batch size reaches the maximum value, the Epoch is incremented by 1, and the next batch of data is continued to be input for training; if the Epoch reaches the set maximum value, the training ends, and the test set is used to evaluate the final performance of the model; if the Epoch does not reach the maximum value, the data is grouped according to the set batch size and sent into the model for continued training.

[0044] Therefore, the present invention adopts the above-mentioned multi-classification method for apple leaf diseases under a complex environmental background. Different downsampling operations are performed on different channels through the multi-scale downsampling module (MSDM), and channel shuffling is used to further fuse the features between different channels, effectively avoiding information loss caused by a single downsampling operation, thereby improving the recognition accuracy of the network; based on the multi-dimensional feature extraction network (MFEN), diverse feature information of apple leaf diseases is captured through a multi-branch structure; a channel separation and shuffling mechanism is introduced into the MFEN to promote feature communication between channels and further improve the feature extraction ability; an improved triple attention module is incorporated into the MFEN, and two cascaded 3×3 convolutions are used to replace the original 7×7 convolution, removing the original normalization operation. While enhancing the linear expression ability of the network, the number of parameters is reduced, the focusing ability of the model on key disease regions is enhanced, and the accuracy and robustness of disease recognition are improved; the LCAMNet of the present invention has both good generality and high efficiency, and is applicable to the apple leaf disease classification task with limited resources and complex backgrounds.

[0045] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Description of the Drawings

[0046] Figure 1 is the flowchart of the data preprocessing of the present invention;

[0047] Figure 2 is the dataset division of the FGVC8 apple leaf disease images of the present invention;

[0048] Figure 3 is the dataset division of the SCEBD apple leaf disease images of the present invention;

[0049] Figure 4 is the output schematic diagram of a 5×5 convolution and two cascaded 3×3 convolutions of the present invention; among them, (a) is the output schematic diagram of a 5×5 convolution; (b) is the output schematic diagram of two cascaded 3×3 convolutions;

[0050] Figure 5 is the structure diagram of the MSDM of the present invention;

[0051] Figure 6It is the MFEN structure diagram of the present invention;

[0052] Figure 7 It is the structure diagram of the improved triple attention mechanism of the present invention;

[0053] Figure 8 It is the LCAMNet structure diagram of the present invention;

[0054] Figure 9 It is the training flow chart of the LCAMNet model of the present invention;

[0055] Figure 10 It is the structure diagram of the feature extraction network and the downsampling network of ShuffleNetV2 in the embodiment of the present invention; among them, (a) is the structure diagram of the feature extraction network; (b) is the structure diagram of the downsampling network;

[0056] Figure 11 It is the three-dimensional visualization of the accuracy rate, floating-point numbers, and the number of parameters of the present invention on SCEBD;

[0057] Figure 12 It is the three-dimensional visualization of the accuracy rate, GFLOPs, and the number of parameters of the present invention on the FGVC8 dataset. Detailed implementation manners

[0058] The technical solutions of the present invention will be further described below with reference to the accompanying drawings and embodiments.

[0059] A multi-classification method for apple leaf diseases based on a complex environmental background of the present invention includes the following steps:

[0060] Step S1, construct a public dataset FGVC8 and a self-built complex environmental background dataset SCEBD, and preprocess the collected image data.

[0061] Step S11, construct a public dataset FGVC8.

[0062] Use the public dataset from the CVPR2021 FGVC8 Plant Pathology Recognition Challenge. This dataset contains 18,632 field-shot apple leaf images. These images were taken at different maturity stages of apples and at different times of the day, with non-uniform backgrounds, and most of the images have a resolution of 2676×4000. The apple leaf images in the dataset cover a variety of disease categories, including Alternaria leaf spot, rust, powdery mildew, scab, and healthy leaves.

[0063] Among them, the number of samples of Alternaria leaf spot, healthy leaves, powdery mildew, rust and scab selected in the present invention are 489, 529, 485, 503 and 504 respectively. After these images are processed, the FGVC8 dataset is formed, and its category distribution is shown in Table 1.

[0064] Table 1 Distribution of the number of picture categories in the FGVC8 dataset

[0065] Category Number Image Type Number of Pictures 0 Alternaria Leaf Spot 489 1 Healthy Leaves 529 2 Powdery Mildew 485 3 Rust 503 4 Scab 504

[0066] Step S12: Self-build a complex environment background dataset SCEBD.

[0067] The present invention also constructs an apple leaf disease dataset under a complex environment background - SCEBD, which uses four data sources. Among them, some Alternaria leaf spot, healthy, rust, powdery mildew and scab come from the FGVC8 dataset, and the number of pictures are 293, 183, 154, 485 and 484 respectively; some mosaic disease comes from Appleleaf9 (a total of 105 pictures), and the images in this dataset are fused with four different apple disease datasets, and the pixel sizes are not uniform; some gray spot disease comes from ATLDSD (a total of 121 pictures), and the images in this dataset are taken with a mobile phone, and the selected gray spot disease pictures come from real cultivated fields, and the pixel size is 256×256.

[0068] In addition, using ordinary smartphones, apple leaf images with Alternaria leaf spot, brown spot, gray spot, healthy, mosaic disease and rust in a certain area are collected, and the numbers are 192, 480, 362, 302, 376 and 328 respectively. The images are collected in real orchards, including complex environment backgrounds such as leaves, weeds, reflective films and fruits. The image resolutions are not uniform, but they can better reflect the complexity and diversity of the images. The image types and corresponding quantity distributions of SCEBD are shown in Table 2.

[0069] Table 2 Distribution of the number of picture categories in SCEBD

[0070]

[0071] Step S13: Data preprocessing.

[0072] As Figure 1 shown, for the collected picture data, data preprocessing is carried out, and the specific process is as follows:

[0073] Step S131: First, through data cleaning, select the picture data that meet the conditions.

[0074] Step S132: Scale all pictures to a size of 224×224 to reduce the computational complexity.

[0075] Step S133: Divide the dataset into a training set, a validation set, and a test set according to the ratio of 7:2:1. Among them, the dataset division of FGVC8 apple leaf disease images is as Figure 2 shown; the dataset division of SCEBD apple leaf disease images is as Figure 3 shown.

[0076] Step S134: Determine whether data augmentation is required. If data augmentation is required, perform data augmentation operations on the training set, while keeping the images in the validation set and the test set unchanged. If data augmentation is not required, skip this step.

[0077] Step S135: Determine whether data normalization processing is required. If normalization is required, set the mean and standard deviation of the RGB three channels of the data. If not, directly complete the data preprocessing.

[0078] Among them, the present invention applies nine data augmentation methods to the training set images to enhance the generalization ability of the network and reduce the influence of noise. Specifically, it includes rotation (90 degrees, 180 degrees, and 270 degrees), Gaussian blur, random flipping (50% probability of horizontal flipping and 50% probability of vertical flipping), contrast enhancement and reduction, and brightness enhancement and reduction. Through these enhancement operations, the number of images in the training set is increased to 10 times that of the original training set. It should be noted that the images in the test set and the validation set are not subjected to data augmentation.

[0079] In addition, in order to avoid unstable model training caused by too large or too small pixel values and reduce the risk of overfitting, the images are normalized. Specifically, the mean values of the red, green, and blue channels are respectively set to [0.485, 0.456, 0.406], and the standard deviation is set to [0.229, 0.224, 0.225]. This normalization method helps to accelerate the convergence of the model, improve the stability of training, and enable the model to more effectively learn image features.

[0080] Step S2: Prove that two cascaded 3×3 convolutions are equal to a 5×5 convolution, and further compare their computational amounts.

[0081] Step S21: Prove that cascaded 3×3 convolutions are equal to 5×5 convolutions.

[0082] Step S211: Assume that the size of the input image is H×W, the size of the convolution kernel is k×k, the stride is 1, and the padding is 0. C in is the number of input channels, and C out is the number of output channels. The output size H′×W′ after convolution is as follows:

[0083]

[0084] When the stride is 1, the formula is simplified to:

[0085] H′ = H - k + 1, W′ = W - k + 1 (2);

[0086] Step S212: For a 5×5 convolutional kernel, assume the input size is H×W, the padding is 0, and the stride is 1. According to the above formula, the output size after a 5×5 convolution is as follows:

[0087] H′ = H - 5 + 1 = H - 4, W′ = W - 5 + 1 = W - 4 (3);

[0088] Therefore, the size of the output feature map of a 5×5 convolutional layer is (H - 4)×(W - 4).

[0089] Step S213: For two consecutive 3×3 convolutions, first pass through the first 3×3 convolution. Let the size of the input image be H×W, then the output size is:

[0090] H′ = H - 3 + 1 = H - 2,,W′ = W - 3 + 1 = W - 2; (4);

[0091] Then, take the output of the first convolution as the input and pass through the second 3×3 convolution. The output size after the second convolution is:

[0092] H″ = H′ - 3 + 1 = H - 4, W″ = W′ - 3 + 1 = W - 4 (5);

[0093] Therefore, the size of the output feature map after two 3×3 convolutions is (H - 4)×(W - 4). This is equal to the size of the feature map of a 5×5 convolution. As Figure 4 shown, the graphical method can also illustrate that the receptive field of a 5×5 convolution is equivalent to two consecutive 3×3 convolutions.

[0094] Step S22: Further compare the computational amounts of cascaded 3×3 convolutions and 5×5 convolutions.

[0095] Step S221: Let the size of the input image be H×W, the size of the output image be H out ×W out , the size of the convolutional kernel is k×k, the stride is 1, and the padding is (k - 1) / 2. C in is the number of input channels, and C out is the number of output channels. When only considering the number of multiplication operations, the formula for the computational amount is:

[0096] C in ×C out ×H out ×W out×k 2 (6);

[0097] Its computational complexity is equivalent to C in ×C out ×H×W×k 2 .

[0098] Step S222: Consider the case of using a 5×5 convolutional kernel, then the computational complexity is:

[0099] O 5×5 = C in ×C out ×H×W×25 (7);

[0100] Step S223: Use two 3×3 convolutional kernels successively. The computational complexity of the first 3×3 convolution is:

[0101]

[0102] The computational complexity of the second 3×3 convolution is:

[0103]

[0104] Therefore, the total computational complexity of two cascaded 3×3 convolutions is as follows:

[0105]

[0106] Step S224: Compare the computational complexity of a 5×5 convolution with that of two cascaded 3×3 convolutions as follows:

[0107]

[0108] It can be seen from the above formula that the ratio of the computational complexity is related to C in and C out . When C in is relatively large, the computational complexity of two 3×3 convolutions is significantly lower than that of a 5×5 convolution. Therefore, using 2 consecutive 3×3 convolutions has the same field of view as a 5×5 convolution, but the computational complexity is lower than that of a 5×5 convolution.

[0109] Step S3: Construct a multi-scale downsampling module (MSDM) to effectively avoid the information loss problem that may be caused by a single downsampling operation by combining multiple downsampling strategies.

[0110] Downsampling operation is a common processing method in convolutional neural networks, mainly used to reduce the spatial dimension of feature maps. Pooling is one of the most commonly used downsampling methods, and common pooling types include max pooling and average pooling. Pooling reduces the size of the feature map by aggregating pixel values within a local region, thereby enhancing the model's robustness to translational variations.

[0111] In addition to pooling, convolution operations are also commonly used for downsampling. When the stride of the convolutional layer is set to 2, each convolution operation skips one pixel, thus halving the spatial size of the feature map. Different from pooling, convolution operations extract image features by learning the parameters of the convolutional kernel, and can retain more useful information during the downsampling process. The feature representation obtained through training in the convolutional layer can more effectively reduce information loss and is more adaptable in different tasks.

[0112] However, in a complex environmental background, the classification performance of apple leaf disease images is often weak. A single downsampling strategy is likely to lead to the loss of image features, thus affecting the classification results. To address this challenge, the present invention proposes a multi-scale downsampling module to reduce the information loss problem caused by a single downsampling operation. As Figure 5 shown, the multi-scale downsampling module contains two branches, aiming to extract richer feature information from multiple scales, thereby improving the classification accuracy.

[0113] The first branch (the blue background area, including path 1 and path 2) extracts the global features (path 1) and significant features (path 2) in the image by applying max pooling and average pooling operations in parallel. Average pooling extracts the global features of the image by averaging the pixel values within the pooling region, thereby reducing the influence of local noise and enabling the model to more effectively capture the overall information of the image, especially the global relationship between the complex background and the disease area.

[0114] Different from average pooling, max pooling extracts local significant features by selecting the maximum value in the pooling window, highlighting the most representative texture and morphological changes in the disease area, which is of great significance for distinguishing the disease spots from the significant areas in the complex environmental background. By concatenating the results of max pooling and average pooling, the complementarity between significant features and global features is achieved.

[0115] Subsequently, channel aggregation is performed using 1×1 convolution, batch normalization (BN), and the ReLU activation function. BN is a commonly used training optimization method that can improve the generalization ability of the model, avoid overfitting, and enhance the robustness of the network. ReLU mainly addresses the problem of linear inseparability in neural networks. By stacking a BN and ReLU after a linear transformation, the entire network can obtain a more powerful learning and fitting ability. This process helps to fuse the feature information from different pooling strategies, ensuring that the network can effectively integrate features at different levels, thereby enhancing the model's comprehensive representation ability of image features.

[0116] The second branch (the dashed box area, including Path 3 and Path 4) is a parallel convolutional downsampling module. In Path 3, the input feature map first passes through a depthwise convolution of 3×3 with a stride of 2 + BN to retain and downsample the overall features. The convolution operation with a stride of 2 effectively reduces the size of the feature map, and at the same time, through depthwise convolution, features can be independently learned on each channel. Then, through a 1×1 conv + BN + ReLU, the feature map adjusts the number of channels through the 1×1 convolutional layer to ensure that the channels of the output feature map match those of Path 1 and Path 2 in the first branch.

[0117] In Path 4, the input feature map first passes through a 1×1 conv + BN + ReLU, then through a depthwise convolution of 3×3 with a stride of 2 + BN, and finally through a 1×1 conv, BN, ReLU. This operation not only extracts local features but also effectively reduces the spatial size of the feature map. By concatenating the results of Path 3 and Path 4, the overall features and local features can be effectively fused.

[0118] After the first branch and the second branch complete their respective downsampling operations, their outputs will be concatenated in the channel dimension to integrate the global and local features extracted by the two branches respectively. The concatenated feature map will be further optimized through a channel shuffle operation to enhance the interaction between channels, thereby promoting better fusion of information. The downsampling process of MSDM, namely Algorithm 1, is shown in Table 3.

[0119] Table 3 Downsampling process of MSDM

[0120]

[0121]

[0122] Step S4: Construct a multi-scale feature extraction module (MFEN) to capture diverse feature information of apple leaf diseases through a multi-branch structure.

[0123] In the context of a complex environmental background, the classification ability of apple leaf diseases is often weak, and it is difficult for a convolutional kernel network of a single size to effectively extract accurate image features. To solve this problem, the present invention designs a multi-scale feature extraction network (MFEN) to extract multi-level features of apple leaf disease images. By introducing multi-scale feature extraction, the diversity of the extracted features can be enriched, thereby improving the classification accuracy.

[0124] The structure of MFEN is as Figure 6 shown. The input feature map first undergoes a channel separation operation, evenly distributing the input number of channels into four independent branches. Subsequently, the feature map will be processed through the following four branches, and the specific process is as follows:

[0125] Step S41, Feature Retention Branch.

[0126] In the Feature Retention Branch, the input features are directly transmitted to the output layer through skip connections to effectively retain low-level feature information, thereby alleviating the common feature degradation problem in deep networks. By retaining the original feature information in the first stage, this branch helps the network more accurately capture the subtle local changes in the image, avoiding the over-abstracting or loss of these key details during the deep convolution process. In the task of leaf disease classification, the diseased areas (such as brown spots, leaf discoloration, etc.) often have similar visual features to background noise (such as light changes, weeds, soil, or vein shadows). In this case, traditional convolutional neural networks may have difficulty distinguishing these subtle differences, especially when the background noise in the image is relatively complex. By introducing the Feature Retention Branch, the network can ensure that the subtle features of the diseased spots are not overlooked.

[0127] Step S42, Local Detail Branch.

[0128] The Local Detail Branch mainly extracts the edges of diseased spots, local texture changes, and tiny diseased areas in apple leaf disease pictures. In a complex environmental background, these local features are often affected by noise and background interference. Therefore, small-scale convolution can effectively focus on the detailed parts of the diseased spots, reducing the interference of the background on the classification results.

[0129] In the Local Detail Branch, the feature map input to the second branch first undergoes a 1×1 convolution to adjust the number of channels of the feature map. Then the feature map passes through the BN layer and the Relu layer. Subsequently, the feature map undergoes a 3×3 depth convolution, BN. Depth convolution performs convolution independently on each input channel and can extract local features on each channel. It is suitable for extracting fine-grained and independent feature information. Finally, the feature map undergoes a 1×1 point convolution to fuse the local features extracted from each channel in the depth convolution. By weighted combination of the features of different channels, it helps the model better understand the global information across channels.

[0130] Step S43, Deep Feature Branch.

[0131] The deep feature branch mainly extracts multi-level representations of apple leaf disease images, such as lesions of different sizes, morphological changes, etc. In the deep feature branch, the feature map input to the third branch first undergoes 1×1 convolution, BN, Relu; then two 3×3 depthwise convolutions, BN; and finally a 1×1 convolution, BN, Relu.

[0132] Based on the proof process in Step S2, using two cascaded 3×3 depthwise convolutions in this branch can significantly reduce network parameters and computational complexity. The two cascaded 3×3 convolutional layers form a deeper feature extraction unit that can obtain information from a richer feature space. This structure can further capture complex local patterns and more levels of details compared to a single 5×5 convolutional layer. By gradually increasing the number of convolutional layers, the network can gradually extract more abstract features while reducing the interference of background noise on the disease area.

[0133] Step S44, Salient Feature Branch.

[0134] In the salient feature branch, the feature map input to the fourth branch first extracts features through a max pooling layer. In a complex background environment, max pooling can effectively enhance the extraction of salient disease features. Especially when the contrast between the background and disease features is low, the pooling operation helps to retain key local information. Then, the feature map is further processed through 1×1 convolution to ensure better fusion of features in different channels, thereby enhancing the overall feature expression ability.

[0135] Step S45, Concatenate the output feature maps of the above four branches. This process effectively integrates features from different scales and levels, ensuring that the model can fully utilize the local, global, and salient features extracted by each branch. The concatenated feature map will contain rich multi-scale information, enhancing the model's overall perception ability of the disease area.

[0136] Step S46, Finally, introduce channel shuffle, aiming to further optimize the information flow and fusion between different channels. The channel shuffle operation rearranges the features of each channel, enabling better information exchange between features from different convolutional channels. The feature extraction process of MFEN, i.e., Algorithm 2, is shown in Table 4.

[0137] Table 4 Feature Extraction Process of MFEN

[0138]

[0139] Step S5, Introduce an improved triple attention mechanism in MFEN to further extract key feature information, thereby enhancing the model's expression ability.

[0140] The triple attention mechanism consists of three branches. The first two branches capture the interaction information between channels and between spaces respectively, and the third branch is used to establish separate spatial attention. The outputs of these three branches are fused to generate the final attention features. Specifically, the shape of the input tensor is C×H×W, and each branch is responsible for aggregating the cross-dimensional interaction features between the spatial dimensions H, W and the channel dimension C.

[0141] For the first branch, the input shape remains C×H×W, which is responsible for aggregating the interaction features of H and W, i.e., the spatial dimensions; for the second branch, the shape of the tensor obtained through rotation is H×C×W, which is responsible for aggregating the interaction features of C and W, i.e., the channel and W spatial dimensions; for the third branch, the shape of the tensor obtained through rotation is W×H×C, which is responsible for aggregating H and C, i.e., the H spatial dimension and the channel. Then the triple tensor is passed to Z-pool.

[0142] Z-pool(χ)=[MaxPool 0d (χ),AvgPool 0d (χ)] (12);

[0143] The Z-pool layer is responsible for reducing the zeroth dimension of the tensor to 2 by concatenating the average pooling and max pooling features in this dimension. This enables the layer to retain the rich representation of the tensor while reducing its depth. The dimensions of the three branches obtained are 2×H×W, 2×C×W, and 2×H×C respectively. Then, through a convolutional layer with a kernel size of 7×7, BN, and Relu, the zeroth dimension of the tensor is reduced to 1. The attention weights are generated by a sigmoid activation layer, then applied to the input tensor, and the rotated tensor is restored to the original input shape. The final attention mechanism is obtained by averaging and summing the results of the three branches.

[0144] On this basis, the present invention improves the original triple attention mechanism. Based on the proof process of step S2, two cascaded 3×3 convolutions are used to replace the original 7×7 convolution for feature extraction. It can maintain sufficient feature extraction ability without increasing the computational burden. Compared with the larger 7×7 convolution kernel, the two cascaded 3×3 convolution kernels can not only extract more abstract features, but also avoid the redundant calculations and higher number of parameters brought by using large convolution kernels.

[0145] In addition, between two 3×3 convolutions, a ReLU activation function is introduced. The ReLU function can effectively control the sparsity of the network, enhance the ability of non-linear transformation, and thus accelerate the gradient calculation and optimization process. By using the ReLU activation function, the mutual dependence between network parameters is reduced, thereby promoting more stable and efficient update of network information. It should be noted that in the present invention, the original batch normalization (BN) module is selected to be removed. Through experimental analysis, it is found that in this improved attention mechanism, the effect of the BN module is not obvious and will increase the computational burden to a certain extent. Therefore, in order to further improve the computational efficiency of the model and reduce unnecessary computational overhead, the BN layer is removed in the new architecture, making the network more concise while maintaining the stability of its training and inference. Through these improvements, not only the computational efficiency of the model is improved, but also the effectiveness of the cross-dimensional attention mechanism in capturing feature interactions and spatial information is ensured. The improved triple attention mechanism structure is as Figure 7 shown. The triple attention mechanism further extracts features, that is, Algorithm 3 is shown in Table 5.

[0146] Table 5 The triple attention mechanism further extracts features

[0147]

[0148] Step S6, based on the above constructed MSDM module, MFEN module and improved triple attention mechanism, establish a lightweight fusion attention multi-branch network (LCAMNet) model.

[0149] The present invention constructs a lightweight fusion attention multi-branch network (LCAMNet) model to solve the problem of apple leaf disease image classification in complex environmental backgrounds. Its overall structure is as Figure 8 shown, specifically including a basic convolution Conv1 of 3×3, a max-pooling layer with a 3×3 pooling kernel, and Stages 2, 3 and 4, as well as a convolution Conv5 of 1×1 size, an average pooling layer and a fully connected layer. Among them, Stages 2, 3 and 4 all contain a multi-scale downsampling module (MSDM) and a multi-scale feature extraction module (MFEN); an improved triple attention mechanism is also added after the MFEN in Stage 4.

[0150] As shown in Table 6, for a multi-channel input image with a size of 224×224×3, the LCAMNet model first performs preliminary processing on the input image through a Conv1 layer with a size of 3×3 and a stride = 2 to extract low-level features. Then, a max-pooling layer with a size of 3×3 and a stride = 2 is used to downsample the image to reduce the dimension of the output features and obtain an output feature map with a size of 56×56×24.

[0151] Then, the feature map goes through Stage 2, Stage 3, and Stage 4 to compress redundant information and gradually extract the features of apple leaf disease images. Each stage includes a cascaded multi-scale downsampling module (MSDM) and a multi-scale feature extraction module (MFEN). The MSDM module is used to downsample the extracted features, improving the computational efficiency while retaining key features. On this basis, the MFEN is stacked to further capture the multi-scale features of the disease. The MSDM and MFEN modules are stacked three times in succession to enhance the model's ability to fuse deep features. Each stacking extracts higher-level features based on the previous layer. After the three stages, the output feature dimensions are: 28×28×48, 14×14×96, and 7×7×192 respectively.

[0152] Next, at the end of Stage 4, by introducing an improved triple attention mechanism, key feature information is further extracted to enhance the model's expressive ability. The output dimension of this mechanism is 7×7×192.

[0153] Finally, a 1×1 convolutional layer Conv5 is connected to adjust the number of channels for feature fusion, and the resulting feature dimension is 7×7×1024. Subsequently, it is processed by a 7×7 average pooling layer to obtain a 1024-dimensional feature vector. The feature vector is classified through a fully connected layer classifier with 1024×8 to help the model make accurate classification decisions and output the final class results.

[0154] Table 6 LCAMNet architecture diagram

[0155]

[0156]

[0157] The training process of the LCAMNet model is as Figure 9 shown. First, data preprocessing is performed, including data cleaning, image size adjustment, dataset division, data augmentation, and data normalization. Then, the hyperparameters in the training process are set, specifically including the number of training epochs (Epoch), batch size (Batch size), optimizer, and loss function, and the model weights are initialized. Then, the data is fed into the model in batches for training.

[0158] During the training process, the model parameters are dynamically adjusted according to the results of the validation set to optimize the model performance. Each time a batch of data is sent to the training set, the Batch size is incremented by 1. If the Batch size reaches the maximum value, the Epoch is incremented by 1, and the next batch of data is continued to be input for training; if the Epoch reaches the set maximum value, the training ends, and the test set is used to evaluate the final performance of the model; if the Epoch does not reach the maximum value, the data is grouped according to the set batch size and sent to the model for continued training.

[0159] Embodiment

[0160] 1. Evaluation metrics of the LCAMNet model of the present invention.

[0161] In this embodiment, four metrics, Accuracy, Precision, Recall, and F1-Score, are used to evaluate the model. Their definitions are as follows: Accuracy is the ratio of the number of samples correctly detected for this class to the total number of samples; Precision is the proportion of samples that the model correctly predicts as a certain class among all samples predicted as this class; Recall is the proportion of samples that the model correctly predicts as a certain class among all samples that actually belong to this class; the F1 value is the harmonic mean of Accuracy and Recall. The calculation formulas for Accuracy, Precision, Recall, and F1 value are as follows:

[0162]

[0163]

[0164] Among them, TP represents the number of samples whose label is a positive sample and the model classification result is also a positive sample; TN represents the number of samples whose label is a negative sample and the model classification result is a negative sample; FP represents the number of samples whose label is a negative example and the model classification result is a positive example sample; FN represents the number of samples whose label is a positive example and the model classification result is a negative example.

[0165] For multi-object classification, there are mainly two statistical methods: macro-average and micro-average. In this embodiment, the macro-average is used as the value of Recall and Accuracy, and the global average is used as the value of Accuracy. Taking Recall as an example, the macro-average is calculated as follows:

[0166]

[0167] 2. Ablation experiment of SCEBD.

[0168] In this embodiment, ablation experiments are carried out on SCEBD to verify the effectiveness of the method proposed in the present invention. LCAMNet is improved based on the structure of the ShuffleNetV2 model. Therefore, several models are designed. Among them, Baseline is the ShuffleNetV2 model, and its feature extraction network and downsampling network are as Figure 10 shown.

[0169] Model1 only changes the feature extraction layer based on ShuffleNetV2. By replacing the Figure 6 MFEN of Figure 10 with the ShuffleNetV2 feature extraction network in (a) of

[0170] Model2 only changes the downsampling layer based on ShuffleNetV2. By replacing the Figure 5 MSDM in Figure 10 with the ShuffleNetV2 downsampling network in (b) of

[0171] Model3 changes both the feature extraction layer and the downsampling layer based on ShuffleNetV2.

[0172] Model4 adds improved triple attention based on Model3.

[0173] As shown in Table 7, GFLOPs represents the floating-point operation count of the model (in billions of operations) and is used to measure the computational complexity of the model; Parameters refers to the number of trainable weights in the model and is usually used to measure the scale of a deep learning model.

[0174] Table 7 Ablation experiments on the SCEBD dataset

[0175]

[0176]

[0177] It can be seen that the performance of Model1 on the test set is better than that of Baseline, indicating that compared with the ordinary feature extraction layer of ShuffleNetV2, MFEN can extract feature information more effectively in the classification of apple leaf disease images in complex environmental backgrounds.

[0178] The test performance of Model2 is better than that of Baseline, indicating that compared with the downsampling method used by ShuffleNetV2, applying MSDM to the downsampling process of apple leaf disease images can better retain feature information, thereby improving the recognition effect.

[0179] Model3 performed better than Model1 and Model2 on the test set, indicating that the combined application of MFEN and MSDM has a synergistic enhancement effect in the apple leaf disease image classification task, and its performance is better than the effect of applying the two separately, that is, Model3>Model1+Model2.

[0180] The test performance of Model 4 is further better than that of Model 3, indicating that although the introduction of the improved triple attention mechanism slightly increases the number of parameters, it significantly improves the model performance.

[0181] In addition, compared with the Baseline (ShuffleNetV2), the improved Model 4 still has a reduction in the number of floating-point operations and the number of parameters, while achieving a greater improvement in classification performance.

[0182] In summary, by gradually introducing MFEN, MSDM and the improved triple attention mechanism, the classification performance of the model continues to improve, verifying the effectiveness of the method proposed in this paper.

[0183] 3. Comparison with classic deep learning networks.

[0184] In this example, the performance of the proposed method is compared with that of classic convolutional neural network models such as VGG, ResNet, ResNext, and DenseNet on a self-built SCEBD dataset. The experimental results are shown in Table 8.

[0185] Table 8 Performance comparison of different deep learning models in SCEBD

[0186] Models Accuracy Recall Precision F1 GFLOPs Parameters VGG13 89.03 89.04 89.15 89.03 11.36G 133.05M ResNet34 88.52 88.54 89.02 88.66 3.68G 21.80M ResNext50 85.97 85.99 86.48 86.12 4.29G 25.03M DenseNet201 90.56 90.56 90.67 90.60 4.39G 20.01M MobileVit 91.84 91.85 91.91 91.83 0.27G 1.27M GoogLeNet 92.09 92.09 92.20 92.11 1.60G 7.01M InceptionResNetV2 91.58 91.59 92.14 91.71 6.50G 55.84M EfficientNetB0 91.33 91.35 91.54 91.34 0.42G 8.43M ShuffleNetV2 88.01 88.00 88.98 88.21 0.04G 1.31M MobileNetV2 89.80 89.79 90.52 89.93 0.33G 3.51M GhostNet 90.31 90.35 91.09 90.45 0.16G 5.18M LCAMNet 92.60 92.64 92.85 92.64 0.03G 1.30M

[0187] The LCAMNet proposed in this paper is based on the ShuffleNetV2 model and integrates technologies such as multi-branch feature extraction, multi-scale downsampling modules, and attention mechanisms. It uses a simpler overall model structure to minimize network parameters. Therefore, among all the models involved in the experiment, the number of parameters of LCAMNet is lower than that of ShuffleNetV2. Compared with classic deep learning networks, LCAMNet has the best classification performance in apple leaf disease images, which shows that LCAMNet is a high-performance, lightweight apple leaf disease image recognition model. Through the three-dimensional visualization analysis of accuracy, GFLOPs and parameter quantity, the results show that LCAMNet has a high performance and lightweight apple leaf disease image recognition model. Figure 11 As shown in Table 8, it further verifies that LCAMNet has excellent classification ability and computational efficiency compared with other deep learning models in Table 8.

[0188] 4. Comparative study on the performance of similar crop disease image classification models.

[0189] In this embodiment, the performance of the proposed method is compared with that of the same type of crop disease image classification model on the SCEBD dataset, as shown in Table 9.

[0190] Table 9 Comparison of the performance of LCAMNet and the same type of crop disease image classification model on SCEBD

[0191] Models Accuracy Recall Precision F1 GFLOPs Parameters Re-GoogLeNet 93.37 93.36 93.79 93.46 2.80G 9.11M ALS-Net 91.07 91.09 91.28 91.13 0.73G 1.18M LBMRNet 90.05 90.06 91.54 90.30 0.17G 0.91M LCAMNet 92.60 92.64 92.85 92.64 0.03G 1.30M

[0192] It can be seen that the performance of the ALS-Net and LBMRNet algorithms is inferior to that of the LCAMNet algorithm. This is because although ALS-Net improves ShuffleNetV2, its multi-scale feature acquisition is only carried out in the initial stage of the network and fails to effectively fuse these features into each stage, thus limiting the feature extraction ability. Although LMBRNet improves the feature extraction layer and the downsampling layer, its multi-residual structure continuously retains useless information, affecting the final recognition effect. In contrast, Re-GoogLeNet optimizes GoogLeNet, introduces an attention mechanism to retain the important features of GoogLeNet, and retains some original features through residual connections. Although its performance is better than that of LCAMNet, its large floating-point number and high number of parameters may not be suitable for resource-constrained scenarios.

[0193] 5. Model comparison based on public datasets.

[0194] In this paper, the proposed method is applied to the public dataset FGVC8 together with several classic deep learning models and the same type of crop disease image classification models for testing to verify the effectiveness and superiority of the method of the present invention. As shown in Table 10, the performance comparison results of the method proposed in the present invention and various classic convolutional neural networks and the same type of crop disease image classification models on the FGVC8 public dataset are shown.

[0195] Table 10 Comparison of the performance of different deep learning models and the same type of crop disease image classification models in the FGVC8 dataset

[0196] Models Accuracy Recall Precision F1 GFLOPs Parameters GoogLeNet 95.31 95.33 95.56 95.39 1.60G 7.01M InceptionResNetV2 94.92 94.93 95.40 95.06 6.50G 55.84M EfficientNetB0 91.80 91.88 92.13 91.92 0.42G 8.43M ShuffleNetV2 91.80 91.96 91.89 91.86 0.04G 1.31M MobileVit 94.53 94.54 95.05 94.63 0.27G 1.27M Re-GoogLeNet 96.48 96.50 96.58 96.49 2.80G 9.11M ALS-Net 92.97 93.03 93.55 93.08 0.73G 1.18M LBMRNet 91.02 91.13 91.75 91.30 0.17G 0.91M LCAMNet 95.31 95.29 95.81 95.39 0.03G 1.30M

[0197] The experimental results show that the recognition accuracy of the method proposed in the present invention on the public dataset is comparable to that of GoogLeNet and InceptionResNetV2, and at the same time, it is significantly lower than these two models in terms of the number of floating-point operations (FLOPs) and the number of model parameters.

[0198] The method proposed by the present invention not only can be comparable to GoogLeNet and InceptionResNetV2 in terms of recognition accuracy, but also has more obvious advantages in terms of computational efficiency, which helps to reduce the consumption of computing resources. In addition, the recognition accuracy of the method proposed by the present invention on the public dataset is significantly higher than that of ShuffleNetV2, and it is slightly lower than ShuffleNetV2 in terms of the number of floating-point operations and the number of parameters, demonstrating that LCAMNet better integrates technologies such as multi-branch feature extraction, multi-scale downsampling module and attention mechanism in ShuffleNetV2, making it applicable to different datasets.

[0199] Furthermore, comparing the method proposed by the present invention with the same type of crop disease image classification models, the experimental results show that although the method proposed by the present invention is lower than Re-GoogLeNet in terms of recognition accuracy, it is much lower than Re-GoogLeNet in terms of the number of floating-point operations and the number of model parameters, reflecting a significant optimization in the consumption of computing resources. This optimization enables the present invention to greatly reduce the computational cost while maintaining a high recognition accuracy, making it more suitable for application scenarios with limited computing resources, such as mobile devices or embedded systems.

[0200] In summary, the experimental results fully verify the effectiveness of the method proposed by the present invention on the public dataset. Compared with some deep learning networks, the present invention significantly reduces the computational complexity while maintaining a competitive recognition accuracy, showing strong practical application value and promotion potential, as shown in Table 10. Through three-dimensional visualization analysis of accuracy, GFLOPs and the number of parameters on FGVC8 as Figure 12 shown, it further verifies the universality and generality of LCAMNet, that is, it has low GFLOPs and the number of parameters, but the accuracy and other indicators are only second to Re-GoogLeNet and GoogLeNet.

[0201] Therefore, the present invention adopts the above-mentioned multi-classification method for apple leaf diseases under complex environmental backgrounds. The multi-scale downsampling module (MSDM) performs different downsampling operations on different channels and further fuses the features between different channels by using channel shuffling, effectively avoiding information loss caused by a single downsampling operation, thereby improving the recognition accuracy of the network. The multi-dimensional feature extraction network (MFEN) captures diverse feature information of apple leaf diseases through a multi-branch structure. By introducing a channel separation and shuffling mechanism into MFEN, the feature communication between channels is promoted, further enhancing the feature extraction ability. An improved triple attention module is incorporated into MFEN, and two cascaded 3×3 convolutions are used to replace the original 7×7 convolution, removing the original normalization operation. While enhancing the linear expression ability of the network, the number of parameters is reduced, the focusing ability of the model on key disease regions is enhanced, and the accuracy and robustness of disease recognition are improved. The LCAMNet of the present invention has both good generality and high efficiency, and is applicable to the apple leaf disease classification task with limited resources and complex backgrounds.

[0202] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that they can still modify or equivalently replace the technical solutions of the present invention, and these modifications or equivalent replacements cannot make the modified technical solutions deviate from the spirit and scope of the technical solutions of the present invention.

Claims

1. A multi-classification method for apple leaf diseases based on a complex environmental background, characterized in that, It includes the following steps: Step S1: Construct the public dataset FGVC8 and the self-built complex environment background dataset SCEBD, and preprocess the collected image data; Step S2: Prove that two cascaded 3×3 convolutions are equivalent to a 5×5 convolution, and further compare their computational amounts; Step S3: Construct a multi-scale downsampling module MSDM by combining multiple downsampling strategies; Step S4: Construct a multi-scale feature extraction module MFEN to capture diverse feature information of apple leaf diseases through a multi-branch structure; Step S5: Introduce an improved triple attention mechanism into MFEN to further extract key feature information; Step S6: Based on the above constructed MSDM module, MFEN module and the improved triple attention mechanism, establish a lightweight fusion attention multi-branch network LCAMNet model.

2. The multi-classification method for apple leaf diseases based on a complex environmental background according to claim 1, wherein In step S3, by combining multiple downsampling strategies, a multi-scale downsampling module MSDM is constructed. The MSDM module contains two branches. The first branch includes two paths. By applying max pooling (i.e., path 1) and average pooling (path 2) in parallel, the global features and salient features in the image are extracted respectively; First, average pooling extracts the global features of the image by averaging the pixel values within the pooling region, reducing the influence of local noise and enabling the model to capture the overall information of the image; Second, max pooling extracts local salient features by selecting the maximum value in the pooling window, highlighting the texture and morphological changes in the disease area; Then, by concatenating the results of max pooling and average pooling, the complementarity between salient features and global features is achieved; Finally, 1×1 convolution, batch normalization BN, and ReLU activation function are used to aggregate the channels.

3. A multi-classification method for apple leaf diseases based on a complex environmental background according to claim 2, characterized in that, In step S3, the second branch includes path 2 and path 3, which are parallel convolutional downsampling modules; In path 3, the input feature map first passes through a 3×3 depth convolution with a stride of 2 + BN to retain and downsample the overall features; then passes through a 1×1 conv + BN + Relu to adjust the number of channels to ensure that the channels of the output feature map match those of path 1 and path 2 of the first branch; In path 4, the input feature map first passes through a 1×1 conv + BN + Relu, then passes through a 3×3 depth convolution with a stride of 2 + BN, and finally passes through a 1×1 conv + BN + Relu to extract local features; by concatenating the results of path 3 and path 4, the overall features and local features are fused; The downsampling outputs of the first branch and the second branch are concatenated in the channel dimension to integrate the global and local features extracted by the two branches. The concatenated feature map is further optimized through channel shuffle operations to enhance the interaction between channels and promote information fusion.

4. A multi-classification method for apple leaf diseases based on a complex environmental background according to claim 1, characterized in that, In step S4, a multi-scale feature extraction module MFEN is constructed. By evenly distributing the number of input channels into four independent branches, diverse feature information of apple leaf diseases is captured. Among them, the four independent branches are: a feature retention branch, a local detail branch, a deep feature branch, and a significant feature branch.

5. A multi-classification method for apple leaf diseases based on a complex environmental background according to claim 4, characterized in that, In the feature retention branch, the input features obtain output features through skip connections, maintaining the low-level information of apple leaf diseases.

6. The multi-classification method for apple leaf diseases based on a complex environmental background according to claim 5, characterized in that, The local detail branch is used to extract the lesion edges, local texture changes, and disease regions of apple leaf disease pictures. In the local detail branch, the feature map input to the second branch first undergoes a 1×1 convolution to adjust the number of channels of the feature map. Then the feature map passes through a BN layer and a Relu layer. Subsequently, the feature map undergoes a 3×3 depth convolution, BN. The depth convolution performs convolution independently on each input channel, extracting local features on each channel, which is suitable for extracting fine-grained and independent feature information. Finally, the feature map undergoes a 1×1 point convolution to fuse the local features extracted from each channel in the depth convolution. By weighted combination of features from different channels, the model captures global information across channels.

7. A multi-classification method for apple leaf diseases based on a complex environmental background according to claim 6, characterized in that The deep feature branch extracts multi-level representations of apple leaf disease pictures. The feature map input to the third branch first undergoes a 1×1 convolution, BN, Relu. Then it undergoes two 3×3 depth convolutions, BN. Finally, it undergoes a 1×1 convolution, BN, Relu. Two cascaded 3×3 depth convolutions are used in this branch to reduce network parameters and computational complexity.

8. A multi-classification method for apple leaf diseases based on a complex environmental background according to claim 7, characterized in that, In the significant feature branch, the feature map input to the fourth branch first extracts features through a max pooling layer. Then, the feature map is further processed through a 1×1 convolution to ensure the fusion of features from different channels and enhance the overall feature expression ability. A concatenation operation is performed on the output feature maps of the above four branches, and channel shuffle is introduced to optimize the information flow and fusion between different channels. The channel shuffle operation realizes information exchange of features from different convolutional channels by rearranging the features of each channel.

9. A multi-classification method for apple leaf diseases based on a complex environmental background according to claim 1, characterized in that, In step S5, the triplet attention mechanism includes three branches. By the first two branches, the interaction information between channels and between spaces is captured respectively. The third branch is used to establish separate spatial attention. The outputs of these three branches are fused to generate the final attention features. Initial tensor X C×H×W , where C is the number of image channels, H is the image height, and W is the image width; for the first branch, the input shape remains C×H×W, which is responsible for aggregating the interaction features of H and W, i.e., the spatial dimensions; for the second branch, the shape of the tensor obtained through rotation is H×C×W, which is responsible for aggregating the interaction features of C and W, i.e., the channel and W spatial dimensions; for the third branch, the shape of the tensor obtained through rotation is W×H×C, which is responsible for aggregating H and C, i.e., the interaction features of the H spatial dimension and the channel; After that, the triplet tensor is passed to Z-pool. The Z-pool layer is responsible for reducing the zeroth dimension of the tensor to 2 by concatenating the average pooling and max pooling features in that dimension. Then, two cascaded 3×3 convolutional layers, Relu, are used to reduce the zeroth dimension of the tensor to 1. The attention weights are generated by a sigmoid activation layer and then applied to the input tensor, and the rotated tensor is restored to the original input shape. The final attention mechanism is obtained by averaging and summing the results of the three branches.

10. A multi-classification method for apple leaf diseases based on a complex environmental background according to claim 1, characterized in that, In step S6, a lightweight fusion attention multi-branch network LCAMNet model is established. The LCAMNet model includes a basic 3×3 convolution Conv1, a max-pooling layer with a 3×3 pooling kernel, and stages 2, 3, and 4, as well as a 1×1 convolution Conv5, an average pooling layer, and a fully connected layer; Among them, stages 2, 3, and 4 all contain a multi-scale downsampling module MSDM and a multi-scale feature extraction module MFEN; an improved triple attention mechanism is stacked after the MFEN in stage 4; For multi-channel input images, the LCAMNet model first preliminarily processes the input images through a Conv1 layer with a size of 3×3 and a stride of 2 to extract low-level features; then, a max-pooling layer with a size of 3×3 and a stride of 2 is used to downsample the images to reduce the dimension of the output features and obtain the output feature map; Then, the feature map passes through stages 2, 3, and 4 to compress redundant information and gradually extract the features of apple leaf disease images; among them, each stage includes a cascaded multi-scale downsampling module MSDM and a multi-scale feature extraction module MFEN; the MSDM module is used to downsample the extracted features to improve the computational efficiency while retaining the key features; on this basis, the MFEN module is stacked to further capture the multi-scale features of the disease; The MSDM and MFEN modules are stacked three times in total, and by introducing an improved triple attention mechanism at the end of stage 4, key feature information is further extracted to enhance the expressive power of the model; Finally, a 1×1 convolutional layer Conv5 is connected to adjust the number of channels for feature fusion, and then it is processed through a 7×7 average pooling layer and a fully connected layer classifier for classification to output the final class result; The training process of the LCAMNet model is as follows: First, data preprocessing is performed, including data cleaning, image size adjustment, dataset division, data augmentation, and data normalization; Next, the hyperparameters in the training process are set, specifically including the number of training epochs Epoch, the batch size Batch size, the optimizer, and the loss function, and the model weights are initialized; Then, the data is fed into the model in batches for training; During the training process, each time a batch of data is sent to the training set, the Batch size is incremented by 1; if the Batch size reaches the maximum value, the Epoch is incremented by 1, and the next batch of data is continued to be input for training; if the Epoch reaches the set maximum value, the training ends, and the test set is used to evaluate the final performance of the model; if the Epoch does not reach the maximum value, the data is continued to be grouped according to the set batch size and fed into the model for continued training.

Citation Information

Patent Citations

  • Plant leaf disease classification method based on artificial intelligence

    CN116486162A

  • Apple leaf disease segmentation and grading system based on attention and feature fusion

    CN118334053A

  • Classification method and device for apple leaf diseases, terminal, medium and program product

    CN119600370A

  • Overhead transmission line defect detection method based on multi-scale feature fusion

    CN119624922A

  • Defect detection method and related device

    WO2024255919A1

Cited By

  • Macadamia nut disease detection and identification method based on image identification

    CN121482468A

  • An image recognition-based macadamia disease detection and identification method

    CN121482468B