Method and system for constructing cross-scale large-kernel convolution corn leaf disease segmentation model based on attention coordination mechanism

By designing a lightweight LKCAFormer network, combined with large-core convolution and coordinated attention mechanism, the problems of insufficient accuracy and high calculation costs in corn leaf disease segmentation are solved, and efficient and accurate disease segmentation is achieved, suitable for resource-constrained equipment.

CN120014275AActive Publication Date: 2025-05-16INNER MONGOLIA AGRICULTURAL UNIVERSITY

Patent Information

Application Number
CN202510103002.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-16
Estimated Expiration
2045-01-22

AI Technical Summary

Technical Problem

The prior art has problems in corn leaf disease segmentation, such as insufficient accuracy, high computational cost and difficulty in deploying on resource-constrained devices, especially in the treatment of complex backgrounds and small lesions.

Method used

A lightweight cross-scale large-core convolution segmentation network LKCAFormer based on coordinated attention mechanism is designed. By combining a multi-level large-core convolution and attention mechanism encoder and a cross-scale attention decoder, it realizes precise segmentation of corn leaf diseases.

Benefits of technology

The IoU in the diseased area was improved by an average of 1.34% compared with the traditional method, while reducing model parameters and improving inference speed. It is suitable for complex backgrounds and small lesions, solving the deployment problem on resource-constrained devices.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014275A_ABST
    Figure CN120014275A_ABST
Patent Text Reader

Abstract

The invention belongs to the field of intelligent agriculture, and discloses a method for constructing a cross-scale large-kernel convolution corn leaf disease segmentation model based on an attention coordination mechanism, the network constructs a large-kernel convolution attention coordination module (LK-COA), global modeling is carried out by using large-kernel convolution to obtain global features, and the global features are obtained by using the LK-COA. And then more tiny spots are extracted through COA attention, and scab segmentation is focused. The module has the capability of obtaining local edge features and detail features through superposition and aggregation, meanwhile, extraction of small spots is enhanced, and the spot adhesion phenomenon is relieved. The CSDEcoder decoder is designed to be used for fusing shallow features rich in detail and edge information and deep features with strong semantic features, then accurate recovery is carried out, and finally a fine segmentation result is output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to but is not limited to the field of smart agriculture, and in particular relates to a method and system for constructing a cross-scale large kernel convolution corn leaf disease segmentation model based on a coordinated attention mechanism. Background Art

[0002] Corn is one of the important food crops in my country and an important raw material for animal husbandry and light industry. However, affected by climate change and environmental factors, the frequency of corn leaf diseases has increased year by year. Diseases on the leaves not only damage photosynthesis and affect the growth of corn plants, but also threaten the quality and yield of corn, causing serious economic losses to growers. Timely and accurate detection and diagnosis of plant diseases are essential for effective disease management. Traditional methods for diagnosing corn leaf diseases rely on manual observation of symptoms and spots, combined with the experience of experts. However, in large-scale farms, manual diagnosis has low efficiency, poor accuracy, high labor intensity and high difficulty in operation. Therefore, automated analysis of corn leaf diseases with the help of computer vision technology can not only improve the efficiency of disease diagnosis, but also help growers to prevent and control diseases more accurately, thereby significantly improving corn yield and quality. For field corn leaf diseases, the background is complex, the disease area is small and the texture is rich, and the disease symptoms are similar, which seriously affects the accuracy of disease segmentation.

[0003] As one of the core architectures of deep learning, convolutional neural network (CNN) has undergone significant evolution and has become a general framework in the agricultural field. Subsequently, CNN-based segmentation networks began to emerge, such as U-Net (Ronneberger et al., 2015), PSPNet (Zhao et al., 2017), SegNet (Sun et al., 2019) and different versions of DeepLab (Chen et al., 2014, 2017, 2019). More and more researchers have used the strong transfer ability and high accuracy of these models to analyze crop leaves and diseases, and achieved good results. Among them, DeepLab V3+ in the DeepLab series is suitable for accurate segmentation in complex scenes with its powerful context information processing capabilities and multi-scale feature extraction, especially in large-scale images with rich background information. U-net is particularly effective in tasks that require accurate segmentation (such as small disease spots and fine structures) due to its high-fidelity restoration of details and excellent performance on small sample data, and has low computational overhead, making it suitable for resource-constrained environments. In recent years, some improved networks have combined the advantages of DeepLab and U-net, which can not only process multi-scale contextual information, but also grasp accurate detail recovery and boundary accuracy, and have efficient computing and low resource consumption. Among them, Divyanth et al. (2023) collected 1050 corn lesion leaves from the Agricultural Research and Education Center of Purdue University in the United States, investigated the advantages and disadvantages of SegNet, U-Net and DeepLabv 3+, and finally selected U-Net and DeepLabv 3+ for the segmentation of corn leaves and lesions respectively. Yong Yang et al. (2023) combined the advantages of U-Net and proposed an improved Deeplabv3+ model to extract multi-scale semantic information in the encoding stage and obtain richer spatial information in the decoding stage, making the segmentation accuracy higher. However, CNN has great limitations in processing long-distance feature dependencies and spatial transformation relationships. The extraction of global features depends on the number of stacked network layers. But the deeper the number of layers, the more serious the network degradation phenomenon. Therefore, it is difficult for CNN-based segmentation networks to strike a balance between network complexity and accuracy.

[0004] In order to capture global features in depth and better extract local features at the same time, some researchers began to apply self-attention to replace convolution to achieve global feature extraction, and proposed the Transformer model. ViT (Dosovitskiy et al., 2010) first adopted the Transformer in the field of computer vision, which abandoned the use of convolutional layers and adopted a pure attention mechanism. Other scholars have expanded the field of image segmentation based on this structure, such as the semantic segmentation model SegFormer (Xie et al., 2021) and PoolFormer (Yu et al., 2022). These models verify that the improved Transformer has superior performance to CNN-based models in segmentation tasks. Transformer can explicitly model global contextual information when providing high-resolution (HR) natural images with complex backgrounds. However, cross-resolution information transfer is not considered, resulting in the inability to generate high-quality segmented images, resulting in poor segmentation performance of the network. In fact, the global feature map is able to obtain more granular information while containing less semantic details, especially for the edges of corn leaf diseases. Local features usually contain stronger semantic representation information, especially for small target disease areas that are difficult to segment. Therefore, maintaining both global and local features plays an important role in building a maize leaf disease segmentation model. Summary of the invention

[0005] In view of the problems existing in the prior art, the present invention provides a lightweight cross-scale large-kernel convolutional segmentation network LKCAFormer based on a coordinated attention mechanism for accurate segmentation of corn leaf diseases.

[0006] The present invention is implemented as follows: a lightweight cross-scale large kernel convolutional segmentation network LKCAFormer based on a coordinated attention mechanism for accurate segmentation of corn leaf diseases, which consists of two main parts:

[0007] (1) An encoder LK-COAT with powerful feature extraction capability based on multi-layer large-kernel convolutional CNN and attention modules; the encoder consists of three layers of LK-COA modules. The convolution operation of each layer with a super-large convolution kernel provides global features. Then, using skip connections, the upper-layer features and the obtained global features are passed to the newly designed coordinated attention block to optimize detailed features and feature aggregation, so as to achieve coarse feature representation and fine feature representation at the same feature scale; the encoder provides both shallow features with rich local details and edge information and deep features with rich global semantics, which can accurately extract the features of leaves and lesions during downsampling;

[0008] (2) Cross-scale attention decoder CSDecoder: Three decoders are designed to receive the low-resolution feature maps containing high-level semantic information transmitted by the encoder, calculate the similarity weights of fine-grained high-frequency global information, and perform attention calculations with the coarse-grained low-frequency feature maps of different scales output by the upper decoder, so as to obtain high-frequency feature maps that are a fusion of fine-grained and coarse-grained features. The obtained feature maps are then upsampled, concatenated, and input into the MLP for nonlinear processing to compensate for the lack of sensitivity of the convolution operation to capturing information. The shallow features are then fused with the deep features, and the aggregated feature maps are directly fed forward to the lightweight segmentation head.

[0009] Furthermore, the LKCAFormer specifically includes:

[0010] The leaf image input to the network is of size 512×512×3; in the encoder, the input image first passes through a feature extraction head consisting of two stacked 3×3 depthwise convolutions for effective feature extraction, which outputs a feature map of 1 / 4 the size of the original image, denoted as F0; this reduces the parameter count of the initial input encoder. Then, feature maps F1, F2, F3 of the original image {1 / 8, 1 / 16, 1 / 32} are obtained through a three-layer encoder. The ultra-large convolution kernel for each feature extraction is k = {(7,9,11), (11,13,15), (15,17,19)}, and the channel dimension is {32,64,128,160}; in the decoder stage, the feature maps F1, F2, F3 obtained from each layer of the encoder are passed to the decoder network for feature fusion and upsampled to the F0 space size, and feature concatenated with F0. Finally, a simple segmentation head module is used to output a 512×512×Ncls segmentation result; Ncls represents the number of pre-designed categories. The encoder includes a feature extraction head, three sets of large-kernel convolution operations and a collaborative attention module.

[0011] Furthermore, the encoder is an LK-COAT encoder, which models global information through large kernel convolution operations, uses a collaborative attention mechanism to capture local features, allows the model to learn the interaction between local and global features, thereby obtaining richer feature representations, and enhances the extraction of edge texture features to achieve fine-grained disease area segmentation of corn leaves;

[0012] Given a feature map F∈RC×H×W, where C is the number of input channels, H and W represent the height and width of the feature map, respectively; to alleviate the high computational cost of depthwise convolution under larger kernel size, the depthwise convolution with large kernel is decomposed into depthwise convolution with small kernel, followed by dilated depthwise convolution with fairly large kernel; the output of the LK-COA module can be obtained by using Equations 1-5;

[0013]

[0014] The symbol * represents the convolution operation, ⊙ represents the Hadamard product; Z in formula (1) C The output feature map is obtained by applying a deep convolution operation W with a convolution kernel of (2d-1)x(2d-1) (d represents the dilation rate) to the input feature F. The dilated convolution is used here to capture the detailed information of corn leaf diseases while compensating for the convolution kernel in formula (2). Grid effect caused by depthwise separable convolution; note Indicates the removal operation; outputs the global spatial information of eliminating background noise with Z C Represents; Keeping the convolution kernel k value below 23 can effectively capture global and local information; when the convolution kernel is greater than 23, it is proved by the paper (Lau et al., 2024) that it will produce high computational complexity and memory usage; the global output feature in formula (3) After average pooling and activation function, the attention weight of the leaf information is obtained and combined with the input feature map F C Perform Hadamard product to obtain the global attention feature map The input feature map F in formula (4) C After the deep separation convolution W with a convolution kernel of 3x3 and the activation function, the attention weight of the perceived disease area is obtained and combined with the global output feature Perform Hadamard product to obtain local attention feature map Finally, formula (5) is completed and By superposition, we finally get a feature map that eliminates background noise, contains high-frequency features of leaf edges and high-frequency features of disease-dense spots, and focuses on global and local details. Such a coding structure can effectively extract global leaf and spot features in position, space, and channel dimensions, enhance the model's ability to represent global features, and reduce the amount of calculation. The designed LK-COAT encoder consists of three LK-COA modules, where the convolution kernels are k=7, 11, 23, and d=1, 2, 3.

[0015] Furthermore, the decoder specifically comprises:

[0016] A network based on an encoder-decoder architecture is built. Therefore, after obtaining the encoder features Afterwards, three CSDecoder decoders are deployed to gradually integrate high-level semantic features and low-level spatial details; for the i-th decoder block, the input contains the encoder features Fi at the same level, the decoder features Fi from the previous decoder block The entire decoder process can be defined as follows

[0017]

[0018] F cls =up(f seg (Cat(F S +F0))) (8)

[0019] In formula (6) represents the i-th decoder feature, f AM is the AM attention module. The feature map is passed into the up of formula (7) f The operation is to upsample to F0 through a bilinear interpolation algorithm, and concatenate multiple feature maps after upsampling, and output feature F through a nonlinear feedforward network. D .

[0020] Furthermore, the cross-scale attention specifically includes:

[0021] The AM attention mechanism is used to optimize the leaf and spot edge segmentation and extract more microscopic spots. The similarity score matrix is ​​calculated in the AM module using formula (9). Specifically, given the input label X∈RH×W×C, the output Z is calculated using a deep convolution with a kernel size of k×k and a Hadamard product, as shown below:

[0022] S=A⊙V(9)

[0023] A=L1F i (10)

[0024] V=L2F i+1 (11)

[0025] Z=W 3×3 (S)+F i+1 (12)

[0026] Where ☉ is the Hadamard product, L1 and L2 are the weight matrices of the two linear layers, and W 3×3 Represents a depthwise convolution with a kernel size of k×k; the above operation enables each spatial position (h, w) to be associated with all pixels in a k×k square area centered at (h, w); information interaction between channels can be achieved through a linear layer; the output of each spatial position is the weighted sum of all pixels in the square area; compared with self-attention, using convolution to establish relationships is more memory-efficient than self-attention, especially when processing high-resolution images.

[0027] Furthermore, the LKCAFormer also includes: the loss function is a method combining cross entropy (CE) and Dice loss, and the specific formula is as follows:

[0028]

[0029] Loss total =0.5*Loss CE +Loss Dice (15)

[0030] The loss function shown in formula (15) combines CE and Dice losses; in the CE loss formula (12), y c represents the true label of the sample in category c, The model outputs the predicted probability for category c. In the Dice loss formula (15), x i It is the probability value that the first element in the prediction graph belongs to a certain type of prospect, y i is the true value of the first element in the label map; Dice is different from CE loss and is not affected by the size of the foreground. CE loss guides Dice loss in network learning; therefore, it is more reasonable to combine these two losses for network learning.

[0031] The present invention also provides a lightweight segmentation system for accurate segmentation of corn leaf diseases, comprising:

[0032] (1) An encoder module based on multi-level large kernel convolution and attention mechanism is used to extract global features and fine-grained features of the input corn leaf image;

[0033] (2) A cross-scale attention decoder module is used to fuse the high-level semantic features and low-level spatial detail features extracted by the encoder and upsample the fused features to the input image size to generate the disease segmentation result;

[0034] (3) A data processing module for model training and inference, which preprocesses the input image to adapt it to the network input and post-processes the output result to optimize the segmentation effect.

[0035] Furthermore, the encoder module uses large kernel convolution operations to globally model the input image and combines it with a collaborative attention mechanism to capture local features in spatial and channel dimensions.

[0036] Furthermore, the decoder module utilizes a cross-scale attention mechanism to interactively fuse high-level semantic features and low-level spatial features, and restores the original resolution of the image by progressive upsampling to generate an accurate segmentation map of the diseased area.

[0037] In combination with the above technical solutions and the technical problems solved, the advantages and positive effects of the technical solutions to be protected by the present invention are as follows:

[0038] First, the present invention designs a lightweight LKCAFormer network for accurate segmentation of corn leaf diseases in the field. In the encoding stage, large kernel convolution and collaborative attention mechanism are combined to capture global features through large kernel convolution, and the collaborative attention mechanism enhances detail extraction. In the decoding stage, the cross-scale attention mechanism is used to fuse features, accurately restore boundaries and details, and effectively improve segmentation accuracy. It can improve the IoU of the diseased area by an average of 1.34% compared with traditional methods. In addition, the model parameters are reduced by 36.7% compared with the classic method Deeplab v3+, and the inference speed is increased by 8.4 times. The robustness of two different data sets, CD&S and Single-CD&S, has been verified, and it is particularly suitable for complex backgrounds and small diseased areas. The proposed solution solves the problems of insufficient accuracy of existing methods in dealing with complex backgrounds or highly similar diseased areas, the limitations of traditional convolutional networks in global feature modeling and small spot segmentation, and the high computational cost and difficulty in deployment on resource-constrained devices, providing an efficient and accurate solution for intelligent agricultural disease detection.

[0039] Second, as auxiliary evidence of the inventiveness of the claims of the present invention, it is also reflected in the following important aspects:

[0040] (1) The technical solution of the present invention solves a technical problem that people have been eager to solve but have never been able to solve successfully:

[0041] In large-scale farms, manual diagnosis of corn leaves is inefficient, inaccurate, labor-intensive, and difficult to operate. Therefore, using computer vision technology to automatically analyze corn leaf diseases can not only improve the efficiency of disease diagnosis, but also help growers prevent and control diseases more accurately. However, traditional algorithms are currently unable to run on edge computing devices. This invention solves the problem of running on edge computing devices through a lightweight model design idea without losing segmentation accuracy, demonstrating technical innovation and practical value.

[0042] (2) The technical solution of the present invention overcomes technical prejudice:

[0043] Existing technologies perform poorly in field environments, and have large errors in the accuracy of segmenting high-similarity diseases and leaves in complex backgrounds. This technical solution overcomes these shortcomings by extracting leaf edges and small disease spots as much as possible through large kernel convolution and collaborative attention mechanism. Traditional models rely too much on high-performance hardware and cannot run efficiently on resource-constrained devices. This solution overcomes this limitation by reducing computing costs through lightweight design. BRIEF DESCRIPTION OF THE DRAWINGS

[0044] Figure 1The overall architecture of LKCAFormer provided by the embodiment of the present invention is as follows: (a) is the overall architecture of LKCAFormer; (b) is the internal structure of LK-COAT, which consists of three LK-COA modules and adds skip connections; (c) is the internal structure of CSDecoder, which integrates features across scales through the attention mechanism;

[0045] Figure 2 is a marking visualization result provided by an embodiment of the present invention;

[0046] Figure 3 is the segmentation result of each method in the gls diseased single leaf test set in the single-CD&S dataset provided by the embodiment of the present invention;

[0047] Figure 4 is the segmentation result of each method under the nls diseased single leaf test set in the single-CD&S dataset provided by the embodiment of the present invention;

[0048] Figure 5 is the segmentation result of each method under the nlb diseased single leaf test set in the single-CD&S dataset provided by the embodiment of the present invention;

[0049] Figure 6 is the segmentation result of each method under the gls disease multi-leaf test set in the CD&S data set provided by the embodiment of the present invention;

[0050] Figure 7 is the segmentation result of each method under the nls disease multi-leaf test set in the CD&S dataset provided by the embodiment of the present invention;

[0051] Figure 8 is the segmentation result of each method under the nlb disease multi-leaf test set in the CD&S dataset provided by the embodiment of the present invention;

[0052] Fig. 9 It is the result of leaf and spot segmentation for each test provided by the embodiment of the present invention. DETAILED DESCRIPTION

[0053] In order to make the purpose, technical solution and advantages of the present invention more clearly understood, the present invention is further described in detail below in conjunction with the embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0054] The embodiment of the present invention provides a lightweight cross-scale large kernel convolutional segmentation network LKCAFormer based on a coordinated attention mechanism for accurate segmentation of corn leaf diseases. The LKCAFormer consists of two main parts:

[0055] (1) An encoder LK-COAT with powerful feature extraction capability based on multi-layer large-kernel convolutional CNN and attention modules; the encoder consists of three layers of LK-COA modules. The convolution operation of each layer with a super-large convolution kernel provides global features. Then, using skip connections, the upper-layer features and the obtained global features are passed to the newly designed coordinated attention block to optimize detailed features and feature aggregation, so as to achieve coarse feature representation and fine feature representation at the same feature scale; the encoder provides both shallow features with rich local details and edge information and deep features with rich global semantics, which can accurately extract the features of leaves and lesions during downsampling;

[0056] (2) Cross-scale attention decoder CSDecoder: Three decoders are designed to receive the low-resolution feature maps containing high-level semantic information transmitted by the encoder, calculate the similarity weights of fine-grained high-frequency global information, and perform attention calculations with the coarse-grained low-frequency feature maps of different scales output by the upper decoder, so as to obtain high-frequency feature maps that are a fusion of fine-grained and coarse-grained features. The obtained feature maps are then upsampled and concatenated, and input into a multi-layer perceptron (MLP) for nonlinear processing to compensate for the lack of sensitivity of the convolution operation to capturing information. The shallow features are then fused with the deep features, and the aggregated feature maps are directly fed forward to the lightweight segmentation head.

[0057] Specifically, the size of the leaf image input to the network is 512×512×3. In the encoder, the input image first passes through a feature extraction head consisting of two stacked 3×3 depthwise convolutions for effective feature extraction, which outputs a feature map of 1 / 4 the size of the original image, denoted as F0. This reduces the parameter count of the initial input encoder. Then, the feature maps F1, F2, F3 of the original image {1 / 8, 1 / 16, 1 / 32} are obtained through a three-layer encoder. The ultra-large convolution kernel for each feature extraction is k = {(7, 9, 11), (11, 13, 15), (15, 17, 19)}, and the channel dimension is {32, 64, 128, 160}. In the decoder stage, the feature maps F1, F2, F3 obtained from each layer of the encoder are passed to the decoder network for feature fusion and upsampled to the F0 spatial size, and feature concatenated with F0, and finally output 512×512×N through a simple segmentation head module. cls The segmentation result. cls Indicates the number of pre-designed categories, which is 3 in this invention. The encoder and decoder of the network will be explained in detail in the rest of this section. The encoder includes a feature extraction head, three sets of large kernel convolution operations and a collaborative attention module (introduced in Section 1), the attention decoder is given in Section 2, and the loss function is introduced in Section 3.

[0058] 1. Super large kernel convolution and coordinated attention

[0059] Due to the high computational complexity of Transformer and the need for significant computing resources, it is not suitable for real-time use in agricultural production sites. However, it is effective in capturing long-distance dependencies and global feature information, which is lacking in convolution operations. In recent years, some scholars have studied the use of ultra-large convolution kernels to improve the global information capture capability of convolutional networks, while using attention mechanisms to enhance channel and spatial feature perception and long-distance dependencies. This scheme can reduce the number of network layers, be more sensitive to local details, optimize parameters, and achieve better segmentation effects. Based on the above analysis, the present invention designs an LK-COAT encoder, which models global information through large-kernel convolution operations, uses a collaborative attention mechanism to capture local features, and allows the model to learn the interaction between local and global features, thereby obtaining richer feature representations and enhancing the extraction of edge texture features to achieve fine-grained disease area segmentation of corn leaves. The LK-COAT encoder is specifically shown in Figure (b). Given a feature map F∈R C ×H×W , where C is the number of input channels, H and W represent the height and width of the feature map, respectively. To alleviate the high computational cost of depthwise convolutions with larger kernel sizes, the depthwise convolution with large kernels is decomposed into a depthwise convolution with small kernels, followed by a dilated depthwise convolution with a fairly large kernel. The output of the LK-COA module can be obtained by using Eqs. 1-5.

[0060]

[0061]

[0062] The symbol * represents the convolution operation, and ⊙ represents the Hadamard product. C The output feature map is obtained by applying a deep convolution operation W with a convolution kernel of (2d-1)×(2d-1) (d represents the dilation rate) to the input feature F. The purpose of using dilated convolution here is to capture the detailed information of corn leaf diseases while compensating for the convolution kernel in formula (2). Grid effect caused by depthwise separable convolution. Note Indicates the removal operation. Output the global spatial information after removing the background noise with Z C Represents. Keeping the convolution kernel k value below 23 can effectively capture global and local information. When the convolution kernel is larger than 23, it is proved in the paper (Lau et al., 2024) that it will produce high computational complexity and memory usage. The global output feature in formula (3) After average pooling and activation function, the attention weight of the leaf information is obtained and combined with the input feature map F C Perform Hadamard product to obtain the global attention feature map The input feature map F in formula (4) C After the deep separation convolution W with a convolution kernel of 3x3 and the activation function, the attention weight of the perceived disease area is obtained and combined with the global output feature Perform Hadamard product to obtain local attention feature map Finally, formula (5) is completed and The final feature map is superimposed to eliminate background noise, including high-frequency features of leaf edges and high-frequency features of disease-intensive spots, and focuses on global and local details. Such a coding structure can effectively extract global leaf and spot features in position, space, and channel dimensions, enhance the model's ability to represent global features, and reduce the amount of calculation. The designed LK-COAT encoder consists of three LK-COA modules, where the convolution kernels are k=7, 11, 23, and d=1, 2, 3.

[0063] 2. Decoder

[0064] As mentioned before, a network based on an encoder-decoder architecture is built. Therefore, after obtaining the encoder features Afterwards, three CSDecoder decoders are deployed to gradually integrate high-level semantic features and low-level spatial details, such as Figure 1 For the i-th decoder block, the input contains the encoder features Fi at the same level, the decoder features from the previous decoder block The entire decoder process can be defined as follows

[0065]

[0066] F cls =up(f seg (Cat(F S +F0))) (8)

[0067] In formula (6) represents the i-th decoder feature, f AM is the AM attention module. The feature map is passed into the up of formula (7) f The operation is to upsample to F0 through a bilinear interpolation algorithm, and concatenate multiple feature maps after upsampling, and output feature F through a nonlinear feedforward network. D .

[0068] 2.1 Cross-Scale Attention

[0069] Due to the different illumination conditions of multiple leaves in the natural environment, the presence of leaf shadows will reduce the accuracy of leaf segmentation. In addition, the similarity between the edge color of the lesion and the leaf color and part of the background color makes it difficult to extract the true outline of the lesion. In addition, in most corn leaf images, the proportion of diseased pixels to the entire image pixels is very small, which makes the extraction of small disease features more difficult. In the encoding stage, the LK-COAT module provides global perception and local feature extraction capabilities, but there is a risk of losing the edge extraction of scattered spots or dense spot areas. Therefore, this section uses the AM attention mechanism to strengthen the optimization of leaf and spot edge segmentation and extract more microscopic spots. The AM module is shown in Figure (c). Formula (9) is used in the AM module to calculate the similarity score matrix. Specifically, given an input label X∈R H×W×C , the output Z is calculated using a depthwise convolution with a kernel size of k×k and a Hadamard product as follows:

[0070] S=A☉V(9)

[0071] A=L1F i (10)

[0072] V=L2F i+1 (11)

[0073] Z=W 3×3 (S)+F i+1 (12)

[0074] Where ☉ is the Hadamard product, L1 and L2 are the weight matrices of the two linear layers, and W 3×3 Represents a deep convolution with a kernel size of k×k. The above operation enables each spatial position (h, w) to be associated with all pixels in a k×k square area centered at (h, w). Information interaction between channels can be achieved through a linear layer. The output of each spatial position is the weighted sum of all pixels in the square area. Compared with self-attention, the model designed by the present invention uses convolution to establish relationships, especially when processing high-resolution images, it saves more memory than self-attention.

[0075] 3. Loss Function

[0076] The present invention adopts a method combining cross entropy (CE) and Dice loss. The specific formula is as follows:

[0077]

[0078] Loss total =0.5*Loss CE +Loss Dice (15)

[0079] The loss function shown in formula (15) combines CE and Dice losses. In the CE loss formula (12), y c represents the true label of the sample in category c, The model outputs the predicted probability for category c. In the Dice loss formula (15), x i It is the probability value that the first element in the prediction graph belongs to a certain type of prospect, y i is the true value of the first element in the label map. Dice is different from CE loss and is not affected by the size of the foreground. However, CE loss guides Dice loss in network learning. Therefore, it is more reasonable to combine these two losses for network learning.

[0080] The LKCAFormer network consists of two main modules: an encoder based on multi-level large kernel convolution and attention modules, and a cross-scale attention decoder. The encoder module extracts global features of the input corn leaf image through multi-level large kernel convolution, and combines the attention mechanism to capture key fine-grained features. The decoder module restores the spatial resolution of the image layer by layer, and uses the cross-scale attention mechanism to fuse the high-level semantic features extracted by the encoder and the low-level spatial detail features, thereby achieving accurate segmentation of corn leaf diseases.

[0081] The encoder uses large kernel convolution operations to model the features of the input image. Large kernel convolution captures the global context information of the image by expanding the receptive field of the convolution kernel while retaining local details. In order to enhance the correlation between features, a collaborative attention mechanism is introduced to simultaneously focus on the feature distribution in spatial and channel dimensions. This mechanism significantly optimizes the interaction between global and local features, enabling the network to more accurately identify diseased areas and ignore irrelevant background.

[0082] The decoder uses a cross-scale attention mechanism to gradually fuse high-level semantic features and low-level spatial detail features extracted in the encoder. At each decoding stage, high-level features are combined with low-level features through cross-scale interactions to restore detail information. The multi-step upsampling operation of the decoder restores the feature map to the same spatial resolution as the input image, generating accurate disease segmentation results, thereby ensuring the delicacy and accuracy of the segmentation boundaries.

[0083] By combining the global modeling capability of large kernel convolution and the feature optimization capability of collaborative attention, LKCAFormer can extract global semantic information while maintaining the integrity of local features. The feature fusion strategy of the cross-scale attention decoder effectively improves the spatial accuracy and semantic consistency of the segmentation results. This method takes into account both efficiency and accuracy on the basis of lightweight design. It is particularly suitable for the processing requirements of complex textures and boundary details in the corn leaf disease segmentation task, providing reliable technical support for crop disease diagnosis.

[0084] In order to verify the actual effect of the lightweight cross-scale large kernel convolution segmentation network (LKCAFormer) based on the coordinated attention mechanism on the field corn leaf disease segmentation, the public CD&S dataset was selected as the experimental data. The segmentation network of the present invention combines large kernel convolution with the coordinated attention mechanism to achieve accurate segmentation of complex diseased areas, solving the problem of insufficient segmentation accuracy of traditional methods when the diseased spot boundaries are blurred.

[0085] Relevant evidence of the technical effects achieved by the embodiments of the present invention.

[0086] 1.1 Dataset

[0087] Data collection

[0088] The proposed model is evaluated on two datasets, including Single-CD&S and CD&S (reference). CD&S is an open and fair corn disease recognition dataset, which contains three common corn leaf diseases: northern leaf blight, northern leaf spot, and gray leaf spot. This dataset collects images in a natural environment. In addition to the diseased leaves in the foreground, there are many diseased leaves in the background. The background is complex and there are interferences. In order to accurately segment the leaves and diseases in the image, some images are extracted from CD&S and only a single leaf and disease area in the image are annotated to construct the Single-CD&S dataset. The original CD&S dataset annotates multiple leaves and diseased areas.

[0089] The study consists of two consecutive parts: (1) The first stage: extracting the target leaf from the complex background. (2) The second stage is to segment the lesions based on the leaf images extracted in the first stage. Therefore, each original image requires leaf, disease, and background labels. The sample data is annotated using the labelme (Russell, Torralba, Murphy, & Freeman, 2008) (https: / / github.com / wkentaro / labelme) tool. The labeled visualization results are as follows Figure 2. In order to alleviate overfitting and improve the robustness and generalization ability of the model, the Augmentor module (Bloice, Roth, & Holzinger, 2019) was used to perform geometric transformations such as random left or right flipping, random cropping, random sampling, and color and brightness enhancement or reduction. In addition, powerful data augmentation methods in the semantic segmentation library MMsegmentation (Contributors, 2020) were also applied. The three corn leaf disease datasets were divided into training and test sets in a ratio of 8:2. In addition, during the training stage, each data was divided into training and validation sets in a ratio of 9:1 for cross-validation. The training set, validation set, and data augmentation details are given in Table 1.

[0090] Table 1 Details of the corn leaf disease dataset used

[0091]

[0092] 1.2 Experimental Setup

[0093] Experimental data details. The experiments are based on the public code library MMSSegmentation Contributors (2020) and Pytorch (Paszke et al., 2019). The model was trained on 2 NVIDIA GTX 4090 GPUs. During training, the images were randomly cropped to 512×512. The AdamW (Loshchilov & Hutter, 2018) optimizer was used for training, using the cos learning rate decay strategy, with the following hyperparameters: momentum 0.9, weight decay 1e-2, batch size 16, epoch 500, initial learning rate 1e-4, and minimum learning rate 1e-7. To prevent overfitting, the descent path rate was set to 0.1.

[0094] Evaluation Metrics. Quantitative metrics for this experiment include: Precision (Zhang et al., 2022), Intersection over Union (IoU) (Xie et al., 2021), Dice coefficient (Garcia-Garcia, Oltz-Escolano, Oprea, Verena-Martínez, & Garcia-Rodriguez, 2017), and Recall (Li et al., 2023). Among them, higher IoU and Dice values ​​generally indicate a higher degree of overlap between the predicted results and the true results, which indicates a more accurate segmentation result.

[0095]

[0096] Among them, TP represents positives that are classified as true positives. TN represents true negatives that are correctly classified. FP represents pixels that are classified as leaves but are actually background. FN represents pixels that are classified as background but are actually leaves.

[0097] 1.3 Comparative Experiment

[0098] In this subsection, the proposed method is compared with popular deep learning semantic segmentation methods, including CNN-based methods U-Net, DeepLab v3+, Transformer-based methods SegFormer, PVT2, lightweight methods TopFormer, AFFormer and SwiftFormer, to further verify the feasibility and effectiveness of LKCAFormer.

[0099] Specifically, U-Net is a simple skip-level connection structure that fuses shallow features and semantic features multiple times. DeepLab v3+ adopts the ASPP spatial pooling pyramid structure to expand the receptive field through hollow convolution. SegFormer uses a hierarchical Transformer block, and the decoder applies a lightweight MLP structure. TopFormer uses layer-by-layer refinement of features to effectively capture global context and avoid loss of details. AFFormer adopts a parallel architecture and uses prototype representation as a specific learnable local description to replace the decoder and retain rich image semantics on high-resolution features. SwiftFormer designs an efficient additive attention mechanism to learn consistent global context at multiple scales. Each method was trained and tested on three corn disease leaf image datasets. Among them, the performance of each method is measured using seven evaluation indicators: Dice, Recall, IoU, Precision, FPS, total parameters, and FLOPs / G. Tables 2, 3, and 4 record the comparison results of the three disease segmentation of the Single-CD&S dataset.

[0100] As can be seen from Table 2, the proposed method shows the best segmentation performance on the gls test set. The proposed method outperforms the CNN models U-net and DeepLab v3+ in terms of segmentation accuracy. Compared with U-net, the IoU of background, leaf and lesion segmentation is 1.14%, 0.6% and 3.15% higher, respectively. DeepLab v3+ is 0.92% lower than the proposed method in IoU of leaf segmentation, 2.54% lower in IoU of lesion segmentation, and 1.77% lower in IoU of background segmentation. Compared with the proposed method, SegFormer is 0.45% lower in IoU of leaf segmentation, 2.83% lower in IoU of disease segmentation, and 0.96% lower in IoU of background segmentation. Compared with PVT2, the proposed method is 0.58% higher in IoU of leaf segmentation, 1.79% higher in IoU of lesion segmentation, and 0.57% higher in IoU of background segmentation. Compared with the segmentation results of lightweight models TopFormer and AFFormer, the proposed methods have improved. In addition, the segmentation accuracy of SwiftFormer is close to and better than other methods, but not as good as the proposed method. Compared with SwiftFormer, the proposed method improves the IoU of leaf segmentation by 0.35%, the IoU of disease segmentation by 0.47%, and the IoU of background segmentation by 0.78%.

[0101] Table 2 Quantitative comparison of CNN-based and Transformer-based SOTA methods on the gls test set in the Single-CD&S dataset

[0102]

[0103] According to Table 3, the proposed method shows the best segmentation performance on the corn leaf nls disease test set. Compared with U-net, the IoU of the leaf and lesion segmentation of the proposed method is improved by 1.08% and 3.51%, respectively, and the IoU of the background segmentation is comparable, with a value of 0.03% increase. The segmentation accuracy of the background, leaf and disease of DeepLab v3+ is not as good as that of the proposed method. The IoU of the background segmentation is 1.24% lower than that of the proposed method, the IoU of the leaf segmentation is 1.74% lower than that of the proposed method, and the IoU of the lesion segmentation is 2.12% lower than that of the proposed method. In addition, the proposed method is 0.72% higher than SegFormer in the IoU of the disease segmentation. The segmentation performance of PVT2 is weaker than that of the proposed method, and its performance in the IoU of the background, leaf and lesion segmentation is 0.21%, 1.01% and 1.08% lower, respectively. The segmentation accuracy of SwiftFormer is higher than that of lightweight methods such as Topformer and AFFormer, but the segmentation accuracy of the proposed method is higher than SwiftFormer by 0.56% in IoU for background segmentation, 0.62% in IoU for leaf segmentation, and 0.95% in IoU for lesion segmentation. In summary, in terms of segmentation accuracy, SwiftFormer has higher accuracy and is worse than other methods, but not as good as the method in this study.

[0104] Table 4 shows the experimental results on the corn leaf nlb disease test set, from which it can be seen that the proposed method shows the best segmentation performance. Among them, Topformer has poor performance in leaf and disease segmentation IoU, which is 4.63% and 7.56% lower than the proposed method. AFFormer is slightly higher than the proposed method in background segmentation IoU, but the accuracy is 1.23% and 0.54% lower in leaf and disease segmentation IoU. Compared with the above methods, the segmentation effects of CNN-based U-net and DeeplabV3+ are similar, but both are lower than the proposed method. The Transformer-based Segformer method is similar to the proposed method in disease segmentation IoU, only 0.13% lower, but in leaf and background segmentation IoU, it is 1.88% and 1.11% lower. The performance of the PVT2 method in the IoU of background, leaf and disease segmentation is 0.61%, 1.61% and 0.78% lower than the proposed method, respectively. SwiftFormer has similar segmentation effects to U-net and DeeplabV3+. In summary, in terms of segmentation accuracy, the methods in this study have improved segmentation accuracy compared to other methods.

[0105] Table 3 Quantitative comparison of CNN-based and Transformer-based SOTA methods on the nls test set in the Single-CD&S dataset

[0106]

[0107] Table 4 Quantitative comparison of CNN-based and Transformer-based SOTA methods on the nlb test set in the Single-CD&S dataset

[0108]

[0109]

[0110] In order to better verify the performance of the proposed method in a real scene, each method was trained and tested on the CD&S dataset with complex background and multiple leaf diseases, and compared with the proposed method. Tables 5, 6 and 7 record the segmentation performance comparison of the proposed method and other methods on the three disease test sets. From Table 5 to Table 2, it can be seen that the performance of all methods in segmenting multiple leaves and diseased areas is worse than that of segmenting single leaves and diseased areas. The best segmentation performance in Table 5 is still the proposed method, which can reach 99.02%, 97.39% and 70.52% in the IoU of background, leaf and disease segmentation. This is 2.55%, 0.65% and 6.83% higher than PVT2, which has the worst disease segmentation performance. Among the lightweight methods, the performance of the Topformer method in the IoU of background, leaf and disease segmentation is 2.15%, 0.83% and 4.6% lower than that of the proposed method, respectively. The segmentation performance of AFFomer and SwiftFormer methods is similar, but on average 0.18%, 0.46% and 1.82% lower than that of the proposed method. Compared with DeepLab v3+, the IoU of the background and disease segmentation of this method is improved by 1.91% and 4.54%, respectively, and the IoU of leaf segmentation is comparable, with an increase of 0.04%. The segmentation accuracy of U-net background, leaves and diseases is not as good as that of the proposed method. The IoU of background segmentation is 1.24% lower than that of the proposed method, the IoU of leaf segmentation is 0.21% lower than that of the proposed method, and the IoU of disease segmentation is 6.19% lower than that of the proposed method.

[0111] Table 6 presents the segmentation performance comparison results of the proposed method and other methods on the nls test set. From the data, the overall segmentation effect is poor, but the proposed method still has the highest segmentation accuracy, reaching 98.02%, 95.19% and 67.28% in the background, leaf and disease segmentation IoU. The worst performer is the lightweight AFFormer, which reaches 91.37%, 87.96% and 59.11% in the background, leaf and disease segmentation IoU. The segmentation accuracy of Segformer background, leaf and disease is not as good as that of the proposed method. The IoU of background segmentation is 1.01% lower than that of the proposed method, the IoU of leaf segmentation is 2.88% lower than that of the proposed method, and the IoU of disease segmentation is 1.65% lower than that of the proposed method. In addition, the IoU of this method is 4.85% higher than that of U-net in disease segmentation, 2.01% higher in leaf segmentation, and 3.24% higher in background segmentation IoU. DeeplabV3+ reduces the IoU of disease segmentation by 5.37%, the IoU of leaf segmentation by 0.84%, and the IoU of background segmentation by 0.82%.

[0112] Table 5 Quantitative comparison of CNN-based and Transformer-based SOTA methods on the gls test set in the CD&S dataset

[0113]

[0114] Table 6 Quantitative comparison of CNN-based and Transformer-based SOTA methods on the nls test set in the CD&S dataset

[0115]

[0116] Table 7 shows the experimental results on the corn leaf nlb disease test set, from which it can be seen that the proposed method shows the best segmentation performance. Among them, the Transformer-based Segformer method is 6.62% lower than the proposed method in disease segmentation IoU, and 2.33% and 2.18% lower in leaf and background segmentation IoU. The performance of the proposed method in background, leaf and disease segmentation IoU is 2.74%, 1.44% and 9.27% ​​higher than the PVT2 method, respectively. The segmentation effects of U-net and DeeplabV3+ are similar, but DeeplabV3+ has better segmentation effects overall, which is 2.01%, 0.54% and 4.2% lower than the proposed method in background, leaf and disease segmentation IoU. The method is close to the lightweight Topformer segmentation performance, which is 3.05%, 1.81% and 2.62% higher in background, leaf and disease segmentation IoU. The segmentation performance of the AFFormer method is poor, with only 95.68%, 95.06% and 59.45% IoU for background, leaf and disease segmentation. In summary, in terms of segmentation accuracy, the methods in this study have improved segmentation accuracy compared with other methods.

[0117] Table 7 Quantitative comparison of CNN-based and Transformer-based SOTA methods on the nlb test set in the CD&S dataset

[0118]

[0119] Table 8 presents the comparison of the methods under the remaining evaluation indicators. As shown in Table 8, the proposed method is 7.36ms higher than U-net in terms of FPS. In addition, the total parameters and floating-point operations of the method are only 12.7% and 0.14% of U-net. PVT2 outperforms other methods in terms of FPS, but the number of parameters is more than twice that of the proposed method. AFFormer has the least number of parameters and floating-point operations, but the FPS is 0.94ms less than the proposed method. Compared with Topformer, the FPS of this method is improved by 6.4ms. In addition, in terms of total parameters and FLOPs, it is 1.46M and 1.05G less than Topformer, respectively. In contrast, through a comprehensive comparison of all parameters, the proposed method has the best segmentation performance and a smaller computational overhead. The algorithm achieves a good balance between segmentation accuracy and inference speed.

[0120] Table 8 Results of different methods on other evaluation indicators

[0121]

[0122] Figure 3 , Figure 4 , Figure 5The segmentation results of each method in the single-CD&S dataset are shown for the single leaf test set with gls, nls, and nlb diseases. Figure 3 As shown in the figure, the white dotted box marks the specific disease area where the color of the lesion is similar to that of the leaves due to the influence of lighting conditions. This area is the key to analyzing the disease segmentation effect. Comparing 2(a) and 2(c), U-net can correctly segment most diseases, but the segmentation is poor in specific areas and the detail information is obviously lost. However, by comparing 2(c) and 2(d), it can be found that the segmentation effect of deeplabv3+ is worse than that of U-net. In 2(e), the segmentation effect of Segformer is better than the previous two methods, but the disease segmentation effect is poor in some edge areas. It has strong global modeling ability and some local detail information is lost. Comparing 3(f), 3(h) and 3(i), AFFormer can remove other noises such as lighting to a certain extent and focus on disease segmentation, but its ability to segment disease edges is poor. Figure 3 The proposed method LKCAFormer in (j) can effectively compensate for the loss of fine-grained information caused by aggregating different resolutions and provide accurate segmentation performance in specific areas. At the same time, the segmentation effect of leaf edge diseases is good. Figure 4 This is the segmentation result of single leaf disease using nls. Since the characteristics of nls are that the disease spots are relatively concentrated and the disease areas are dispersed, the segmentation differences between all methods are small. Among them, the AFFormer method in 3(h) has poor segmentation of the disease area and only segments a small number of obvious disease spots. Methods 4(c), 4(d), 4(f) and 4(g) can basically segment dense disease spots, but their ability to segment the edges of disease spots is poor and there are incorrect segmentations. 4(e) has poor segmentation effect on independent disease spots on the edge. There are differences between 4(i) and 4(j) in the obvious stripes on the leaves. SwiftFormer mixes the stripe color with the disease spots, while the proposed method is more accurate in segmentation and has better segmentation effect on the disease spots on the edge of the leaves. Figure 5 The nlb single leaf disease segmentation results are presented. The white dotted box marks the specific area with dense lesions. Compared with 5(a) and 5(c), the U-net method has poor lesion segmentation effect in the specific area. Methods 5(d), 5(e), 5(f) and 5(i) have poor lesion segmentation in the specific area. At the same time, the above methods mistake the main veins of the leaves with similar color to the lesions as lesions, resulting in incorrect segmentation. Compared with 5(h) and 5(j), the proposed method can segment more dense lesions in specific areas, and has better effect on the edge segmentation of lesions in larger areas around the main veins. From the above experimental results, it can be seen that the LKCAFormer method can not only segment the edge area of ​​the leaf more clearly, but also segment the edge of the spots more accurately, and has better segmentation performance.

[0123] Figure 6 , Figure 7 , Figure 8 The segmentation results of each method in the CD&S dataset are shown for the gls, nls, and nlb diseased multi-leaf test sets. Figure 6 The gls disease shown in the figure is densely distributed, and affected by the light, the color of the reflected light from some leaves is similar to the color of the lesions, which is easy to be segmented incorrectly. Comparing 6(a) and 6(c), the segmentation of the dense lesion area is relatively complete, but there are leaves that are mistakenly segmented as lesions. At the same time, affected by the light, the segmentation effect of the lesions in the shadow part is poor. 6h) has the best segmentation effect for the lesions in the shadow part, but compared with 8(b), it is found that the marked lesions are not correctly segmented. 6(f) and 6(g) have poor segmentation effects on the edges of dense lesions, and there are incorrect segmentations. The method proposed in 6(j) has a good overall segmentation effect. For 6(b), most of the marked lesions can be correctly segmented and there are fewer incorrect segmentations. Figure 7 The results of nls segmentation of multiple leaf diseases in an environment where the background and leaf colors are similar are presented. Since there are grass and corn plants in the background, which are similar in color to the leaves, methods 7(c), 7(d), 7(g), 7(h) and 7(i) all have different degrees of incorrect segmentation in leaf segmentation. Other methods such as 7(e) and 7(f) fail to effectively segment leaves that are affected by light and are located at the edge, while the proposed method 7(j) correctly segments leaves compared to 7(b). In disease segmentation, since nls lesions are scattered and small, all methods have similar performance in segmenting relatively large lesions, but the proposed method 7(j) is relatively good at segmenting small lesions. Other methods such as 7(c) and 7(d) have the problem of incorrectly segmenting lesions in disease-free areas, and methods like 7(g) and 7(i) incorrectly segment the background as lesions. Figure 8 The nlb multi-leaf disease segmentation results are presented. Since the color of the background is affected by the light and is similar to the leaves, the leaves are incorrectly segmented, such as 8(e), 8(f) and 8(h). By comparing the proposed method 8(j) with 8(c) and 8(d), it can be seen that the performance of the lesion segmentation is similar, and most of the labeled lesions can be segmented out, but 8(j) is better in the lesion edge segmentation and leaf edge segmentation. From the above experimental results, it can be seen that the LKCAFormer method can eliminate background noise as much as possible in the CD&S test set with complex backgrounds and multiple leaves, focus on the leaf area, and segment the edges of lesions of different forms that are either scattered or dense, showing better segmentation effects.

[0124] 1.4 Ablation Experiment

[0125] In this section, five sets of ablation experiments are designed to verify the adaptability of all LK-COA modules to different model architectures and their effectiveness in optimizing global feature modeling and detail feature methods. Specifically, in Test 1, LKCAFormer-TR is a model in which the LK-COA module is removed and replaced by three Transformer blocks. Test 2 removes the traditional Transformer block and adds an LK-COA module. Test 3 adds two LK-COA modules. Test 4 adds three LK-COA modules, which is the proposed model, and Test 5 adds four LK-COA modules. In addition, the ablation experiments in this section use three disease datasets in single-CD&S. The evaluation results of the ablation study are recorded in Table 9.

[0126] Table 9 Ablation study results of three corn leaf datasets

[0127]

[0128] As shown in Table 9, by comparing Experiment 1 and Experiment 2, after replacing the traditional transformer block in the model with the proposed LK-COA module, the segmentation accuracy dropped significantly. However, when the LK-COA module was added to three, that is, the model of Experiment 4, the IOU of the lesion segmentation under the three corn disease test sets was improved by 3.11%, 1.26% and 1.89% respectively compared with Experiment 1. On the contrary, when the LK-COA module was added to four, the IOU of the lesion segmentation in Experiment 5 decreased, but the accuracy of background and leaf segmentation was improved, which was attributed to the global feature perception ability of the large kernel convolution. In terms of the comprehensive segmentation effects of diseases, leaves and background, the encoder with three LK-COA modules superimposed is more balanced in improving the segmentation accuracy. While taking into account the perception of global features, it can also better optimize the model's ability to extract details and edge features, thereby effectively improving the segmentation performance of background, leaves and lesions. Fig. 9 The results of leaf and spot segmentation for each test are shown. Figure 8 In the two rows of Test 1 and Test 4, the replacement of the LK-COA module can enable the model to extract more microscopic spots and make the segmentation of leaves and spot edges more refined. From the comparison of this row of Test 4 with other test results, the proposed method can not only segment the contours of leaves and lesions more clearly, but also extract more small lesion areas, clearly segment the lesion edges, and alleviate the adhesion of small lesions in dense lesion areas. The experimental results show that this method improves the segmentation effect of different corn leaf lesions.

[0129] It should be noted that the embodiments of the present invention can be implemented by hardware, software, or a combination of software and hardware. The hardware part can be implemented using dedicated logic; the software part can be stored in a memory and executed by an appropriate instruction execution system, such as a microprocessor or dedicated design hardware. It can be understood by a person of ordinary skill in the art that the above-mentioned devices and methods can be implemented using computer executable instructions and / or contained in a processor control code, such as a carrier medium such as a disk, CD or DVD-ROM, a programmable memory such as a read-only memory (firmware), or a data carrier such as an optical or electronic signal carrier. Such code is provided on the carrier medium. The device and its modules of the present invention can be implemented by hardware circuits such as very large-scale integrated circuits or gate arrays, semiconductors such as logic chips, transistors, etc., or programmable hardware devices such as field programmable gate arrays, programmable logic devices, etc., can also be implemented by software executed by various types of processors, and can also be implemented by a combination of the above-mentioned hardware circuits and software, such as firmware.

[0130] The above description is only a specific implementation mode of the present invention, but the protection scope of the present invention is not limited thereto. Any modification, equivalent substitution and improvement made by any technician familiar with the technical field within the technical scope disclosed by the present invention and within the spirit and principle of the present invention should be covered by the protection scope of the present invention.

Claims

1. A method for constructing a cross-scale large kernel convolution corn leaf disease segmentation model based on a coordinated attention mechanism, characterized in that: The following steps are involved: a. Encoder based on multi-level large kernel convolution and attention modules to extract global features and fine-grained features; b. Cross-scale attention decoder, used to fuse global semantic information with local feature information; (2) A large kernel convolution operation is used to model the global features of the input leaf image, and a collaborative attention mechanism is used to capture local features and optimize the interaction between global and local features; (3) In the decoder stage, the high-level semantic features and low-level spatial detail features extracted by the encoder are gradually integrated, the features are fused through a cross-scale attention mechanism, and then upsampled to the resolution of the input image to generate disease segmentation results.

2. The method for accurate segmentation of corn leaf diseases according to claim 1, characterized in that: The size of the input leaf image is 512×512×3. The encoder first extracts features through a feature extraction head consisting of two stacked three-by-three depthwise convolutions, and outputs a feature map that is one-fourth the size of the original image, thereby effectively reducing the number of parameters of the encoder.

3. The method for accurate segmentation of corn leaf diseases according to claim 1, characterized in that: The encoder consists of three layers of large kernel convolution and collaborative attention modules. The large kernel convolution reduces the computational cost by decomposing the large kernel into multiple small kernel convolutions and extending the depth convolution. At the same time, the collaborative attention mechanism is combined to capture the high-frequency features of the leaf edge and the fine-grained features of the disease spots.

4. The method for accurate segmentation of corn leaf diseases according to claim 1, characterized in that: The decoder consists of three cross-scale attention decoding modules, each of which receives the encoder features at the same level and the features from the previous decoder. It calculates the similarity weights of fine-grained high-frequency information through the attention mechanism, and fuses it with the coarse-grained low-frequency information before upsampling it to the original spatial size.

5. The method for accurate segmentation of corn leaf diseases according to claim 1, characterized in that: The cross-scale attention mechanism introduces deep convolution with a square kernel size into the input feature map to calculate the weighted values ​​in the spatial area, optimizes the global and local feature extraction capabilities of the feature map by using information interaction between channels, and establishes the relationship between pixels through convolution to reduce the memory overhead of high-resolution image processing.

6. The method for accurate segmentation of corn leaf diseases according to claim 1, characterized in that: The loss function is designed by combining cross entropy loss and Dice loss, where: The cross entropy loss is used to calculate the error between the predicted probability of each pixel belonging to the target category in the model output and the true label; Dice loss is used to enhance the segmentation ability of small target areas by calculating the overlap between the predicted image and the true label image in the target category; The combination of cross entropy loss and Dice loss cooperate with each other during the network training process to improve the segmentation accuracy.

7. The method for accurate segmentation of corn leaf diseases according to claim 1, characterized in that: The fused features output by the decoder are nonlinearly processed through the designed lightweight segmentation head module, and a segmentation result map consistent with the size of the input image is directly generated. The segmentation result includes multiple predefined categories, each category representing a disease type or a healthy area.

8. A lightweight segmentation system for accurate segmentation of corn leaf diseases, characterized in that: include: (1) An encoder module based on multi-level large kernel convolution and attention mechanism is used to extract global features and fine-grained features of the input corn leaf image; (2) A cross-scale attention decoder module is used to fuse the high-level semantic features and low-level spatial detail features extracted by the encoder and upsample the fused features to the input image size to generate the disease segmentation result; (3) A data processing module for model training and inference, which preprocesses the input image to adapt it to the network input and post-processes the output result to optimize the segmentation effect.

9. The system according to claim 8, characterized in that The encoder module uses large kernel convolution operations to globally model the input image and combines it with a collaborative attention mechanism to capture local features in both spatial and channel dimensions.

10. The system according to claim 8, characterized in that The decoder module uses a cross-scale attention mechanism to interactively fuse high-level semantic features and low-level spatial features, and restores the original resolution of the image by progressive upsampling to generate an accurate segmentation map of the diseased area.

Citation Information

Patent Citations

  • Remote sensing image fusion method based on large kernel attention mechanism for multi-scale feature enhancement

    CN114936995A

  • Rubber disease image recognition method, mobile device and embedded device

    CN115546611A

  • Multi-attention-combined cross-scale remote sensing image cultivated land extraction method

    CN116844039A

  • Fine-grained image classification method, system and device based on space attention and medium

    CN117726874A

  • Retinal blood vessel image segmentation method based on large kernel convolution

    CN117830325A

Cited By

  • Image segmentation method based on deep learning remote sensing image

    CN120580436A

  • Image segmentation method based on deep learning remote sensing image

    CN120580436B

  • Medical image segmentation method based on dynamic region aggregation convolution

    CN121121094A

  • A medical image segmentation method based on dynamic region aggregation convolution

    CN121121094B