Cotton leaf verticillium wilt segmentation model based on YOLO11 framework

By introducing Swin Transformer and multi-scale feature fusion attention mechanism on the YOLO11 framework, combined with sliding window and attention mechanism, the missegment problem of cotton Verticillium Worm leaf disease segmentation model in complex field backgrounds is solved, and an integrated solution with high precision and real-time is achieved.

CN120125830APending Publication Date: 2025-06-10SANYA NATIONAL INSTITUTE OF SOUTHERN BREEDING CHINESE ACADEMY OF AGRICULTURAL SCIENCES

Patent Information

Application Number
CN202510625856.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-15
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art cotton Verticillium Worm leaf disease segmentation model has high missegment rate and missed detection of small lesions in complex field backgrounds, and its performance plummeted in small sample scenarios, making it difficult to meet the real-time monitoring needs in fields.

Method used

The cotton leaf verticillium worries segmentation model based on YOLO11 framework is adopted, and a multi-scale feature fusion attention mechanism and sliding window mechanism are introduced through Swin Transformer as a feature extraction module, combining channel and spatial attention mechanisms, and using sliding weighted loss function.

Benefits of technology

It significantly improves the accuracy of crop diseased leaves detection and segmentation in complex farmland scenarios, reduces the error detection rate, improves the performance of small lesions detection, enhances the robustness of the model for light changes, occlusion and similar color interference, and is suitable for real-time monitoring in the field.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125830A_ABST
    Figure CN120125830A_ABST
Patent Text Reader

Abstract

The invention relates to the field of agricultural disease intelligent detection and computer vision interdiscipline, in particular to a cotton leaf verticillium wilt segmentation model based on a YOLO11 framework, and adopts the technical scheme that Swin Transform is adopted as a feature extraction module, and a layered framework and a sliding window mechanism are introduced; a multi-scale feature fusion attention mechanism is adopted as an attention mechanism module, features of different scales are extracted by using five different branches, then a comprehensive feature map is spliced in a channel dimension, and finally, a channel attention mechanism and a space attention mechanism are used to calibrate the merged feature map in a parallel mode. A sliding weighted loss function is adopted as a loss function; wherein the step of calibrating the merged feature map comprises the steps of obtaining an enhanced feature map through fusion after channel weighting and space weighting are achieved, and generating a final output feature map through dimensionality reduction and integration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the interdisciplinary field of intelligent detection of agricultural diseases and computer vision, and particularly to a cotton leaf verticillium wilt segmentation model based on the YOLO11 framework. Background Art

[0002] Cotton verticillium wilt is a devastating disease caused by Verticillium dahliae, resulting in a cotton field yield reduction of 10% - 80%. Traditional manual surveys have problems such as low efficiency and strong subjectivity. Disease segmentation technology based on computer vision can provide pixel-level lesion localization and quantitative disease assessment, which is crucial for precise prevention and control and resistance breeding. However, the recognition of complex field backgrounds (such as soil, weeds, and plastic film interference) and small lesions (occupying <5% of the leaf area) remains a technical challenge.

[0003] Under the interference of complex backgrounds, the segmentation accuracy of traditional models is insufficient. Under the influence of interference factors such as plastic film reflection and weeds in cotton fields, the mis-segmentation rate is relatively high, and the lesion boundaries are blurred, making it difficult to identify the transition zone between healthy and diseased areas. There is a situation of missed detection of small lesions, and the existing solutions have insufficient detection ability for lesions with a smaller area.

[0004] Developing a high-precision instance segmentation model for the diseased part of verticillium wilt leaves that can adapt to complex field backgrounds, realizing pixel-scale lesion localization and segmentation, and further realizing disease grading and disease assessment is of great significance for promoting the cultivation of cotton verticillium wilt-resistant varieties, assisting in precise disease prevention and control, and improving the level of agricultural intelligence.

[0005] Convolutional neural networks (CNNs) have good image feature extraction capabilities and are superior to other methods in terms of accuracy and efficiency. CNN-based segmentation can not only identify different classes but also provide other information such as the spatial location of these classes. The BLSNet based on U-Net can estimate the severity of rice bacterial leaf streak to improve the accuracy of lesion area segmentation. The full-resolution convolutional network (FrCNnet) model based on convolutional neural networks can accurately segment mango leaf damage. The hybrid segmentation algorithm based on image masking and REC (IMHSA) can solve the problems of limited datasets and overlapping convolutional neural network models during the classification process. The instance segmentation models Mask R-CNN and Mask Scoreing R-CNN segmented and identified peach diseases, extracting the names, locations, and areas of the diseases. Variants of U-net, SegNet, PSPNet, FPN, and DeepLabv3+ were applied to estimate the disease severity of soybean rust and wheat brown leaf spot.

[0006] The existing technologies mainly have problems such as a high mis-segmentation rate under complex background interference, prominent problems of missed detection of small lesions, significant contradictions between real-time performance and accuracy, and limited model generalization ability.

[0007] Traditional segmentation models (such as the combination of DeepLabv3+ and K-means clustering) have an incorrect segmentation rate of over 15% in the complex background of cotton fields (soil, plastic film, weeds, etc.). In particular, it is difficult to accurately distinguish the blurred transition areas of leaves caused by Verticillium wilt (the boundary between the diseased area and the healthy area). The existing YOLOv8-seg model has a recall rate of only 61.5% for lesions with an area less than 50 px, and it has not been optimized for the specific small lesion morphologies (such as dot-like and star-like) of cotton Verticillium wilt. Although models introducing attention mechanisms (such as CBAM and CA) improve the segmentation accuracy, the computational complexity increases by more than 40%, resulting in an inference speed of less than 10 FPS1 on embedded devices (such as Jetson Nano), making it difficult to meet the requirements of field real-time monitoring. Existing solutions rely on large-scale labeled data (tens of thousands of samples), while in actual small-sample field scenarios (<1000 images), the model performance drops sharply, and the problems of over-segmentation and under-segmentation coexist.

[0008] In view of this, we propose a segmentation model for cotton leaf Verticillium wilt based on the YOLO11 framework to solve the existing problems. Summary of the Invention

[0009] The purpose of the present invention is to provide a segmentation model for cotton leaf Verticillium wilt based on the YOLO11 framework to solve the problems raised in the above background technology.

[0010] To achieve the above object, the present invention provides the following technical solution: A segmentation model for cotton leaf Verticillium wilt based on the YOLO11 framework, including a feature extraction module, an attention mechanism module, and a loss function. The Swin Transformer is used as the feature extraction module, introducing a hierarchical architecture and a sliding window mechanism; the multi-scale feature fusion attention mechanism is used as the attention mechanism module, using five different branches to extract features of different scales, then stitching them together into a comprehensive feature map in the channel dimension, and finally using the channel attention mechanism and the spatial attention mechanism to calibrate the merged feature map in a parallel manner; the sliding weighted loss function is used as the loss function; wherein, the steps of calibrating the merged feature map include achieving channel weighting and spatial weighting and then obtaining an enhanced feature map through fusion, and generating the final output feature map through dimensionality reduction and integration.

[0011] Further, among the five different branches of the multi-scale feature fusion attention mechanism, the first branch uses a 1×1 convolutional kernel to directly extract features without changing the spatial scale; the second to the fourth branches all use 3×3 convolutional kernels, and the dilation rates are 6, 12, and 18 in sequence, gradually expanding the receptive field step by step to capture more extensive context information; the fifth branch is used to extract context features, and a global average pooling branch is used to enhance the model's understanding ability of the overall layout.

[0012] Furthermore, in the channel weighting implementation of the multi-scale feature fusion attention mechanism, first, global average pooling is performed on the comprehensive feature map to obtain the global features of each channel; then, the importance weights of each channel are learned through two fully connected layers, where the two fully connected layers use the ReLU and Sigmoid activation functions respectively; finally, the weights are multiplied with the original feature map channel by channel to achieve channel weighting.

[0013] Furthermore, in the spatial weighting implementation of the multi-scale feature fusion attention mechanism, first, global pooling is performed on the comprehensive feature map in the channel dimension in the spatial attention mechanism to obtain the spatial feature map; then, the importance weights of each spatial position are learned through a 1×1 convolution and the Sigmoid activation function; finally, the weights are multiplied with the original feature map element by element to achieve spatial weighting.

[0014] Furthermore, in the process of obtaining the enhanced feature map by the multi-scale feature fusion attention mechanism, the output feature maps of the channel attention and spatial attention mechanisms are fused through an element-wise addition operation to finally obtain the enhanced feature map.

[0015] Furthermore, the feature maps calibrated by channels and spaces respectively are subjected to an element-wise addition operation with the original merged feature map to integrate and enhance the features.

[0016] Furthermore, in the process of generating the final output feature map by the multi-scale feature fusion attention mechanism, the enhanced feature map is reduced in dimension and integrated through a 1×1 convolutional layer to generate the final output feature map.

[0017] Furthermore, in the feature extraction module, the Swin Transformer contains 4 stages. Except for the first stage, each stage will first perform a downsampling operation through the Patch Merging layer to reduce the resolution of the input feature map, expand the receptive field layer by layer, and thus obtain global information.

[0018] Furthermore, first, the Patch Partition layer is used to divide the input RGB image into several non-overlapping patches. Each patch is regarded as a token, and its features are composed of the concatenation of the original pixel values. By applying the LinearEmbedding layer, the original pixel value features are mapped to any dimension. Then, several Swin Transformer modules are applied to the patch tokens while keeping the number of tokens unchanged, constituting the first stage; the Patch Merging layer is used to reduce the number of tokens, and the Swin Transformer module is continued to be used for feature transformation, constituting the second stage; the operations constituting the second stage are repeated twice, corresponding to the third stage and the fourth stage respectively.

[0019] Furthermore, each Swin Transformer block includes a window-based multi-head self-attention mechanism, layer normalization, and a multi-layer perceptron; the multi-head self-attention mechanism module uses window-based attention to capture the dependencies within the local region of the input sequence, layer normalization is performed for normalization, and the multi-layer perceptron module uses multi-layer non-linear transformations to capture the complex patterns in the data.

[0020] Compared with the prior art, the beneficial effects of the present invention are: The multi-scale feature fusion attention mechanism of the present invention significantly improves the accuracy of crop diseased leaf detection and segmentation in complex farmland scenes by integrating channel and spatial attention mechanisms and combining multi-scale feature extraction with a parallel calibration strategy. Its core design includes multi-scale feature extraction with five branches: Branch 1 uses 1×1 convolution to retain local details, Branches 2-4 gradually expand the receptive field through 3×3 dilated convolutions (dilation rates 6, 12, 18) to capture multi-level contexts, and Branch 5 uses global average pooling to extract global semantic features; after the features of all branches are concatenated in the channel dimension, they are respectively calibrated through parallel channel attention (global pooling + fully connected weight learning) and spatial attention (channel pooling + 1×1 convolution weight learning), and finally, dimensionality reduction and fusion are performed through residual connection and 1×1 convolution to balance local details and global semantics. This mechanism combines progressive dilated convolution with global-local feature fusion and adopts a parallel dual-attention structure, which not only avoids the information attenuation problem of traditional serial attention but also enhances feature stability through residual connection, thus achieving a robust expression of diseased leaf features under complex background interference.

[0021] Compared with the existing attention mechanisms, the multi-scale feature fusion attention mechanism has more advantages in terms of efficiency and performance: compared with the squeeze-and-excitation attention mechanism module that only focuses on channel weights or the convolutional block attention mechanism that serially processes channel-spatial attention, its parallel dual-attention design can retain both spatial-channel independence and reduce information loss; compared with the computationally intensive self-attention model, the multi-scale feature fusion attention mechanism realizes long-range dependence modeling at low cost through dilated convolution and global pooling, making it more suitable for dense prediction tasks; and compared with the selective kernel network that dynamically selects convolutional kernels, the multi-scale feature fusion attention mechanism explicitly introduces spatial weight calibration, making it more applicable to the precise localization of lesion edges and low-contrast regions in farmland scenes. Through multi-scale feature fusion and residual optimization, this mechanism significantly improves the robustness of the model against illumination changes, occlusion, and similar color interference while maintaining a low computational cost, providing an efficient solution for fine-grained analysis of agricultural images.

[0022] The present invention has significant advantages compared with the prior art. First, in terms of the segmentation accuracy of small targets in complex backgrounds, a dynamic multi-scale feature fusion attention module is adopted in combination with a dual-branch feature pyramid, and background noise is suppressed through the collaborative action of channel attention and spatial attention, improving the detection performance of small lesions and reducing the false detection rate; second, to address the problem of sample imbalance, a dynamic focal weighted loss function is proposed and a 3-fold loss weight is assigned to difficult samples. At the same time, the local feature learning ability is enhanced through a sliding window feature extraction module. Compared with traditional methods, no additional spectral equipment is required and the problems of over-segmentation and under-segmentation can be effectively alleviated; finally, in terms of cross-scene generalization ability, based on a multi-modal feature fusion strategy combined with an attention mechanism, it dynamically adapts to different environments, and the dependence on labeled data is reduced through transfer learning, making it more advantageous than existing models in terms of the amount of labeled data and dynamic environment adaptability.

[0023] Therefore, the present invention applies the multi-scale feature fusion attention mechanism to the cotton Verticillium wilt leaf segmentation model, realizing the effective segmentation of small target lesions in complex field backgrounds, improving the performance of the cotton Verticillium wilt leaf segmentation model. The feature extraction module with a sliding window and the feature fusion module with a pyramid structure effectively improve the target feature extraction ability in complex field backgrounds. The multi-scale feature fusion attention mechanism enhances feature expression by integrating channel and spatial attention mechanisms. The sliding weighted loss function balances simple and difficult samples in the samples, enhancing the recognition ability of small target lesions. Compared with other attention mechanisms, the multi-scale feature fusion attention mechanism achieves a better improvement in model performance. The model precision reaches 88.2%, and the performance of the benchmark model is improved. The combination of all modules has the best improvement effect on the model, and the model precision reaches 91.2%.

[0024] In summary, by combining multi-scale feature fusion and dynamic attention mechanism, the complex background interference and missed detection of small lesions in the current technology in the field scenario of cotton Verticillium wilt are solved, pixel-level lesion segmentation in the complex field scenario is achieved, and key technical support for precise pesticide application and resistance breeding is provided. Brief Description of the Drawings

[0025] Figure 1 It is the architecture diagram of the cotton leaf Verticillium wilt segmentation network based on the YOLO11 framework of the present invention; Figure 2 It is the module composition diagram of CBS, C3K2, SPPF, and C2PSA of the present invention; Figure 3 It is the Swin Transformer architecture diagram of the feature extraction module of the present invention; Figure 4 It is the structural diagram of the multi-scale feature fusion attention mechanism of the attention mechanism module of the present invention; Figure 5 It is the effect comparison diagram of the cotton leaf Verticillium wilt segmentation model of the present invention. Detailed Embodiment

[0026] The technical solution of the present invention will be further described below in conjunction with the accompanying drawings and specific embodiments. Embodiment 1

[0027] Aiming at the limitations of field manual investigation of cotton Verticillium wilt, the difficulty in identifying small lesions, the imbalance between the number of simple samples and difficult samples, etc., an instance segmentation model for the diseased parts of cotton Verticillium wilt leaves based on a computer vision framework is adopted, and a variety of strategies are used to optimize the model network structure to further improve the segmentation accuracy of small target lesions, improve the model performance while reducing the computational complexity. A multi-scale feature fusion attention module is proposed and compared with other existing attention mechanisms. Aiming at the characteristics of complex environmental background information and small lesion area, a sliding window feature extraction module, a multi-scale feature fusion attention mechanism, and a sliding weighted loss function network are adopted, and ablation experiments are used to compare the role of different modules in improving the model performance. The relevant research results have good guiding and reference significance for improving the efficiency of field investigation of Verticillium wilt, mining cotton Verticillium wilt resistance genes and cultivating resistant varieties, as well as precise pesticide application and disease prevention and control.

[0028] The present invention is based on the latest version of YOLO11 for pixel-level segmentation of the diseased areas of cotton Verticillium wilt leaves. The YOLO framework has an open and flexible design concept, providing models of different sizes, enabling it to effectively integrate various tasks and applications in different scenarios, and can easily adapt to hardware platforms with different network connection distances and computing performances, ranging from edge computing to cloud computing. YOLO11 has made significant improvements in the architecture and training methods. Based on the good performance of the existing YOLO versions, it further enhances the performance boundary of the YOLO framework with more advanced accuracy, speed, and efficiency. As a current cutting-edge image segmentation algorithm, YOLO11 provides a versatile choice for a wide range of computer vision tasks in the field of agricultural pest and disease image segmentation.

[0029] As Figure 1 shown, in terms of the network structure, the YOLO11 network consists of a backbone network, a neck layer, and a head layer. The backbone network extracts image features and is mainly composed of CBS (Convolution-BatchNorm-SiLU), C3K2 (an optimization of the CSP Bottleneck structure, which combines the characteristics of the Bottleneck module and the Cross Stage Partial connection), SPPF (SpatialPyramid Pooling-Fast), and C2PSA (Cross Stage Partial with Pyramid Squeeze Attention). The neck layer is used for multi-scale feature fusion of feature maps and is mainly composed of Upsample, Concat, C3K2, and CBS. The head layer is used to predict bounding boxes and classes. In addition, Swin Transformer (a vision Transformer architecture based on the self-attention mechanism) is used as a feature extraction module, and a multi-scale feature fusion attention mechanism (MFFA) is proposed to enhance the attention to cotton Verticillium wilt targets from the perspective of multi-channel and multi-spatial scale fusion.

[0030] As Figure 2As shown, the CBS module is a fundamental and important component, mainly composed of three parts: Conv (convolutional layer), BN (Batch Normalization), and SiLU (activation function), which realizes feature extraction and non-linear transformation of the input feature map. The CBS module is the basic unit in the network, and a more complex backbone network structure with powerful feature extraction capabilities is constructed by stacking multiple CBS modules. The C3K2 module is an important feature extraction component, including two Convs, Split, n C3Ks (convolutional kernels with multiple scales), and Concat, adopting variable convolutional kernels and channel separation strategies. The C3K2 module introduces the multi-scale convolutional kernel C3K, where K is an adjustable convolutional kernel size, such as 3×3, 5×5, etc., and captures more extensive context information by expanding the receptive field, which is suitable for scenes with complex backgrounds. The C3K2 module divides the input features into two parts, which are respectively subjected to deep feature extraction through ordinary convolution operations and multiple C3K structures, and finally the two parts of the features are concatenated and fused, which can effectively extract deep features. The SPPF module is a module that enhances the feature extraction ability, including two Convs, 3 MaxPools (maximum pooling), and Concat, which expands the receptive field of the network by introducing multi-scale feature map pooling. The SPPF uses pooling operations with different scales, such as 5×5, 3×3, 1×1 pooling, to perform pooling operations on the feature map and concatenate them, enabling the model to capture features of different scales without significantly increasing the computational cost, so as to enhance the detection ability for objects of different spatial scales. The C2PSA module includes two Convs, Split, n PSAs (Position-wise Self-Attention), and Concat.

[0031] To better adapt to the specific requirements in the segmentation of the diseased area of cotton Verticillium wilt leaves, the YOLOv11 model was selected to balance the speed and accuracy performance of the segmentation model. To address these issues while achieving lightweight design and high precision, the following improvements were made to YOLO11, and the improved model was named CVWSeg.

[0032] In the cotton Verticillium wilt leaf image segmentation network of the present invention, Swin Transformer is used as the feature extraction module. By introducing a hierarchical architecture and a sliding window mechanism, Swin Transformer achieves a balance between performance and efficiency, and solves the problem of high computational complexity of traditional Transformer models in computer vision tasks. Due to its advanced performance in vision tasks, Swin Transformer is commonly used as the backbone network of many vision model architectures and is widely applied to vision tasks such as image classification, object detection, and segmentation. Swin Transformer provides different versions of large, medium, and small models, and the appropriate one can be freely selected for use.

[0033] As Figure 3 shown, the entire model of the Swin Transformer architecture adopts a hierarchical design and consists of a total of 4 stages. Except for the first stage, each stage will first perform downsampling operations through the Patch Merging layer to reduce the resolution of the input feature map, gradually expand the receptive field layer by layer, and thus obtain global information. First, the PatchPartition layer is used to divide the input RGB image into several non-overlapping Patches. Each Patch is regarded as a Token, and its features are composed of the concatenation of the original pixel values. Taking a 4×4 Patch as an example, the feature dimension of each Patch is 4×4×3 = 96. By applying the Linear Embedding layer, the original pixel value features are mapped to any dimension (denoted as C). Then, several Swin Transformer modules are applied to the Patch Tokens while keeping the number of Tokens unchanged (H / 4×W / 4), forming the first stage. In order to form a hierarchical representation, as the network depth increases, the number of Tokens is reduced through the Patch Merging layer, and the Swin Transformer module is continued to be used for feature transformation, forming the second stage. This process is repeated twice, corresponding to the third stage and the fourth stage respectively, and the final output resolution is H / 32×W / 32. Among them, H and W are the height and width of the original image respectively.

[0034] Each Swin Transformer block consists of several sub - modules, including window - based multi - head self - attention mechanism (W - MSA), layer normalization (LN), multi - layer perceptron (MLP), and spatial multi - scale attention (SW - MSA). The W - MSA module uses window - based attention to capture dependencies within local regions of the input sequence, LN performs normalization, and the MLP module uses multi - layer non - linear transformations to capture complex patterns in the data. By stacking two consecutive Swin Transformer Blocks, the model can capture a broader context by combining local and global dependencies. Two consecutively connected Swin Transformer Blocks are expressed as follows:

[0035] Swin Transformer is constructed by replacing the standard multi - head attention (MSA) module in the Transformer block with a module based on moving windows, while other layers remain unchanged. The shifted window partitioning method alternates between two partitioning configurations in consecutive Swin Transformer blocks. In the I - th layer, regular window partitioning is adopted, and self - attention is calculated within each window. In the (I + 1) - th layer, shifted window partitioning is used, resulting in new windows. The self - attention calculation in the new windows crosses the boundaries of the windows in the I - th layer, providing connections between them.

[0036] In image processing tasks, there are targets of different shapes and sizes. The morphological characteristics of Verticillium wilt lesions on cotton vary greatly. There are some lesions with large areas and consistent characteristics, as well as some lesions with small areas and large variations in characteristics. These small - sized lesions are often ignored by the model in image segmentation tasks, thus reducing the model performance. The perception method of traditional convolutional layers is relatively fixed, which limits its ability to capture target features at different channels and different spatial scales. Channel features represent different types of features extracted from the input image, such as global information about a specific pattern or attribute in the image. Spatial features are the feature information of pixels in the spatial distribution, such as various information like edges, textures, and shapes. Field cotton Verticillium wilt images have complex background information and small target sizes, making it difficult for traditional convolutional layers to highlight key regions. Therefore, by introducing an attention mechanism, the model's ability to recognize small - sized Verticillium wilt targets is enhanced. By dynamically focusing on channels and different spatial scales with different degrees of importance, the most suitable channels and spatial scales for target features and sizes are found, thereby improving the quality of feature extraction.

[0037] The present invention proposes a multi-scale feature fusion attention mechanism to enhance the attention to the target of Verticillium wilt of cotton from the perspective of multi-channel and multi-space scale fusion. Comparative studies are carried out with existing attention mechanism modules that are widely used, including the Convolutional Block Attention Module (CBAM), the Squeeze and Extraction (SE), and the Coordinate Attention (CA). CBAM is an attention module for feed-forward convolutional neural networks, which cascades the channel attention mechanism and the spatial attention mechanism to effectively refine the features. The SE attention mechanism uses a global average pooling operation in the squeezing stage to compress the input feature map into a vector, and uses the sigmoid function to transform the elements in the vector in the excitation step, improving the model's performance by enhancing the channel features of the feature map. CA is suitable for mobile networks. It decomposes the channel attention in the horizontal and vertical directions to aggregate features from different directions in order to capture long-range dependencies, enhancing the model's ability to localize targets and improving the image segmentation task.

[0038] MFFA enhances feature representation by integrating channel and spatial attention mechanisms and capturing both detailed information and broad context information in the image. This is crucial for the detection and segmentation of diseased crop leaves in the complex background of the field environment. As Figure 4As shown, five different branches are used to extract features of different scales, and then they are concatenated into a comprehensive feature map in the channel dimension. Finally, a channel attention mechanism and a spatial attention mechanism are used to calibrate the merged feature map in a parallel manner. The first branch uses a 1×1 convolutional kernel to directly extract features without changing the spatial scale. The second to the fourth branches all use 3×3 convolutional kernels, and the dilation rates are 6, 12, and 18 in sequence, gradually expanding the receptive field step by step to capture more extensive context information. The fifth branch is used to extract context features, and a global average pooling branch is used to enhance the model's understanding ability of the overall layout. First, global average pooling (Global Pooling) is performed on the comprehensive feature map to obtain the global features of each channel. Then, the importance weights of each channel are learned through two fully connected layers (using ReLU and Sigmoid activation functions respectively). Finally, these weights are multiplied by the original feature map channel by channel to achieve channel weighting. Secondly, in terms of the spatial attention mechanism, global pooling is performed on the comprehensive feature map in the channel dimension to obtain the spatial feature map. Through a 1×1 convolution and a Sigmoid activation function, the importance weights of each spatial position are learned. These weights are multiplied by the original feature map element by element to achieve spatial weighting. Finally, the output feature maps of the channel attention and spatial attention mechanisms are fused through an element-wise addition (Add) operation to finally obtain the enhanced feature map. At this time, the feature maps calibrated by the channel and the spatial respectively are added to the original merged feature map element by element to integrate and enhance the relevant features. The enhanced feature map finally passes through a 1×1 convolutional layer for dimensionality reduction and integration to generate the final output feature map.

[0039] The present invention uses the Slide Loss function as the loss function in the cotton Verticillium wilt leaf image segmentation network to solve the problem of unbalanced dataset samples. In image segmentation, the problem of unbalanced numbers of hard samples and easy samples in the dataset is often encountered. Easy samples generally have obvious features and clear boundaries. Hard samples mostly have problems such as occlusion, deformation, blurring, and small size. Generally speaking, there are more easy samples and fewer hard samples in the dataset. In addition, during the model training process, the loss value of easy samples is lower than that of hard samples. Therefore, the segmentation model performs better in the target segmentation of easy samples, but has lower segmentation performance for hard samples. In the segmentation of the diseased parts of cotton leaves infected with Verticillium wilt, the disease characteristics of Verticillium wilt are diverse. There are both large areas of withered symptoms and small, scattered spots. At the same time, the Verticillium wilt symptoms distributed on the leaf edges are mixed with the field background. All these pose difficulties for the Verticillium wilt segmentation model. The Slide Loss function is a piecewise function that uses the intersection over union (IoU) of the predicted bounding box and the ground truth bounding box to distinguish easy samples and hard samples. By assigning relatively higher weights to samples with IoU near the average value during the model training process, the model can pay more attention to hard samples, thereby improving the model performance. The calculation formula of the Slide Loss function is as follows:

[0040] where x is the sample IoU, and the threshold μ is the average value of the IoU values of all bounding boxes. The Slide Loss function can adaptively learn the positive and negative sampling threshold parameters. Samples with IoU less than the threshold are regarded as negative samples, and samples with IoU greater than the threshold are regarded as positive samples. However, due to the above-mentioned problem of unbalanced samples, samples near the boundary often have relatively large loss values. By setting a high weight near the threshold to increase the loss value of hard samples, the model can pay more attention to the hard samples that are misclassified. Through the optimization of the loss function, the segmentation model of the diseased parts of cotton leaves infected with Verticillium wilt can better learn hard samples, make more full use of these samples to train the network, and finally improve the model performance.

[0041] The present invention uses the CVWSeg model improved based on YOLOv11 to segment the healthy parts and diseased parts of cotton leaves infected with Verticillium wilt. The model effect is as Figure 5 shown. There are still many cases of misdetection and missed detection in the detection results of the basic model (YOLOv11) for the diseased parts of cotton leaves infected with Verticillium wilt. Specifically, Figure 5Taking the leaves in [the specific context] as an example, it can be found that the basic model misses the detection of some disease-affected parts with larger areas and only detects some of the disease-affected parts. At the same time, some soil backgrounds similar to the disease-affected parts are wrongly detected as disease-affected parts. In addition, it can be found that the basic model has poor recognition and segmentation ability for disease-affected parts with smaller areas, and the segmentation result of the model does not fit well with the shape of the disease-affected parts. The improved model has better optimized the above problems. The segmentation result area is clear and closely fits the disease-affected parts. In particular, the segmented disease-affected area is no longer a fragmented and scattered area, but a continuous and complete area. From the segmentation result, it can be seen that the improved model has a more independent and complete segmentation effect on the leaves. The healthy area and the disease-affected area can fit completely, forming the whole cotton leaf. The research results show that the CVWSeg model can effectively segment the cotton Verticillium wilt leaves in the complex field environment, providing a basis for the assessment of the Verticillium wilt disease level.

[0042] The above specific embodiments are only several preferred embodiments of the present invention. Based on the technical solution of the present invention and the relevant inspirations of the above embodiments, those skilled in the art can make various alternative improvements and combinations to the above specific embodiments.

Claims

1. A cotton leaf Verticillium wilt segmentation model based on the YOLO11 framework, including a feature extraction module, an attention mechanism module, and a loss function, characterized in that: Swin Transformer is used as the feature extraction module, and a hierarchical architecture and a sliding window mechanism are introduced. A multi-scale feature fusion attention mechanism is used as the attention mechanism module, and five different branches are used to extract features of different scales, which are then spliced ​​into a comprehensive feature map in the channel dimension. Finally, the channel attention mechanism and the spatial attention mechanism are used to calibrate the merged feature map in parallel. A sliding weighted loss function is used as the loss function. Among them, the step of calibrating the merged feature map includes realizing channel weighting and spatial weighting to obtain an enhanced feature map through fusion, and generating the final output feature map through dimensionality reduction and integration. Moreover, in the feature extraction module, Swin Transformer contains 4 stages. Except for the first stage, each stage will first perform a downsampling operation through the PatchMerging layer to reduce the resolution of the input feature map, expand the receptive field layer by layer, and then obtain global information.

2. According to claim 1, a cotton leaf Verticillium wilt segmentation model based on the YOLO11 framework, in five different branches of the multi-scale feature fusion attention mechanism, is characterized in that: The first branch uses a 1×1 convolution kernel to directly extract features without changing the spatial scale; The second to fourth branches all use 3×3 convolution kernels with void rates of 6, 12, and 18, respectively, gradually expanding the receptive field to capture a wider range of contextual information; the fifth branch is used to extract contextual features and enhance the model's ability to understand the overall layout through a global average pooling branch.

3. According to the cotton leaf Verticillium wilt segmentation model based on the YOLO11 framework according to claim 1, in the channel weighted implementation of the multi-scale feature fusion attention mechanism, it is characterized in that: First, the comprehensive feature map is globally averaged pooled to obtain the global features of each channel; then, the importance weight of each channel is learned through two fully connected layers, where the two fully connected layers use ReLU and Sigmoid activation functions respectively; finally, the weight is multiplied with the original feature map channel by channel to achieve channel weighting.

4. According to the cotton leaf Verticillium wilt segmentation model based on the YOLO11 framework according to claim 1, in the spatial weighted implementation of the multi-scale feature fusion attention mechanism, it is characterized in that: First, in terms of the spatial attention mechanism, the comprehensive feature map is globally pooled in the channel dimension to obtain the spatial feature map; then, the importance weight of each spatial position is learned through a 1×1 convolution and a Sigmoid activation function; finally, the weight is multiplied element-by-element with the original feature map to achieve spatial weighting.

5. According to the cotton leaf Verticillium wilt segmentation model based on the YOLO11 framework of claim 1, in the feature map process after the multi-scale feature fusion attention mechanism is enhanced, it is characterized in that: The output feature maps of the channel attention and spatial attention mechanisms are fused through element-wise addition operations to obtain the enhanced feature maps.

6. The cotton leaf Verticillium wilt segmentation model based on the YOLO11 framework according to claim 5, characterized in that: The feature maps that have been calibrated channel-wise and spatially are element-wise added to the original merged feature maps to integrate and enhance the features.

7. According to the cotton leaf Verticillium wilt segmentation model based on the YOLO11 framework of claim 1, in the process of generating the final output feature map by the multi-scale feature fusion attention mechanism, it is characterized in that: The enhanced feature map is reduced in dimension and integrated through a 1×1 convolutional layer to generate the final output feature map.

8. According to the cotton leaf Verticillium wilt segmentation model based on the YOLO11 framework as claimed in claim 1, the four stages included in SwinTransformer are characterized in that: First, the Patch Partition layer is used to divide the input RGB image into several non-overlapping Patches. Each Patch is regarded as a Token, whose features are composed of the concatenation of the original pixel values. The original pixel value features are mapped to any dimension by applying the Linear Embedding layer. Then, several SwinTransformer modules are applied to the Patch Token, and the number of Tokens is kept unchanged, which constitutes the first stage. The number of Tokens is reduced by the PatchMerging layer, and the Swin Transformer module is continued to be used for feature transformation, which constitutes the second stage. The operations constituting the second stage are repeated twice, corresponding to the third and fourth stages respectively.

9. The cotton leaf Verticillium wilt segmentation model based on the YOLO11 framework according to claim 1, characterized in that: Each Swin Transformer block includes a window-based multi-head self-attention mechanism, layer normalization, and a multi-layer perceptron; the multi-head self-attention mechanism module uses window-based attention to capture dependencies within local regions of the input sequence, layer normalization is performed for normalization, and the multi-layer perceptron module uses multiple layers of nonlinear transformations to capture complex patterns in the data.

Citation Information

Patent Citations

  • Medical image segmentation method fusing multi-scale features and multi-attention mechanism based on Swin Transform

    CN116416434A

  • Leaf vegetable harvesting method, system and equipment based on deep learning and storage medium

    CN118470525A

  • Remote sensing target detection method and device based on multi-scale and efficient attention mechanism

    CN119600274A

  • Cell instance segmentation method and system based on YOLOv11 improvement

    CN119888729A

Cited By

  • Image segmentation method based on space attention mechanism

    CN120374992A

  • Image segmentation method based on spatial attention mechanism

    CN120374992B

  • Improved soybean seed test system and method based on Transform-faster-Rcnn

    CN120431413A

  • Jersey cattle body condition scoring method and system based on tail root area segmentation

    CN120451714A

  • Production environment mobile phone illegal use detection method and system based on YOLOv12 optimization

    CN120689760A