A method for detecting cotton leaf disease
Patent Information
- Application Number
- CN202611235146.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-08-14
- Publication Date
- 2026-09-22
AI Technical Summary
[0006]针对现有棉花叶片病害检测技术存在大小病斑识别不均衡、复杂田间背景易引发漏检、误检,检测精度有待提升的缺陷,本发明的目的在于提供一种棉花叶片病害检测方法,均衡不同尺度病斑识别效果,降低田间环境带来的检测误差,提升病害整体检测精度
[0048]ARFF模块通过多分支特征融合方式增强不同尺度病害区域的空间特征表达能力,提高模型对小尺寸病斑和弱纹理病害特征的提取能力;D-SALA模块通过动态极性注意力机制增强关键病害区域特征响应,降低复杂背景对检测结果的影响,提高病害目标识别准确性;二者结合能够充分融合浅层细节信息与深层语义信息,提升模型在复杂环境下对棉花叶片病害目标的检测精度和稳定性,减少漏检和误检现象。
Smart Images

Figure CN122799286A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of crop leaf disease detection technology, and more particularly to a method for detecting cotton leaf diseases. Background Technology
[0002] Intelligent detection of crop leaf diseases is an important research direction in the field of smart agriculture and plant protection. Cotton, as an important economic crop, suffers from leaf diseases that affect normal plant growth, reduce yield and quality. Therefore, the development of automatic identification and accurate detection of cotton leaf diseases has significant application value.
[0003] With the development of machine vision and deep learning technologies, image-based intelligent detection methods have been widely applied in the field of crop disease identification. Traditional cotton disease detection mainly relies on manual inspection, which suffers from low detection efficiency, high labor costs, and strong subjectivity in disease judgment, making it difficult to meet the needs of rapid and accurate detection in large-scale agricultural production environments. In recent years, deep learning-based target detection methods can automatically learn the features of diseased areas, achieving disease category identification and location positioning, effectively improving detection efficiency.
[0004] However, existing methods for detecting cotton leaf diseases still have certain shortcomings: on the one hand, leaf texture, light changes, and background interference factors in complex background environments can easily affect the model's accurate localization of diseased areas; on the other hand, cotton lesions usually have significant size differences, weak edge features, and small area of early lesions. Existing detection models are insufficient in extracting multi-scale disease features and have limited ability to express features of key disease areas, which can easily lead to missed detections and false detections, resulting in room for improvement in detection accuracy and stability.
[0005] Therefore, how to enhance the model's ability to extract disease features at different scales and improve the detection accuracy of cotton leaf diseases under complex backgrounds has become an urgent problem to be solved in the field of intelligent disease detection. Summary of the Invention
[0006] To address the shortcomings of existing cotton leaf disease detection technologies, such as uneven identification of lesions of different sizes, the susceptibility to missed or false detections due to complex field backgrounds, and the need to improve detection accuracy, this invention aims to provide a cotton leaf disease detection method that balances the identification effect of lesions of different sizes, reduces detection errors caused by the field environment, and improves the overall detection accuracy of diseases.
[0007] To achieve the above objectives, the technical solution adopted by the present invention is as follows: The technical solution of the present invention will be described in detail with reference to specific embodiments. A method for detecting cotton leaf diseases includes the following steps:
[0008] S1: Acquire and label images of cotton leaf diseases to construct a cotton leaf disease detection image sample set; wherein, the cotton leaf disease image dataset includes images of healthy cotton leaves and images of cotton leaves with different disease types, the disease types including wilt, leaf roll, gray mold, leaf spot and wilt, and the image sample set is divided into training set, validation set and test set according to a preset ratio;
[0009] S2: A cotton leaf disease detection model was built based on the DEIM network, using HGNetv2 as the backbone network to output four-scale feature maps. , , , ;Will , The feature maps are fed into the ARFF adaptive receptive field fusion module. Enhanced features are generated after processing by the ARFF module. , Intermediate enhancement features are generated after processing by the ARFF module. ,Will Further input into the D-SALA dynamic spatial adaptive polar linear attention module generates enhanced features after dynamic polar attention calculation. ;Will and The input detection head completes the localization and classification of disease targets, and the enhanced features are used to achieve the localization and classification of disease targets, resulting in an improved detection model;
[0010] S3: Use the cotton leaf disease dataset constructed in step S1 to train the improved cotton leaf disease detection model in step S2.
[0011] S4: Conduct ablation experiments to verify the performance improvement effects of the ARFF module and the D-SALA module used alone and in combination, thus verifying the effectiveness of the trained model.
[0012] Furthermore, in the above technical solution, the cotton leaf disease detection model includes a feature extraction network, a feature enhancement fusion network, and a detection prediction network.
[0013] The feature extraction network uses HGNetv2 as the backbone network to extract features layer by layer from the input cotton leaf image, obtaining multi-level feature maps containing information at different spatial scales. By extracting feature information at different scales, rich semantic information and spatial detail information are provided for subsequent disease target detection.
[0014] Furthermore, the feature enhancement fusion network includes an ARFF adaptive receptive field fusion module. The ARFF module enhances the multi-scale feature information output by the backbone network. By constructing feature extraction paths with different receptive ranges, the network can acquire feature information at corresponding scales based on changes in the size of the lesion target, improving the model's representation ability for lesion regions of different sizes. Specifically, for early-stage lesion regions with smaller sizes and weaker texture features, the ARFF module can enhance the response of local detail features; for larger lesion regions, it can retain more complete spatial structural information, thereby improving the multi-scale lesion target detection effect.
[0015] The ARFF module uses convolutional branches of different scales and depths to extract features from the input features.
[0016] Three scale branches are set up, each using:
[0017]
[0018] For any scale branch, the RepDWLite structure is first used for feature extraction.
[0019] RepDWLite consists of: a horizontal depthwise convolution branch; a vertical depthwise convolution branch; a 3×3 depthwise convolution branch; and an identity mapping branch.
[0020] RepDWLite fusion adds the three spatial features and the identity mapping result element-wise:
[0021]
[0022] The enhanced spatial features at the i-th scale are obtained.
[0023] in: It is indicated that large receptive fields are simulated by using large horizontal and vertical convolutional kernels to extract spatial structural information such as lesion edges, textures, and shapes; Representation: Local spatial features; Indicates: Identity mapping feature.
[0024] Add a dynamic gating structure BranchSEGate after each scale branch.
[0025] Subsequently, the BranchSEGate dynamic gating mechanism was used to adjust the channel weights of features at different scales:
[0026]
[0027]
[0028] Fusing features enhanced at multiple scales:
[0029]
[0030] After normalization and activation operations, the ARFF module output is obtained.
[0031] Furthermore, the feature enhancement fusion network also includes a D-SALA dynamic spatial adaptive polar linear attention module. The D-SALA module, by introducing a dynamic spatial feature modeling mechanism, further optimizes the fused feature information, enabling the model to adaptively focus on key region features related to the disease target and reduce the impact of irrelevant background information on the detection results, thereby improving the ability to distinguish disease target features in complex background environments.
[0032] Features obtained by enhancing deep features P5 using the ARFF module As input to the D-SALA module:
[0033]
[0034] Where: B represents the input batch size; C represents the number of feature channels; H' and W' represent the spatial height and width, respectively.
[0035] As an input feature of the D-SALA module, it is defined as follows:
[0036]
[0037] The D-SALA module first expands the input features spatially and obtains query features, key features, and value features through linear mapping. Building upon this, a dynamic spatial location encoding mechanism is introduced to adaptively adjust the spatial information based on the current input feature size.
[0038]
[0039] in, Represents dynamic spatial location encoding. This represents the initial learnable space encoding. This represents a spatial interpolation operation. By dynamically generating spatial location information, attention calculations can adapt to changes in the location of disease targets at different scales.
[0040] Subsequently, the dynamic spatial location encoding is fused into the key features:
[0041]
[0042] Furthermore, to improve the model's ability to model the differences between the target disease area and the complex background area, the D-SALA module introduces a dynamic polarity mapping mechanism, which describes information about different response directions through positive and negative polarity features. The calculation process is as follows:
[0043] Q+=(ReLU(Qd))ρ; Q−=(ReLU(−Qd))ρ; K+=(ReLU(Kd))ρ; K−=(ReLU(−Kd))ρ.
[0044] Here, ρ represents the dynamic polarity response factor, used to control the degree of polarity feature enhancement. By simultaneously establishing positive and negative polarity response relationships, the network can enhance the effective features of the diseased area while suppressing background noise interference.
[0045]
[0046] in, This represents the enhanced features after dynamic gating. This represents the combined global and local features after fusion. This represents a dynamic gating feature. By dynamically adjusting the feature weights of different regions, information on key diseased areas can be further enhanced.
[0047] After the above processing, the D-SALA module outputs an enhanced feature map, which is then input into the subsequent detection head to complete the classification and localization of disease targets.
[0048] The ARFF module enhances the spatial feature representation of disease areas at different scales through multi-branch feature fusion, improving the model's ability to extract features of small-sized lesions and weakly textured diseases. The D-SALA module enhances the feature response of key disease areas through dynamic polarity attention mechanism, reducing the impact of complex backgrounds on detection results and improving the accuracy of disease target identification. The combination of the two can fully integrate shallow detail information and deep semantic information, improving the model's detection accuracy and stability of cotton leaf disease targets in complex environments, and reducing missed detections and false detections. Attached Figure Description
[0049] Figure 1 This is a flowchart of the method described in this invention;
[0050] Figure 2 Original images of cotton leaf diseases used in this invention;
[0051] Figure 3 This is a schematic diagram of the detection network used in the cotton leaf disease detection method of the present invention.
[0052] Figure 4 A comparison of the missed detection improvement effects between the DEIM network and the cotton leaf disease detection model of this invention;
[0053] Figure 5 This is a comparison chart showing the improvement in false detection between the DEIM network and the cotton leaf disease detection model of this invention. Detailed Implementation
[0054] To make the objectives, technical solutions, and advantages of the embodiments of this application clearer, the technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are some embodiments of this application, but not all embodiments.
[0055] The technical solution of the present invention will be further described in detail below through embodiments and in conjunction with the accompanying drawings.
[0056] See Figure 1 A method for detecting cotton leaf diseases includes the following steps:
[0057] S1: Acquire cotton leaf disease images and annotate the images to obtain corresponding disease category information and target location information, and construct a cotton leaf disease detection image sample set; acquire cotton leaf disease images and annotate them to construct a cotton leaf disease detection image sample set; wherein, the cotton leaf disease image dataset includes images of healthy cotton leaves and images of cotton leaves with different disease types, the disease types including wilt, leaf roll, gray mold, leaf spot and wilt, and divide the image sample set into training set, validation set and test set;
[0058] S2: Construct a cotton leaf disease detection model based on the DEIM network, using HGNetv2 as the DEIM backbone network to output four-scale feature maps. , , , Select , The feature map is input into the ARFF adaptive receptive field fusion module, which adopts a multi-branch image frame adaptive structure to realize multi-scale feature enhancement fusion, thereby improving the model's ability to extract multi-scale lesion features of cotton. The D-SALA dynamic spatial adaptive polar linear attention module is introduced to strengthen the expression of small lesion features and improve the detection accuracy of cotton leaf diseases. Based on the enhanced features, the disease target localization and classification are completed, resulting in an improved cotton leaf disease detection model.
[0059] S3: Use the cotton leaf disease dataset constructed in step S1 to train the improved cotton leaf disease detection model in step S2.
[0060] S4: Conduct ablation experiments to verify the performance improvement effects of the ARFF adaptive receptive field fusion module and the D-SALA dynamic spatial adaptive polar linear attention module when used alone and in combination, thus completing the effectiveness verification of the cotton leaf disease detection model after training in step S3.
[0061] The selected image data includes information from real agricultural scenarios such as changes in natural lighting, leaf occlusion, overlapping branches and leaves, and complex background environments. After screening and manual annotation, a cotton leaf disease dataset was formed, providing data support for the subsequent training and performance verification of disease detection models based on the DEIM network. Figure 2 As shown, Figure 2 The original image samples of cotton leaf diseases used in this invention are shown, including typical images of different types of cotton leaf diseases, which are used to reflect the morphological characteristics of leaf disease targets and the differences in background environment in actual scenarios.
[0062] The cotton leaf disease detection method uses an HGNetv2 feature extraction network as the backbone network in step S2, and the backbone network sequentially outputs four-scale feature maps. , , , Feature map Output of ARFF adaptive receptive field fusion module Feature map First, the ARFF adaptive receptive field fusion module is connected, followed by the D-SALA dynamic spatial adaptive polarity linear attention module to perform feature enhancement and fusion processing and output. When training the improved cotton leaf disease detection model in step S3, the training parameters used include: setting the training cycle to 200 rounds, the batch size to 8, and the input image pixel to 320×320.
[0063] and As input features of the ARFF module, when the input features are hour: ;
[0064] When the input features are hour: .
[0065] For ease of description, the input features of the ARFF module will be uniformly referred to as follows: .
[0066] This embodiment sets up three scale branches, which are respectively:
[0067]
[0068] The i-th scale branch is represented as:
[0069]
[0070] Three convolutional kernels of different scales, among which: =3, =5, =7; Different scale branches are used to obtain spatial features within different receptive field ranges.
[0071] For any scale branch i, feature extraction is first performed using the RepDWLite structure.
[0072] RepDWLite consists of: a horizontal depthwise convolution branch; a vertical depthwise convolution branch; a 3×3 depthwise convolution branch; and an identity mapping branch.
[0073] For the lateral depthwise convolution calculation, for the scale parameter ki, the following is first performed:
[0074]
[0075] Where DConv represents depthwise convolution.
[0076] Vertical depthwise convolution calculations input the results of horizontal convolutions into the vertical convolution branch:
[0077]
[0078] Further obtain spatial information in the vertical direction. Therefore:
[0079]
[0080] Local spatial feature computation, while utilizing 3×3 depthwise convolution branches:
[0081]
[0082] Identity mapping preserves the original information, utilizing the Identity branch:
[0083]
[0084] RepDWLite fusion adds the above three spatial features and the identity mapping result element-wise:
[0085]
[0086] The enhanced spatial features at the i-th scale are obtained.
[0087] Branches of different scales need to be merged to unify the output channels of each branch. Channel conversion is performed using 1×1 convolution.
[0088]
[0089] in:
[0090]
[0091] i∈{1,2,3} represents the output features of convolution branches at different scales.
[0092] A dynamic gating structure, BranchSEGate, is added after each scale branch. For the i-th scale feature: Perform global average pooling.
[0093] Global information compression, the statistical information of the c-th channel is represented as follows:
[0094]
[0095] c represents the feature channel index, and H' and W' represent the spatial dimensions of the input features after convolution.
[0096] Obtain the channel description vector:
[0097]
[0098] Subsequently, the channel description vector Si is input into a two-layer fully connected network to model the channel relationships.
[0099] The first and second layers are respectively:
[0100]
[0101]
[0102] in:
[0103]
[0104]
[0105] r represents the channel compression ratio.
[0106] Obtain channel weights using the Sigmoid function:
[0107]
[0108] in: The range of each element: 0 < i<1 indicates the importance of each channel in the current scale branch.
[0109] The original scale features are recalibrated using the obtained weights:
[0110]
[0111] in: This indicates element-wise multiplication in the channel dimension.
[0112] The dynamically adjusted features at multiple scales are then concatenated along the channel direction using the following formula:
[0113]
[0114] If the number of output channels is not divisible by the number of branches, then add a tail branch: ,final:
[0115]
[0116] The fusion characteristics were obtained:
[0117]
[0118] Group Normalization is applied to the fused features:
[0119]
[0120] Activation function via GELU:
[0121]
[0122] When processing hour, When processing hour, This yields the final enhanced feature output.
[0123] After processing by the ARFF module described above, an enhanced feature map incorporating multi-scale spatial information is obtained. This includes features targeting deep features. Features obtained after ARFF module enhancement The feature is further optimized by inputting the D-SALA dynamic spatial adaptive polar linear attention module.
[0124] Specifically, the output of the ARFF module Layer-enhanced features are used as input to the D-SALA module:
[0125]
[0126] Where: B represents the input batch size; C represents the number of feature channels; H' and W' represent the spatial height and width, respectively.
[0127] As an input feature of the D-SALA module, it is defined as follows:
[0128]
[0129] Expanding the spatial dimensions:
[0130]
[0131] get: Where: N = H' × W', representing the number of spatial locations.
[0132] Generate query features and gated features through linear mapping:
[0133]
[0134] in: Indicates query characteristics; This indicates a dynamic gating feature.
[0135] Then key and value are generated:
[0136]
[0137] in: .
[0138] Dynamic spatial location coding enhancement, initial learnable spatial coding is Dynamically adjust based on the current input space size:
[0139]
[0140] in: This represents dynamic spatial location encoding, which is then fused into the Key: ,in: This represents the enhanced key feature after fusing dynamic spatial location information.
[0141] Dynamic polarity mapping first requires calculating the dynamic normalization scale:
[0142]
[0143] in: Indicates dynamic scaling parameters; This represents a learnable scale adjustment parameter; This represents a nonlinear activation function used to ensure that the scaling parameter is positive.
[0144] Then: , ,in: This indicates the query features after dynamic scaling. This represents the key features after incorporating dynamic spatial locations and adjusting for scale.
[0145] Formula for calculating dynamic polarity response factor:
[0146]
[0147] in: Indicates the dynamic polarity response factor; Indicates the polarity adjustment coefficient; Indicates the learnable polarity parameter; This represents the Sigmoid activation function.
[0148] The query features and key features, after dynamic scaling normalization, are mapped to positive and negative polarity spaces, respectively. For the query features... The positive polarity characteristic is calculated as follows:
[0149]
[0150] The negative polarity characteristic is calculated as follows:
[0151]
[0152] For key features The positive polarity characteristic is calculated as follows:
[0153]
[0154] The negative polarity characteristic is calculated as follows:
[0155]
[0156] in Used to control the degree of polarity response enhancement.
[0157] Combining positive and negative polarity query features, similarity polarity query is represented as:
[0158]
[0159] Complementary polarity lookup is represented as:
[0160]
[0161] Simultaneously, the key features are combined to obtain the polar bond representation:
[0162]
[0163] Calculation of similarity polarity relationship, calculation of similarity polarity normalization factor:
[0164]
[0165] in, This represents the polar bond features after averaging along the sequence dimension. This represents a minimal constant to prevent the denominator from being zero.
[0166] Divide the value feature V equally along the channel dimension and Two parts.
[0167] Calculate similarity polarity key value aggregation information:
[0168]
[0169] Similar polarity enhancement features were obtained:
[0170]
[0171] The complementary polarity normalization factor is:
[0172]
[0173] Key-value aggregation is represented as:
[0174]
[0175] The complementary polarity enhancement feature is obtained:
[0176]
[0177] The two polarity relationship features are fused:
[0178]
[0179] in This represents the polarity attention characteristics after fusion.
[0180] Remapping value features in sequence form to spatial structure features:
[0181]
[0182] V represents the value characteristic during the attention calculation process; This represents a two-dimensional value feature after spatial dimensional rearrangement. This indicates a spatial dimension transformation operation.
[0183] Two-dimensional spatial value features are input into the depthwise convolution branch, and spatial neighborhood information is extracted through local convolution operations, then flattened to align with the sequence dimensions:
[0184]
[0185] in Indicates local spatial enhancement features; This indicates a depthwise separable convolution operation.
[0186] The obtained global polarity attention features and Perform element-by-element fusion:
[0187]
[0188] in This represents the combined global and local features after fusion.
[0189] Utilizing dynamic gating features The fused features are dynamically adjusted using the following calculation formula:
[0190]
[0191] in This indicates element-wise multiplication.
[0192] The gated and enhanced features are then input into the linear projection layer:
[0193]
[0194] in Y represents the output mapping matrix; Y represents the sequence form characteristics of the D-SALA module output.
[0195] Output features preserve sequence structure:
[0196]
[0197] Where: N = H × W, This indicates the output sequence characteristics of the D-SALA module.
[0198] Spatial dimension recovery of output sequence features:
[0199]
[0200] The final two-dimensional spatial features are obtained:
[0201]
[0202] in This represents the enhanced feature map that is the final output of the D-SALA module; this feature map is input into subsequent detection heads to complete target localization and category prediction.
[0203] The cotton leaf disease detection method, the training strategy for the ablation experiment used in step S4 includes the following steps:
[0204] S41: Input the cotton leaf disease detection dataset, and obtain the target detection model under different structural configurations by optimizing the model parameters;
[0205] S42: The ablation experiment was trained in four groups. Group 1 used the original DEIM detection network as the base model. Group 2 added the ARFF module to the original DEIM detection network to verify the effectiveness of the ARFF module in enhancing multi-scale disease features. Group 3 added the D-SALA module to the original DEIM detection network to verify the improvement effect of the D-SALA module on the feature relationship modeling ability in complex scenarios. Group 4 used an improved DEIM network with both the ARFF and D-SALA modules added to verify the improvement effect of the synergistic effect of each improved module on the detection performance of cotton leaf diseases.
[0206] S43: Statistical analysis will be performed on the detection results of the following four groups of ablation experiments to verify the effectiveness of the improved cotton leaf disease detection model.
[0207] 1 × × 87.8 89.4 88.2 77.4 2 √ × 89.5 89.9 90.0 79.8 3 × √ 90.8 89.7 88.9 78.0 4 √ √ 92.5 91.7 92.1 81.4
[0208] Ablation test data show that, compared with the basic model, the complete detection model has an accuracy of 92.5%, a recall of 91.7%, an AP50 accuracy of 92.1%, and an AP50-95 accuracy of 81.4%. The above experimental data prove that the ARFF and D-SALA modules added in this invention can synergistically improve the overall detection accuracy of cotton leaf diseases.
[0209] like Figure 3 As shown, the detection network used in this invention mainly includes a feature extraction network, a multi-scale feature fusion structure, an ARFF adaptive receptive field fusion module, a D-SALA dynamic spatial adaptive polarity linear attention module, and a detection prediction module. The input cotton leaf image is processed by the feature extraction network to obtain multi-scale feature information. Then, the ARFF module enhances the spatial feature representation of disease areas at different scales, and the D-SALA module further extracts key disease area features. Finally, the detection prediction module outputs disease category information and target bounding box location information, achieving the identification and localization of cotton leaf disease targets.
[0210] like Figure 4 As shown, by comparing the detection results of the DEIM network and the detection model of this invention in the cotton leaf disease detection task, the differences in the identification of disease targets between the two models can be observed. The model of this invention enhances the multi-scale disease feature expression through the ARFF module and improves the feature response capability of key areas by combining the D-SALA module, which can effectively reduce the problem of some small-sized lesions and weak textured disease targets not being detected, and improve the detection completeness of disease targets.
[0211] like Figure 5 As shown, by comparing the detection results of the DEIM network and the detection model of this invention, it can be found that the model of this invention has better target discrimination ability under complex background conditions. By enhancing the features of key regions through the D-SALA module and combining it with a multi-scale feature fusion structure to achieve full fusion of shallow detail information and deep semantic information, false detection caused by background interference can be effectively reduced, and the accuracy of cotton leaf disease target identification can be improved.
[0212] This solution effectively addresses the problems of insufficient extraction of small lesions in conventional detection networks and the tendency for missed or incorrect detections in complex field backgrounds. It is adapted to field scenarios with high light levels, weed cover, and lesion sizes that vary greatly, accurately distinguishing healthy leaves from cotton diseases, and achieving automatic disease identification and precise lesion location. This provides comprehensive technical support for rapid inspection and early prevention and control of cotton diseases.
Claims
1. A method for detecting cotton leaf diseases, characterized in that, Includes the following steps: S1: Obtain and label images of cotton leaf diseases to construct a cotton leaf disease detection image sample set; wherein, the cotton leaf disease image dataset includes images of healthy cotton leaves and images of cotton leaves with different disease types, including wilt, leaf roll, gray mold, leaf spot and wilt, and the image sample set is divided into training set, validation set and test set; S2: Construct a cotton leaf disease detection network based on the DEIM end-to-end target detection network, using HGNetv2 as the backbone feature extraction network, to perform multi-stage feature extraction on the input cotton leaf disease image and output a four-scale feature map. , , as well as ; Selecting the fourth-scale feature map and fifth-scale feature map As an enhanced object, The feature input ARFF adaptive receptive field fusion module constructs spatial receptive regions of different scales through a multi-branch image frame adaptive structure, extracts the features of disease areas in different spatial ranges, and realizes multi-scale lesion feature enhancement fusion by utilizing the information interaction between receptive fields of different scales. Will The features are input into the ARFF module for high-level semantic feature enhancement. Then The D-SALA Dynamic Spatial Adaptive Polarity Linear Attention module is used to enhance the modeling of key disease areas in complex field backgrounds through dynamic spatial location encoding, positive and negative polarity feature interaction, and dynamic polarity response adjustment mechanism. in, Enhanced features are generated after processing by the ARFF module. , Enhanced features are generated through joint processing by the ARFF module and the D-SALA module. This will enhance the features and The input detection head completes the location and category prediction of disease targets, and constructs an improved cotton leaf disease detection network; S3: Use the cotton leaf disease dataset constructed in step S1 to train the improved cotton leaf disease detection network constructed in step S2, and obtain the trained cotton leaf disease detection model. S4: The trained cotton leaf disease detection model was used to detect images in the test set. Ablation experiments were conducted by setting different module combinations to verify the effects of introducing the ARFF adaptive receptive field fusion module and the D-SALA dynamic spatial adaptive polar linear attention module individually and in combination on the performance of the detection model. This was used to verify the effectiveness of the constructed cotton leaf disease detection model.
2. The method for detecting cotton leaf diseases according to claim 1, characterized in that, In step S2, the DEIM network uses HGNetv2 as the backbone feature extraction network. Through a multi-stage feature extraction structure, it performs step-by-step feature encoding on the input cotton leaf disease image to obtain feature representations at different spatial scales. Where I represents the input image of cotton leaf disease; This represents the HGNetv2 feature extraction process; , , as well as These represent the feature maps output at different scales; among them, the fourth scale feature map is selected. and the fifth-scale feature map As an enhanced feature input, it is used to further improve the feature representation ability of disease target areas at different scales.
3. The method for detecting cotton leaf diseases according to claim 1, characterized in that, The ARFF module is used to process the mesoscale features output by the backbone network. Perform multi-scale spatial feature enhancement processing; Specifically, the output of the backbone network Layer features are used as input to the ARFF module: As an input feature of the ARFF module, it is defined as follows: Input feature map The input is fed into multiple parallel feature extraction branches within the ARFF module. In this embodiment, three scale branches are set up, each employing the following methods: Acquire spatial features within different receptive field ranges; For the i-th scale branch, spatial feature enhancement is performed using the RepDWLite structure: Subsequently, the BranchSEGate dynamic gating mechanism was used to adjust the channel weights of features at different scales: Fusing features enhanced at multiple scales: If the number of output channels is not divisible by the number of branches, then add a tail branch: ,final: After normalization and activation operations, the ARFF module output is obtained: The final enhanced mesoscale features are obtained: in, This represents the mesoscale feature map enhanced by the ARFF module, which is used for subsequent feature fusion and detection.
4. The method for detecting cotton leaf diseases according to claim 3, characterized in that, The ARFF module is further used to process the deep features output by the backbone network. Enhancement processing was performed to improve the ability to express the characteristics of cotton leaf diseases under complex natural environments; Specifically, the output of the backbone network Layer features are used as input to the ARFF module: As an input feature of the ARFF module, it is defined as follows: Input feature map The input is fed into multiple parallel feature extraction branches within the ARFF module. In this embodiment, three scale branches are set up, each employing the following methods: Acquire spatial features within different receptive field ranges; For the i-th scale branch, spatial feature enhancement is performed using the RepDWLite structure: Subsequently, the BranchSEGate dynamic gating mechanism was used to adjust the channel weights of features at different scales: Fusing features enhanced at multiple scales: If the number of output channels is not divisible by the number of branches, then add a tail branch: ,final: After normalization and activation operations, the ARFF module output is obtained: The final enhanced deep features are obtained: in, This represents the deep feature map enhanced by the ARFF module, and serves as the input feature for the subsequent D-SALA module.
5. The method for detecting cotton leaf diseases according to claim 1, characterized in that, The D-SALA module is used to enhance the deep features output by the ARFF module. Further feature optimization was carried out to enhance the spatial feature representation ability of cotton leaf diseases under complex natural environments; Specifically, the enhanced features output by the ARFF module are used as input to the D-SALA module: Spatial dimensional expansion of the input features: Obtain sequence-form features: Where N=H' W' represents the number of spatial locations; Query features, key features, value features, and dynamic gating features are obtained through linear mapping: A dynamic spatial location coding mechanism is introduced to generate dynamic location codes based on the input feature size: And integrate into key features: in: This represents the enhanced key feature after fusing dynamic spatial location information; Positive and negative polarity features are obtained through dynamic scaling and polarity mapping mechanisms, and similar and complementary polarity attention features are calculated. The polarity bond features after averaging along the sequence dimension are represented as follows: The value feature V is divided equally along the channel dimension. and Two parts; By fusing global polarity enhancement features with local spatial features, the spatial structure of the value features is first restored: V represents the value characteristic during the attention calculation process; This represents a two-dimensional value feature after spatial dimensional rearrangement. This indicates a spatial dimension transformation operation; Extracting local spatial information through depthwise convolution: in Indicates local spatial enhancement features; This represents a depthwise separable convolution operation; Finally, the D-SALA output features are obtained through output mapping: And restore the spatial dimension: in, This is the enhanced feature map output by the D-SALA module.
6. The method for detecting cotton leaf diseases according to claim 1, characterized in that, The ablation experiment training strategy includes the following steps: S41: Input cotton leaf disease dataset, and obtain target detection models under different structural configurations by optimizing model parameters; S42: The ablation experiment was trained in four groups. Group 1 used the original DEIM detection network as the base model. Group 2 added the ARFF module to the original DEIM detection network to verify the effectiveness of the ARFF module in enhancing multi-scale disease features. Group 3 added the D-SALA module to the original DEIM detection network to verify the improvement effect of the D-SALA module on the feature relationship modeling ability in complex scenarios. Group 4 used an improved DEIM network with both the ARFF and D-SALA modules added to verify the improvement effect of the synergistic effect of each improved module on the detection performance of cotton leaf diseases. S43: Statistical analysis was performed on the detection results of the above four groups of ablation experiments to verify the effectiveness of the improved cotton leaf disease detection model.