Field weed detection method based on improved YOLOv11 network

The improved YOLOv11 network addresses low detection accuracy in complex agricultural scenes by enhancing feature extraction and occlusion handling, resulting in precise weed detection.

CN120318690APending Publication Date: 2025-07-15XIDIAN UNIV
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510461931.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

The prior art has low detection accuracy in complex field scenarios with dense overlap and obscured small-sized weeds, especially poor detection of small-sized weeds and obscured weeds.

Method used

Using the improved YOLOv11 network model, improvements in the reversible feature extraction columnar subnet, feature fusion subnet and detection subnet are improved through reversible feature extraction columnar subnet, including the reversible feature extraction columnar subnet to replace the backbone feature extraction subnet, and the separable nuclear attention module and occlusion perception attention module are added to improve the feature extraction and occlusion area processing capabilities.

Benefits of technology

It effectively improves the detection accuracy of small-sized weeds and blocked weeds, and improves the detection effect of the model in complex field scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318690A_ABST
    Figure CN120318690A_ABST
Patent Text Reader

Abstract

The invention provides a field weed detection method based on an improved YOLOv11 network. The method comprises the following implementation steps: acquiring a training sample set and a test sample set; the YOLOv11 network model is improved, and iterative training is carried out on the YOLOv11 network model; and obtaining a field weed detection result. According to the invention, the reversible feature extraction columnar sub-network gradually performs deep feature extraction on each training sample through a plurality of feature extraction branches, so that the detection of large-size weeds is ensured, and meanwhile, the effective detection of shielded small-size weeds is realized; the separable kernel attention module weights the importance of different regions of the feature map, so that the model better pays attention to the information of the target region; and meanwhile, the occlusion perception attention module enhances the feature response of the occluded region in the fusion feature map, so that the model can more effectively process the occluded information, and the detection precision is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of image processing, and relates to a method for detecting field weeds based on an improved YOLOv11 network, which can be used in the field of smart agriculture. Background Art

[0002] Field weed management is a key link to ensure the yield and quality of crops. Weeds compete with crops for nutrients, water, light, and space resources, seriously affecting the growth of crops, resulting in reduced crop yields and degraded quality. Accurately and efficiently detecting field weeds is of great significance for timely taking targeted weeding measures, reducing the use of chemical herbicides, lowering agricultural production costs, and protecting the ecological environment.

[0003] Field weed detection is a technology for predicting the species information, location information, and confidence information of weeds in the field. The performance of field weed detection methods based on image processing and machine learning still faces challenges of misjudgment and missed judgment in complex field scenarios with densely overlapping and occluded small-sized weeds. How to improve the detection accuracy of weeds in complex environments is the focus of field weed detection methods. For example, a patent application with the application publication number CN118314438A and the name "A Method for Detecting Field Weeds Based on YOLOv8 Optimization" discloses a method for detecting field weeds based on YOLOv8 optimization. In this method, the convolution in the repeatedly stacked modules of the backbone network C2F module is replaced with a dynamic snake convolution Dynamic Snake Convolution, effectively reducing the loss of fine-grained information, and a plug-and-play lightweight channel attention mechanism SE module is added to the YOLO-head detection head to enable the network to better focus on important feature channels, thereby improving the detection accuracy of the model. However, the SE module only weights through channel attention, lacks attention to occluded areas, and the backbone feature extraction network after replacing Dynamic Snake Convolution with the C2F module has poor ability to extract deep features of weeds, affecting the further improvement of detection accuracy in complex field scenarios containing densely overlapping and occluded small-sized weeds. Summary of the Invention

[0004] The purpose of the present invention is to propose a method for detecting field weeds based on an improved YOLOv11 network in view of the deficiencies of the above-mentioned existing technologies, so as to solve the problem of low detection accuracy in complex field scenarios containing densely overlapping and occluded small-sized weeds existing in the prior art.

[0005] To achieve the above purpose, the technical solutions adopted by the present invention include the following steps:

[0006] (1) Obtain a training sample set and a test sample set:

[0007] Preprocess the obtained L RGB images including multiple weed categories, label the weeds in the preprocessed images, and then form a training sample set with the preprocessed M images and their labels, and form a test sample set with the remaining L - M preprocessed images and their labels, where L ≥ 2400.

[0008] (2) Improve the YOLOv11 network model:

[0009] Obtain the YOLOv11 network including a cascaded backbone feature extraction sub-network, a feature pyramid module, and a detection head module, and replace the backbone feature extraction sub-network with a reversible feature extraction columnar sub-network to achieve the improvement of the backbone feature extraction sub-network; add a separable kernel attention module at the input end of the feature pyramid module to form a feature fusion sub-network to achieve the improvement of the feature pyramid module; add an occlusion-aware attention module at the input end of the detection head module to form a detection sub-network to achieve the improvement of the detection head module, and then obtain the improved YOLOv11 network model W.

[0010] (3) Iteratively train the improved YOLOv11 network model:

[0011] Iteratively train the improved YOLOv11 network model W with the training sample set to obtain the trained network model W. * ;

[0012] (4) Obtain the detection results of field weeds:

[0013] Use the test sample set as the input of the trained network model W * to perform forward propagation and obtain the type information, location information, and confidence information of the weeds in each test sample.

[0014] Compared with the prior art, the present invention has the following advantages:

[0015] 1. In the process of iteratively training the improved YOLOv11 network model and obtaining the detection results, the reversible feature extraction columnar sub-network gradually performs deep feature extraction on each training sample through multiple feature extraction branches, effectively detecting occluded small-sized weeds while ensuring the detection accuracy of large-sized weeds, avoiding the defect that the backbone feature extraction sub-network in the prior art has poor ability to extract deep features of weeds, and effectively improving the detection accuracy of weeds.

[0016] 2. In the separable kernel attention module of the feature fusion sub-network of the present invention, the importance of different regions of the feature map is weighted, enabling the model to better focus on the information in the target region; and in the occlusion-aware attention module of the detection sub-network, the feature response in the occluded region of the fused feature map is enhanced, enabling the model to more effectively process the occluded information and further improving the detection accuracy. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] Figure 1 It is a flowchart for the implementation of the present invention.

[0018] Figure 2 It is a schematic structural diagram of the present invention based on an improved YOLOv11 network model.

[0019] Figure 3 It is a schematic structural diagram of the feature extraction branch in the reversible feature extraction column sub-network of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0020] The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.

[0021] Refer to Figure 1 , the present invention includes the following steps:

[0022] Step 1) Obtain a training sample set and a test sample set:

[0023] Preprocess the obtained L RGB images including multiple weed categories, label the weeds in the preprocessed images, then form a training sample set with the preprocessed M images and their labels, and form a test sample set with the remaining L - M preprocessed images and their labels;

[0024] Among them, the specific steps for preprocessing the L RGB images are as follows:

[0025] (1a) Perform data augmentation on each RGB image. First, randomly select Q images for vertical mirror flipping, and then randomly select P images for clockwise rotation by a random angle θ to obtain the enhanced RGB image I'(R, G, B), where θ ≤ 45°, Q ≤ 1800, and P ≤ 2000;

[0026] (1b) Perform a normalization operation on the enhanced RGB image I'(R, G, B), and its operation formula is:

[0027]

[0028] Among them, I'(R, G, B) represents the enhanced RGB image, μ(R, G, B) represents the mean of the three channels in the enhanced RGB image, σ(R, G, B) represents the standard deviation of the three channels in the enhanced RGB image, and I”(R, G, B) is the preprocessed RGB image;

[0029] Then, label the preprocessed image I”(R, G, B) to obtain L corresponding labels. Combine the labeled M images and the labels corresponding to each image to form a training sample set, and combine the remaining L - M images and labels to form a test sample set; in this embodiment, L = 4500 and M = 3000.

[0030] Step 2) Improve the YOLOv11 network model:

[0031] (2a) The original YOLOv11 network model includes a cascaded backbone feature extraction sub-network, a feature pyramid module, and a detection head module, where:

[0032] The backbone feature extraction sub-network includes N cascaded feature extraction modules, and each feature extraction module includes a cascaded standard convolution module and a cross-stage residual convolution module; and a feature output layer is connected to the output end of each feature extraction module, where N ≥ 3;

[0033] The feature pyramid module includes N feature fusion branches cascaded with N feature output layers respectively. Each feature fusion branch includes a cascaded standard convolution module, a bilinear interpolation module, and a feature splicing module, and the output end of the standard convolution module is connected to the output end of the bilinear interpolation module with a residual connection;

[0034] The detection head module includes multiple parallel multi-scale classification branch decoupling heads and multi-scale regression branch decoupling heads,

[0035] In this embodiment, N = 5.

[0036] (2b) Refer to Figure 2 and Figure 3 , the improved YOLOv11 network model includes a cascaded reversible feature extraction columnar sub-network, a feature fusion sub-network, and a detection sub-network, where:

[0037] The reversible feature extraction columnar sub-network includes K parallel feature extraction branches, and a cross-branch reversible interaction module is loaded between adjacent branches; each branch includes S depthwise separable convolution modules and S cross-stage residual convolution modules connected alternately, and the output end of the s-th depthwise separable convolution module is connected to the output end of the s-th cross-stage residual convolution module with a residual connection, where K = 5 and S = 5;

[0038] Feature fusion sub-network, where the separable kernel attention module includes a cascaded depthwise separable convolution module and a standard convolution module;

[0039] Detection sub-network, where the occlusion-aware attention module includes a cascaded depthwise separable convolution module and an occlusion enhancement module, and the input end and the output end of the depthwise separable convolution module are connected with residuals.

[0040] Step 3) Iteratively train the improved YOLOv11 network model:

[0041] (3a) Initialize the iteration number as t, the maximum iteration number as T, T≥300, and the weights ω in the current improved YOLOv11 network model W t and let t = 1. In this embodiment, T = 500; t

[0042] (3b) The training samples are sent into the feature extraction branch of the k-th reversible structure. The depthwise separable convolution module in the k-th reversible feature extraction branch performs feature extraction at different scales on each training sample; the cross-stage residual convolution module performs feature enhancement on the extracted multi-scale feature maps; while the enhanced feature maps and the feature maps at different scales extracted by the depthwise separable convolution module are sent into the depthwise separable convolution module at the next level of this branch for further feature extraction, they are also sent into the cross-branch reversible interaction module. The cross-branch reversible interaction module pools the enhanced feature maps to obtain the pooled enhanced feature maps, which are sent into the previous depthwise separable convolution module of the (k + 1)-th branch to achieve reversible transfer of features. Finally, the S cross-stage residual convolution modules of the K-th branch output S feature maps at different scales; in the prior art, the backbone feature extraction sub-network adopts a cascaded single-branch structure, and there is a problem of feature information loss during the network transmission. At the same time, the depth feature extraction ability of the single-branch structure is poor, which in turn affects the detection accuracy. The reversible feature extraction columnar sub-network uses the feature extraction branch with a reversible structure as a unit to transmit feature information, which not only ensures feature decoupling but also ensures that the feature information is not lost during the network transmission. The entire network includes multiple feature extraction branches, and the feature extraction branches are connected by cross-branch reversible interaction modules. By repeatedly sending the input into the feature extraction branches, the texture details and semantic information are gradually separated, and then deeper feature information is extracted, effectively improving the detection progress.

[0043] (3c) The depthwise separable convolution module in the separable kernel attention module extracts region importance feature information from the S feature maps of different scales output by the Kth branch. The standard convolution module performs a weighting operation on the feature maps from which the region importance feature information is extracted according to the extracted region importance feature information, obtaining S weighted feature maps of different scales. The standard convolution module and the bilinear interpolation module in the feature pyramid module perform downsampling and upsampling on the S weighted feature maps of different scales respectively; the feature concatenation module concatenates the downsampling and upsampling results to obtain a multi-scale fusion feature map. The prior art directly sends the multi-scale feature maps output by the backbone feature extraction subnetwork into the feature pyramid module for fusion, treating all regions including the target weed and the background equally, making it easy for the weed target feature information in the fusion feature map to be contaminated with background feature information, resulting in the model being unable to effectively learn and recognize the weed target features, especially having a greater impact on small-sized weeds, and thus leading to a low detection accuracy of the model for small-sized weeds. The separable kernel attention module weights the importance of different regions of the feature map before the feature maps of different scales are sent into the feature pyramid module for fusion, helping the model distinguish the target weed region from the background region, enabling the model to better focus on the feature information of the target weed region, and thus improving the detection accuracy.

[0044] (3d) The depthwise separable convolution module in the occlusion-aware attention module extracts occlusion relationship features from the multi-scale fusion feature map. The multi-scale fusion feature map output by the feature fusion subnetwork and the occlusion relationship features are fed as inputs into the occlusion enhancement module. The occlusion enhancement module enhances the input according to the extracted occlusion relationship features to obtain an occlusion-aware enhanced multi-scale fusion feature map. The multi-scale classification branch decoupling head in the detection head module predicts the class information and confidence information of the occlusion-aware enhanced multi-scale fusion feature map, and at the same time the multi-scale regression branch decoupling head predicts the position information of the occlusion-aware enhanced multi-scale fusion feature map, obtaining the prediction result Y of the weed in each training sample. m , where different combinations of the multi-scale classification branch decoupling head and the multi-scale regression branch decoupling head predict the feature maps of different sizes respectively; the prior art directly sends the multi-scale fusion feature map into the detection head module for classification and regression prediction. The features in the occluded region of the feature map are severely interfered by the occluder, making it impossible for the model to effectively identify and process the feature information of the occluded weed, and thus leading to a low detection accuracy of the model for the occluded weed. The occlusion-aware attention module enhances the feature response in the occluded region, dynamically adjusts the focus in the feature map, increases the weight of the occluded region in the feature map, enabling the model to more effectively process the occluded information, and thus improving the detection accuracy.

[0045] (3e) Through the label y of each training sample mand the corresponding predicted result Y m Calculate W t of the loss value Loss t , and then through Loss t update the weight ω t to obtain the improved YOLOv11 network model W after this iteration t , where:

[0046] Loss t = λ loc Loss loc t + λ cls Loss cls t

[0047]

[0048] where, Loss loc t , λ cls respectively represent the localization loss and the classification loss, and λ loc , λ cls respectively represent the weight coefficients of Loss loc t , λ cls ; GIoU(·) represents the Generalized Intersection over Union loss;

[0049] (3f) Judge whether t = T holds. If so, obtain the trained network model W * , otherwise, let t = t + 1, W t = W, and execute step (3b);

[0050] Step 4) Obtain the field weed detection result:

[0051] Use the test sample set as the input of the trained network model W * to perform forward propagation, and obtain the weed type information, location information, and confidence information in each test sample.

Claims

1. A method for detecting field weeds based on an improved YOLOv11 network, characterized in that, It includes the following steps: (1) Obtain a training sample set and a test sample set: Preprocess the acquired L RGB images including multiple weed categories, label the weeds in the preprocessed images, then form a training sample set with the preprocessed M images and their labels, and form a test sample set with the remaining L - M preprocessed images and their labels, where L ≥ 2400. (2) Improve the YOLOv11 network model: Obtain the YOLOv11 network including a cascaded backbone feature extraction sub-network, a feature pyramid module, and a detection head module, and replace the backbone feature extraction sub-network with a reversible feature extraction columnar sub-network to achieve the improvement of the backbone feature extraction sub-network; add a separable kernel attention module at the input end of the feature pyramid module to form a feature fusion sub-network to achieve the improvement of the feature pyramid module; add an occlusion-aware attention module at the input end of the detection head module to form a detection sub-network to achieve the improvement of the detection head module, and obtain the improved YOLOv11 network model W; (3) Iteratively train the improved YOLOv11 network model: Iteratively train the improved YOLOv11 network model W with the training sample set to obtain the trained network model W * ; (4) Obtain the detection results of field weeds: Use the test sample set as the input of the trained network model W * to perform forward propagation, and obtain the weed type information, location information, and confidence information in each test sample.

2. The method according to claim 1, wherein The preprocessing of the obtained L RGB images including multiple weed categories in step (1) is achieved as follows: Perform random mirror flipping in the vertical direction on each RGB image, and perform random rotation on the mirror-flipped image to achieve data augmentation of the RGB images, and then normalize the data-augmented images to obtain L preprocessed RGB images.

3. The method according to claim 1, wherein The YOLOv11 network model described in step (2), where: The backbone feature extraction sub-network includes N cascaded feature extraction modules, and each feature extraction module includes a cascaded standard convolution module and a cross-stage residual convolution module; and a feature output layer is connected to the output end of each feature extraction module, where N≥3; The feature pyramid module includes N feature fusion branches respectively cascaded with N feature output layers. Each feature fusion branch includes a cascaded standard convolution module, a bilinear interpolation module, and a feature splicing module, and the output end of the standard convolution module is connected to the output end of the bilinear interpolation module in a residual manner; The detection head module includes multiple parallel multi-scale classification branch decoupling heads and multi-scale regression branch decoupling heads.

4. The method according to claim 3, characterized in that, The improved YOLOv11 network model W described in step (2), where: The reversible feature extraction columnar sub-network includes K parallel feature extraction branches, and a cross-branch reversible interaction module is loaded between adjacent branches; each branch includes S depthwise separable convolution modules and S cross-stage residual convolution modules connected alternately, and the output end of the s-th depthwise separable convolution module is connected to the output end of the s-th cross-stage residual convolution module in a residual manner, where K≥5 and S≥3; The feature fusion sub-network, and the separable kernel attention module therein includes a cascaded depthwise separable convolution module and a standard convolution module; The detection sub-network, and the occlusion-aware attention module therein includes a cascaded depthwise separable convolution module and an occlusion enhancement module, and the input end of the depthwise separable convolution module is connected to the output end in a residual manner.

5. The method according to claim 4, wherein The iterative training of the improved YOLOv11 network model described in step (3) is achieved as follows: (3a) Initialize the number of iterations as t, the maximum number of iterations as T, where T ≥ 300, and the weights of the current improved YOLOv11 network model W t are ω t , and set t = 1; (3b) The reversible feature extraction columnar sub-network performs deep feature extraction on each training sample; Feature The fusion sub-network performs multi-scale feature fusion on the extracted deep feature information after weighting according to the importance of different regions; The detection sub-network performs occlusion relationship feature enhancement on the fused multi-scale fused feature map and then makes a prediction to obtain the prediction result Y of the weeds in each training sample m ; (3c) Through the label y of each training sample m and the corresponding prediction result Y m calculate the loss value Loss t of W t , and then use Loss t to update the weight ω t to obtain the improved YOLOv11 network model W after this iteration t ; (3d) Determine whether t = T holds. If so, obtain the trained network model W * , otherwise, set t = t + 1, and W t = W, and execute step (3b).

6. The method according to claim 4, characterized in that, The reversible feature extraction column sub-network described in step (3b) performs deep feature extraction on each training sample, and the implementation steps are as follows: S depthwise separable convolution modules in each feature extraction branch perform feature extraction of different scales on each training sample; S cross-stage residual convolution modules perform feature enhancement on the extracted multi-scale feature maps; the cross-branch reversible interaction module performs pooling on the feature-enhanced feature maps to obtain S different-scale feature maps output by the k-th branch.

7. The method according to claim 4, wherein The feature fusion sub-network described in step (3b) performs multi-scale feature fusion on the extracted deep feature information after weighting according to the importance of different regions, and the implementation steps are as follows: The depthwise separable convolution module in the separable kernel attention module extracts region importance feature information from the S different-scale feature maps output by the K-th branch; The standard convolution module performs a weighting operation on the feature maps extracting region importance feature information according to the extracted region importance feature information to obtain S weighted different-scale feature maps; The standard convolution module and the bilinear interpolation module in the feature pyramid module perform downsampling and upsampling on the S weighted different-scale feature maps respectively; the feature concatenation module concatenates the downsampling and upsampling results to obtain a multi-scale fusion feature map.

8. The method according to claim 4, wherein The detection sub-network described in step (3b) performs prediction after performing occlusion relationship feature enhancement on the fused multi-scale fusion feature map, and the implementation steps are as follows: The depthwise separable convolution module in the occlusion-aware attention module extracts occlusion relationship features from the multi-scale fusion feature map, and the occlusion enhancement module enhances the multi-scale fusion feature map according to the extracted occlusion relationship features to obtain an occlusion-aware enhanced multi-scale fusion feature map; The multi-scale classification branch decoupling head in the detection head module predicts the category information and confidence information of the multi-scale fusion feature map with enhanced occlusion perception, while the multi-scale regression branch decoupling head predicts the position information of the multi-scale fusion feature map with enhanced occlusion perception, obtaining the prediction result Y including the category, confidence, and position information of weeds in each training sample m 。 9. The method according to claim 4, characterized in that W as described in step (3c) t loss value Loss t The calculation formula is as follows: Loss t = λ loc Loss loc t + λ cls Loss cls t Among them, Loss loc t and Loss cls t represent the localization loss and the classification loss respectively. λ loc and λ cls represent the weight coefficients of Loss loc t and Loss cls t respectively; GIoU(·) represents the Generalized Intersection over Union loss.

10. The method according to claim 4, characterized in that, The update of the weight ω described in step (3c) t is carried out, and the update formula is as follows: Among them, η represents the gradient descent parameter, and ω t+1 represents the updated result of ω t , and represents the partial derivative operation.

Citation Information

Patent Citations

  • Field weed detection method based on YOLOv8 optimization

    CN118314438A