Winter wheat weed segmentation method based on unmanned aerial vehicle RGB data
Patent Information
- Application Number
- CN202610713956.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-05-22
- Publication Date
- 2026-09-29
AI Technical Summary
[0006]为了解决现有冬小麦出苗期杂草分割方法存在的分割精度低、小尺度杂草漏检率高、特征利用不充分等问题,本发明提供了一种基于无人机RGB数据的冬小麦杂草分割方法
[0026](1)本发明提出的SSMR-Net杂草分割模型,通过在编码器-解码器架构中集成残差模块、ASPP模块和scSE机制的跳跃连接,有效提升了模型对复杂场景的适应能力,在冬小麦出苗期杂草与小麦颜色相近、形态相似、生长重叠的环境中实现了高精度像素级语义分割;
Smart Images

Figure CN122841985A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of agricultural informatization and computer vision technology, and in particular to a method for segmenting winter wheat weeds based on RGB data from unmanned aerial vehicles (UAVs). Background Technology
[0002] Weeds are a major factor affecting wheat yield and quality, especially during the winter wheat seedling stage when competition between weeds and wheat is most intense. Timely and effective weed management is crucial for ensuring healthy wheat growth and high and stable yields. Currently, chemical herbicides remain the primary means of weed management in the field, but inappropriate application not only wastes resources but also causes negative impacts such as soil pollution and ecological damage. Precision herbicide application technology has gained widespread attention due to its cost-effectiveness and environmental advantages. The success of precision herbicide application depends on the accurate identification and location of weeds; therefore, it is necessary to develop efficient and accurate weed identification strategies to support sustainable agricultural production.
[0003] In recent years, the rapid development of artificial intelligence and computer vision technologies has driven the transformation of traditional agriculture towards precision and intelligent agriculture. In the field of weed identification, semantic segmentation methods based on deep learning have shown great potential. Remote sensing monitoring systems based on drone platforms can quickly acquire high-resolution farmland images. Compared to aircraft or satellite remote sensing, drones offer significant advantages such as lower cost, greater operational flexibility, and higher image resolution. Although multispectral cameras have unique advantages in vegetation analysis, RGB cameras are more practical due to their low cost, ease of data processing, and high image quality, making them particularly suitable for large-scale application in agricultural extension.
[0004] While some existing methods for crop and weed segmentation based on UAV RGB imagery exist, they still suffer from the following technical limitations: First, during the wheat seedling stage, weeds and wheat are small, similar in shape, and have complex backgrounds, making it difficult for existing semantic segmentation models to effectively distinguish between them, especially in areas where weeds and wheat overlap, where segmentation accuracy drops significantly. Second, weeds in wheat fields exhibit uneven distribution patterns with significant differences in size scale. Traditional encoder-decoder architectures such as U-Net have limited capabilities in multi-scale feature extraction, resulting in poor recognition of small-sized weeds and a tendency to miss detections. Third, in existing methods, the shallow spatial information of the encoder and the deep semantic information of the decoder are directly fused through simple skip connections, leading to a semantic level mismatch and low feature utilization, which limits further improvement in segmentation accuracy.
[0005] Therefore, there is an urgent need for a weed segmentation method that can handle the complex scenarios during the winter wheat seedling stage, take into account the identification of weeds at different scales, and has high segmentation accuracy. Summary of the Invention
[0006] To address the problems of low segmentation accuracy, high miss rate of small-scale weeds, and insufficient utilization of features in existing winter wheat weed segmentation methods during the seedling stage, this invention provides a winter wheat weed segmentation method based on UAV RGB data.
[0007] To achieve the above objectives, the present invention is implemented according to the following technical solution:
[0008] A method for segmenting winter wheat weeds based on UAV RGB data includes the following steps:
[0009] S1. Collect RGB image data of farmland during the winter wheat seedling stage using a drone platform equipped with an RGB lens, and perform data preprocessing and annotation to construct a training dataset;
[0010] S2. Construct a weed segmentation model SSMR-Net (Single-Stage Multi-Scale Residual Network). The weed segmentation model is based on the encoder-decoder architecture U-Net. In this encoder-decoder architecture, a residual module is introduced to extract deep features, a hollow spatial pyramid pooling module is introduced to capture multi-scale contextual information, a fusion upsampling module is introduced to achieve semantic alignment between shallow and deep features, and the scSE mechanism is used to recalibrate the skip connections to generate multi-scale spatial feature representations.
[0011] S3. The weed segmentation model SSMR-Net constructed in step S2 is combined with the AcrossFeature Mapping Attention (AFMA) module. The AFMA module segments the original image and feature map into image blocks of equal size, calculates the correlation between feature blocks of different feature layers, captures the spatial dependency between weed targets at different scales, and generates attention maps to adjust the model output.
[0012] S4. The weed segmentation model is trained using a joint loss function consisting of a segmentation loss function and a cross-feature mapping attention loss function, wherein the segmentation loss function assigns different median frequency weights to different categories to alleviate the class imbalance problem.
[0013] S5. Input the UAV RGB image of the winter wheat field to be segmented into the trained weed segmentation model, and output the pixel-level semantic segmentation results of winter wheat and weeds.
[0014] Furthermore, the data preprocessing in step S1 includes image size normalization adjustment, the data annotation is manually performed using polygon semantic tagging, and the dataset construction also employs data augmentation techniques including horizontal flipping, random scaling and cropping, Gaussian noise addition, and perspective transformation to expand the training dataset.
[0015] Furthermore, in step S2, the residual module adopts an enhanced residual block structure and enhances the feature representation through convolution operations; the dilated spatial pyramid pooling module adopts 3×3 dilated convolutions with dilation rates of 1, 3, and 5, and combines them with depthwise separable convolutions to reduce the number of parameters; the fusion upsampling module introduces a spatial and channel squeezing excitation mechanism to adaptively fuse the encoder's feature map with the decoder's upsampled feature map.
[0016] Furthermore, the segmentation loss function adopts the median frequency-balanced weighted sigmoid cross-entropy loss, as shown in formula (1):
[0017] (1)
[0018] in, This represents the true value (0 or 1) of category k at pixel (h,w). This represents the model's final predicted probability for category k. This represents the median frequency weight assigned to category k. By assigning median frequency weights, the model gives greater attention to the weed category, effectively mitigating the class imbalance problem between wheat and weeds in the segmentation task.
[0019] For the AFMA module, the cross-feature map attention loss function adopts the mean squared error loss, as shown in Equation (2):
[0020] (2)
[0021] in, and These are the predicted AFMA and the optimal AFMA, respectively.
[0022] The overall training loss consists of the segmentation loss and the AFMA loss, as shown in formula (3):
[0023] (3).
[0024] Furthermore, after outputting the segmentation results in step S5, the method also includes generating a visual map of weed distribution based on the segmentation results to support precise pesticide application.
[0025] Compared with the prior art, the present invention has the following beneficial effects:
[0026] (1) The SSMR-Net weed segmentation model proposed in this invention effectively improves the model’s adaptability to complex scenarios by integrating residual modules, ASPP modules and scSE mechanism skip connections in the encoder-decoder architecture. It achieves high-precision pixel-level semantic segmentation in an environment where weeds and wheat are similar in color, shape and growth during the winter wheat seedling stage.
[0027] (2) This invention organically combines SSMR-Net with the Cross Feature Map Attention Module (AFMA), which significantly improves the model's ability to segment small-sized weeds by capturing the spatial dependencies and feature block correlations between weed targets of different scales, and effectively reduces the weed false detection rate.
[0028] (3) The joint loss function used in this invention alleviates the class imbalance problem by allocating median frequency weights, and guides the model to learn the optimal feature association by using AFMA loss, which takes into account both segmentation accuracy and feature representation quality, and improves the overall segmentation performance of the model.
[0029] (4) This invention achieves weed segmentation based on low-cost RGB drone images, avoiding the high cost of specialized equipment such as multispectral imaging. It has the outstanding advantages of simple operation, easy data processing, and easy promotion and application, and can provide reliable technical support for precision pesticide application. Attached Figure Description
[0030] Figure 1 This is an example diagram of a drone data acquisition device and data in an embodiment of the present invention.
[0031] Figure 2 This is an example diagram of data annotation in an embodiment of the present invention.
[0032] Figure 3 This is an example diagram of a dataset challenge in an embodiment of the present invention.
[0033] Figure 4 This is a structural diagram of the weed segmentation model SSMR-Net in an embodiment of the present invention.
[0034] Figure 5 The following is a structural diagram of the fusion upsampling and connection module in an embodiment of the present invention: (a) is the fusion upsampling module; (b) is the connection module.
[0035] Figure 6 Here is a structural diagram of the cross-feature map attention module (AFMA) in an embodiment of the present invention: (A) is the generation part of cross-feature map attention; (B) is the adjustment of the decoder output using the obtained attention map; (C) is the generation of the optimal AFMA.
[0036] Figure 7These are example diagrams showing weed segmentation results under different environments in embodiments of the present invention. Detailed Implementation
[0037] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to embodiments. The specific embodiments described herein are for illustrative purposes only and are not intended to limit the invention.
[0038] This embodiment exemplifies a method for weed segmentation of winter wheat based on UAV RGB data. The study area is located at the Smart Agriculture Demonstration Base of Jiaozuo Academy of Agricultural and Forestry Sciences in Henan Province. The winter wheat was in the seedling stage after sowing, and the data was collected in December 2023.
[0039] 1. Study Area and Data Collection
[0040] The study area for this embodiment is located at the Smart Agriculture Demonstration Base of the Jiaozuo Academy of Agricultural and Forestry Sciences in Henan Province, where winter wheat was in the seedling stage after sowing. Data collection was conducted in December 2023.
[0041] like Figure 1 As shown, image acquisition was conducted using a DJI Phantom 4 drone platform equipped with a high-resolution RGB lens. The image resolution is 5472×3648 pixels. The drone flew between 9:00 AM and 12:00 PM under clear and partly cloudy weather conditions, with a flight altitude set at 5 meters to ensure clear field of view and sufficient ground spatial resolution. During the acquisition process, areas with different wheat distribution densities and different lighting conditions were selected for image acquisition to ensure data diversity and representativeness.
[0042] 2. Dataset Construction and Augmentation
[0043] During the winter wheat emergence period, a total of 453 original images were collected from different areas within the experimental plot. Considering the limitations of computing resources, the original image size was adjusted to 640×480 pixels. To balance the speed and accuracy of model training, 2830 valid images were obtained through image cropping and random sampling, and were divided into a training set (2264 images), a validation set (283 images), and a test set (283 images) at a ratio of 80%, 10%, and 10%, respectively.
[0044] like Figure 2 As shown, before training the weed segmentation model, the images were manually semantically labeled using Labelme software, with categories labeled using polygons. Subsequently, the labeled information was stored in JSON format and converted to PascalVOC format to generate binary mask images of wheat and weeds.
[0045] To enhance the robustness and generalization ability of the model, this invention employs various data augmentation techniques, including horizontal flipping, random scaling and cropping, Gaussian noise addition, and perspective transformation. Gaussian noise is used to simulate image quality variations by randomly adding Gaussian-distributed pixel noise to the image, enhancing the model's adaptability to different imaging qualities. Perspective transformation projects a planar image onto a specified plane using a projection matrix, performing a random four-point perspective transformation to simulate image effects under different shooting angles, significantly increasing the spatial diversity of the data.
[0046] 3. Model Structure
[0047] like Figure 4 As shown, the weed segmentation model SSMR-Net (Single-Stage Multi-Scale Residual Network) proposed in this invention is built based on the U-Net encoder-decoder architecture. This network consists of an improved encoder, decoder, ASPP module, and cross-feature map attention module (AFMA).
[0048] The encoder section is an improvement upon the residual units in SuperU-Net, replacing standard convolutional modules with enhanced residual blocks. These enhanced residual blocks, through a combination of residual connections and convolutional operations, effectively address the degradation problem in deep networks, ensuring stable feature transmission in deep networks and reducing information loss during forward propagation. The decoder section includes a fusion upsampling module and a residual module. The fusion upsampling module employs a spatial and channel squeeze excitation (scSE) mechanism to aggregate the divergent feature maps extracted by the encoder into similar features, facilitating the establishment of effective skip connections with the encoder.
[0049] like Figure 4 As shown, the ASPP module bridges the encoder and decoder, employing dilated convolutions with dilation rates of 1, 3, and 5 and a kernel size of 3×3, and introducing depthwise separable convolutions to reduce the number of parameters. This design allows the model to associate information from a broad field of view in the previous layer, effectively preventing the loss of small target features during information transmission. The skip connections introduce the scSE mechanism to recalibrate the shallow features of the encoder and the deep features of the decoder, reducing the semantic gap between feature layers of the same depth and improving the efficiency of direct skip connections.
[0050] like Figure 5As shown, the structure of the fusion upsampling and connection modules is as follows: the recalibrated feature map of the encoder is passed to the decoder through long skip connections. The feature maps of the encoder and decoder are first input into the scSE module. The encoded feature map is processed by the channel-dimensional cSE attention mechanism and the spatial-dimensional sSE attention mechanism. In the channel attention branch, the feature map dimension is compressed from [C, H, W] to [C, 1, 1] through global average pooling. Then, two 1×1 convolutions are used for information processing (dimensionality reduction and dimensionality increase). After normalization by the sigmoid function, it is multiplied with the original feature map channel by channel to obtain the channel-calibrated feature map. In the spatial attention branch, a weight matrix of [1, H, W] is obtained through a convolutional layer with 1 output channel and a kernel size of 1×1. After sigmoid normalization, it is multiplied with the original feature map in the spatial dimension to obtain the spatially enhanced feature map. Finally, the features of the channel branch and the spatial branch are added along the channel dimension to adjust the network features, emphasizing important features and suppressing the influence of irrelevant features, ultimately improving the semantic segmentation effect.
[0051] like Figure 6 As shown, the Cross Feature Map Attention Module (AFMA) consists of three components: Cross Feature Map Attention Generation, Output Adjustment, and Optimal AFMA Generation.
[0052] Figure 6 Part (A) illustrates the generation process of cross-feature map attention. AFMA consists of two branches: the first branch performs convolution on the original feature map with 64 7×7 kernels, followed by a 1×1 kernel, and finally divides it into a series of flat image patches of fixed size; the second branch performs similarity operations on the feature map, first converting it into a feature representation for each class through a 1×1 convolution of the number of output classes, and then dividing it into a series of flat image patches of fixed size. Subsequently, the association between each image patch of the original image and the feature map associated with the k-th class is calculated through dot product, ultimately generating AFMA.
[0053] Figure 6 Part (B) illustrates the process of adjusting the decoder output using the obtained attention map. First, the decoder output is compressed to a specified size using fixed-size pooling. Then, each channel is divided into a series of flat image blocks, and the AFMA is multiplied by the result. Finally, the product is folded back to the original output size, thereby adjusting the decoder's final output.
[0054] Figure 6 Part (C) shows the optimal AFMA generation method: the image is split into blocks from the real label image and then split into blocks after pooling. The results of the processing are multiplied by a dot product to obtain the optimal AFMA.
[0055] 4. Loss Function
[0056] The overall loss function consists of the standard loss function and the AFMA loss function. The standard loss aims to minimize the difference between the ground truth segmentation mask and the prediction, while the AFMA loss minimizes the difference between the learned AFMA and the optimal AFMA. For the segmentation loss, we use the median frequency-balanced weighted sigmoid cross-entropy loss for training, as shown in Equation (1):
[0057] (1)
[0058] in, This represents the true value (0 or 1) of category k at pixel (h,w). This represents the model's final predicted probability for category k. This represents the median frequency weight assigned to category k. By assigning median frequency weights, the model gives greater attention to the weed category, effectively mitigating the class imbalance problem between wheat and weeds in the segmentation task.
[0059] For the AFMA module, the cross-feature map attention loss function is trained using mean squared error (MSE) loss, as shown in Equation (2):
[0060] (2)
[0061] in, and These are the predicted AFMA and the optimal AFMA, respectively.
[0062] The overall training loss consists of the segmentation loss and the AFMA loss, as shown in formula (3):
[0063] (3)
[0064] This joint loss function balances segmentation accuracy and feature mapping quality during the optimization process, guiding the model to learn accurate pixel classification and optimal inter-feature block relationships simultaneously, thereby significantly improving the final semantic segmentation performance.
[0065] 5. Model Training and Evaluation
[0066] This embodiment uses a combination of metrics to evaluate model performance: Intersection over Union (IOU), Precision, Mean Intersection over Union (MIOU), and Average Frames Per Second (aFPS). IOU measures the overlap between the predicted mask and the ground truth segmentation mask. A higher IOU score indicates a greater similarity between the predicted segmentation region and the ground truth label. Precision refers to the proportion of pixels correctly predicted by the model for a specific class, i.e., the ratio of the number of true pixels of that class to the total number of pixels predicted as belonging to that class. MIOU provides an overall view of the model's performance across the entire dataset, offering a more comprehensive performance evaluation, especially when dealing with imbalanced classes, and better reflecting the model's true performance. Furthermore, MIOU is measured by averaging the MIOU for each class, providing a weighted average of the MIOU for each class. aFPS is an important metric for evaluating the performance of semantic segmentation models in practical applications. Model evaluation needs to comprehensively consider the balance between accuracy and inference speed.
[0067] (4)
[0068] (5)
[0069] (6)
[0070] Model training was performed using the PyTorch deep learning framework, employing the Adam optimizer for parameter optimization. The initial learning rate was set to 0.001, and a cosine annealing learning rate decay strategy was used for a total of 200 epochs. During training, model parameters were saved, and the best-performing model parameters on the validation set were selected for evaluation on the test set.
[0071] To verify the effectiveness of the AFMA module, a comparison was made with the eight different models mentioned above. Table 1 and Figure 7 The results and graphs show the evaluation results of different models after using AFMA.
[0072] Table 1. Comparison of segmentation results using different methods combined with AFMA
[0073] SSMR-Net 0.685 0.741 0.944 0.980 0.870 2.24 <![CDATA[SSMR-Net AFMA ]]> 0.696(1.6%↑) 0.774 0.923 0.978 0.865 1.92 DeepLabv3 0.653 0.685 0.771 0.942 0.789 1.18 <![CDATA[DeepLabv3 AFMA ]]> 0.669(2.5%↑) 0.719 0.790 0.947 0.802 1.15 PSPNet 0.673 0.716 0.779 0.946 0.799 1.63 <![CDATA[PSPNet AFMA ]]> 0.681(1.2%↑) 0.734 0.785 0.947 0.804 1.58 LinkNet 0.675 0.722 0.884 0.967 0.842 3.67 <![CDATA[LinkNet AFMA ]]> 0.693(2.7%↑) 0.745 0.891 0.970 0.851 3.43 MANet 0.671 0.712 0.890 0.969 0.843 2.22 <![CDATA[MANet AFMA ]]> 0.677(0.9%↑) 0.727 0.901 0.971 0.849 2.17 Super UNet 0.675 0.740 0.898 0.971 0.848 6.54 <![CDATA[Super UNet AFMA ]]> 0.684(1.3%↑) 0.756 0.897 0.971 0.851 4.54 UNet++ 0.675 0.727 0.893 0.971 0.846 0.94 <![CDATA[UNet++ AFMA ]]> 0.693(2.7%↑) 0.742 0.920 0.975 0.863 0.93 ResUNet. 0.678 0.717 0.895 0.970 0.848 3.32 <![CDATA[ResUNet. AFMA ]]> 0.692(2.1%↑) 0.741 0.907 0.973 0.857 3.10
[0074] From Table 1 and Figure 7It can be seen that, compared with the base model without AFMA, applying AFMA to the baseline network significantly improves the IOU score for weeds. Specifically, when AFMA is combined with the proposed SSMR-Net, DeepLabv3, PSPNet, LinkNet, MANet, Super UNet, UNet++, and ResUNet base segmentation networks, the IOU for weeds is improved by 1.6%, 2.5%, 1.2%, 2.7%, 0.9%, 1.3%, 2.7%, and 2.1%, respectively. Improvements are also seen in wheat segmentation, but for soil segmentation, AFMA actually reduces accuracy, possibly due to the presence of dissimilar components in the soil, such as abundant fallen leaves, gravel, and different soil types. These factors lead to significant feature differences between different parts of the soil, weakening the correlation between features and thus affecting the model's segmentation performance. Furthermore, the impact of AFMA on the model's average frame rate (aFPS) is negligible, indicating limited performance overhead and maintaining overall efficiency.
[0075] The winter wheat weed segmentation method based on UAV RGB data proposed in this embodiment effectively solves the technical challenges of weed segmentation during the winter wheat seedling stage, such as morphological similarity, scale diversity, and complex background, by integrating an enhanced residual module, an ASPP module, a jump connection mechanism based on scSE, and an AFMA module into the U-Net framework. Experimental results show that this method outperforms existing mainstream models in key indicators such as weed identification accuracy (0.774) and intersection-over-union ratio (0.696), verifying the effectiveness and advancement of the technical solution of this invention. This method can provide a high-precision weed distribution map for precision pesticide application systems and has significant practical application value.
[0076] The technical solutions of the present invention are not limited to the specific embodiments described above. Any technical modifications made in accordance with the technical solutions of the present invention fall within the protection scope of the present invention.
Claims
1. A method for segmenting winter wheat weeds based on UAV RGB data, characterized in that, Includes the following steps: S1. Collect RGB image data of farmland during the winter wheat seedling stage using a drone platform equipped with an RGB lens, and perform data preprocessing and annotation to construct a training dataset; S2. Construct a weed segmentation model. The weed segmentation model is based on an encoder-decoder architecture. In this encoder-decoder architecture, a residual module is introduced to extract deep features, a hollow spatial pyramid pooling module is introduced to capture multi-scale contextual information, a fusion upsampling module is introduced to achieve semantic alignment between shallow and deep features, and the scSE mechanism is used to recalibrate the skip connections to generate multi-scale spatial feature representations. S3. The weed segmentation model constructed in step S2 is combined with the cross-feature map attention module. The cross-feature map attention module divides the original image and feature map into image blocks of equal size, calculates the feature block correlation between different feature layers, captures the spatial dependency between weed targets at different scales, and generates an attention map to adjust the model output. S4. The weed segmentation model is trained using a joint loss function consisting of a segmentation loss function and a cross-feature mapping attention loss function, wherein the segmentation loss function assigns different median frequency weights to different categories to alleviate the class imbalance problem. S5. Input the UAV RGB image of the winter wheat field to be segmented into the trained weed segmentation model, and output the pixel-level semantic segmentation results of winter wheat and weeds.
2. The method for segmenting winter wheat weeds based on UAV RGB data according to claim 1, characterized in that, The data preprocessing in step S1 includes image size normalization adjustment, the data annotation is manually performed using polygon semantic tagging, and the dataset construction also uses data augmentation techniques including horizontal flipping, random scaling and cropping, Gaussian noise addition, and perspective transformation to expand the training dataset.
3. The method for segmenting winter wheat weeds based on UAV RGB data according to claim 1, characterized in that, In step S2, the residual module adopts an enhanced residual block structure, which enhances the feature representation through convolution operations. The hollow spatial pyramid pooling module uses 3×3 hollow convolutions with dilation rates of 1, 3, and 5, and combines them with depthwise separable convolutions to reduce the number of parameters; the fusion upsampling module introduces a spatial and channel squeezing excitation mechanism to adaptively fuse the encoder's feature map with the decoder's upsampled feature map.
4. The method for segmenting winter wheat weeds based on UAV RGB data according to claim 1, characterized in that, The segmentation loss function adopts the median frequency-balanced weighted sigmoid cross-entropy loss, as shown in formula (1): (1) in, This represents the true value (0 or 1) of category k at pixel (h,w). This represents the model's final predicted probability for category k. This represents the median frequency weight assigned to category k. By assigning median frequency weights, the model gives greater attention to the weed category, effectively mitigating the class imbalance problem between wheat and weeds in the segmentation task. For the AFMA module, the cross-feature map attention loss function adopts the mean squared error loss, as shown in Equation (2): (2) in, and These are the predicted AFMA and the optimal AFMA, respectively. The overall training loss consists of the segmentation loss and the AFMA loss, as shown in formula (3): (3)。 5. The method for segmenting winter wheat weeds based on UAV RGB data according to claim 1, characterized in that, After outputting the segmentation results in step S5, the method also includes generating a visual map of weed distribution based on the segmentation results to support precise pesticide application.