Weed density estimation model and estimation method based on deep learning model
Through the CombinedUNet model and CBAM attention mechanism, combined with global and local feature extraction, the problem of low detection accuracy of existing systems in low density and small-scale areas is solved, and high-precision weed density estimation and grid management are achieved.
Patent Information
- Application Number
- CN202510295056.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-11
AI Technical Summary
Existing weed monitoring systems based on drone images have low detection accuracy in low-density or small-scale areas, and lack grid density estimation and dynamic spray control for different areas of farmland.
The CombinedUNet model architecture is adopted, combining global branches and local branches, and weed density maps are output through feature extraction, fusion and decoding, and features are enhanced by CBAM attention mechanism, and model training is carried out in combination with weakly supervised learning.
In a complex background, weed detection accuracy and robustness are improved, and missed detection is reduced, and it is suitable for farmland under different soil colors and light intensity, reducing labeling costs, and achieving high-precision grid-based weed density estimation.
Smart Images

Figure CN120298923A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the fields of computer vision, remote sensing technology, and agricultural science and technology, and particularly relates to a weed density estimation model and an estimation method based on a deep learning model. Background Art
[0002] The rise of unmanned aerial vehicle (UAV) image acquisition technology has provided new possibilities for weed monitoring. UAVs can quickly cover large areas of farmland and collect high-resolution images, providing data support for real-time monitoring of weed distribution. Combining deep learning technology, automatic identification and density estimation of weeds can be achieved, providing a basis for the management and control of weeds in precision agriculture. However, existing systems mainly focus on image classification or simple object detection. These systems perform well in low-density weed areas, but in areas with dense weeds or small scales, their detection accuracy drops significantly. At the same time, they mainly focus on the overall analysis of images and lack grid density estimation and dynamic spraying control for different regions of farmland.
[0003] Most UAV image-based weed monitoring systems do not consider the impact of image resolution on recognition accuracy. In practical applications, the flight altitude and device resolution of UAVs are limited, resulting in insufficient image clarity, especially the features of small weeds are not obvious, which affects the detection effect of deep learning models. Therefore, a detection model with good effect is needed to achieve grid weed density estimation. Summary of the Invention
[0004] Technical Solution: To solve the above technical problems, the present invention provides a weed density estimation model based on a deep learning model. The technical solution is as follows:
[0005] The model is a CombinedUNet model architecture, including an input processing module, a feature extraction module for the global branch and the local branch, a feature fusion module, a decoding module, and an output density map module;
[0006] The model inputs the global image and the local image through the input processing module, then extracts features from the input images under the feature extraction modules of the global branch and the local branch, then performs splicing and feature fusion through the feature fusion module, and the decoding module performs step-by-step decoding, and finally outputs a density map.
[0007] As an improvement, the input sizes of the global image and the local image are both 3×H×W, where the global image is denoted as X global ∈R 3×H×W , and the local image is denoted as X local ∈R 3×H×W , H is the height, and W is the width.
[0008] As an improvement, the global branch is used to extract low-resolution global semantic features, denoted as: b global, [e global [1],e global [2],e global [3],e global [4],] = UNetGlobalBranch(X global );
[0009] The local branch is used to extract high-resolution local detail features, denoted as:
[0010] b local, [e local [1],e local [2],e local [3],e local [4],] = UNetLocalBranch(X local ).
[0011] As an improvement, the specific method for the feature fusion module to perform feature fusion is: (1) Feature alignment: Use bilinear interpolation to adjust the global branch feature size to be the same as the output size of the local branch: b global →F.interpolate(b global , size = b local.shape [2:]); (2) Feature concatenation: Concatenate the local feature and the global feature, denoted as b combined = Concat(b local , b global_resized , dim = 1), and the number of channels after concatenation is 2048; (3) Use the channel attention mechanism and spatial attention mechanism of CBAM for enhancement and output the fused feature; (4) Use 1×1 convolution for dimensionality reduction processing and output the feature size.
[0012] As an improvement, the channel attention mechanism is a processing method for generating channel weights through global average pooling and max pooling, and the formula is: F CA = F·σ(MLP(GAP(F))), where MLP: fully connected layer, GAP(F): global average pooling;
[0013] The spatial attention mechanism is to generate spatial weights through average pooling and max pooling combined with convolution, and the formula is: F SA = F CA ·σ(Conv2D(Concat(Mean(F CA ), Max(F CA )))));
[0014] The feature b after the enhanced feature is output enhanced, denoted as b enhanced = CBAM(b combined ).
[0015] As an improvement, the specific steps of the channel attention mechanism are as follows
[0016] (1) Set the goal and generate channel weights according to the assigned weights of each channel
[0017] (2) Input the feature, where the feature is X ∈ R C×H×W , C: number of channels; H, W: height and width of the feature map
[0018] (3) Perform global pooling operation and output two channel description vectors, the vectors are output z avg , z max ∈ R C ;
[0019] (4) Share a fully connected network, input z avg and z max into a shared two-layer fully connected network respectively to obtain two weight results
[0020] (5) Weightedly sum the two weight results to get the final channel weight, s c = s avg + s max ; Then multiply each channel of the input feature XX by the corresponding weight to generate an enhanced feature map X CA = X ⊙ s c , where ⊙ is element-wise multiplication of channels
[0021] As an improvement, the specific steps of the spatial attention mechanism are as follows
[0022] (1) Set the goal, highlight the weed information in specific regions of the image in the spatial dimension, and generate a spatial weight formula
[0023] (2) Input the feature, where the feature is X CA ∈ R C×H×W , C: number of channels; H, W: height and width of the feature map
[0024] (3) Perform pooling processing in the spatial dimension and output two spatial feature vectors, the formula is X avg , X max ∈ R 1×H×W ;
[0025] (4) Concatenation and convolution processing, concatenate the two spatial feature maps in the spatial dimension to form a 2-channel feature map: X S = Concat(X avg, X max ), X S ∈R 2×H×W ;
[0026] (5) Input a convolutional layer with a convolution kernel of k×k to generate a spatial weight map W S ∈R 1×H×W , and then apply the spatial weight to the input feature map, X SA = X CA ⊙W S , where ⊙ is element-wise multiplication of channels.
[0027] Meanwhile, the present invention also provides a method for estimating weed density based on a deep learning model. The specific steps of the method are as follows:
[0028] (1) Input image patches
[0029] Collect and obtain image patches in image format, and each image patch corresponds to the GPS positioning information of the target;
[0030] (2) Image preprocessing
[0031] Perform normalization processing on the image patches in (1), scale the pixel values from [0, 255] to [0, 1] to unify the input data distribution; then perform tensor conversion, and convert the image into tensor format using a deep learning framework; then perform batch loading to organize multiple image patches into a batch;
[0032] (3) Build a model
[0033] Build the above-mentioned weed density estimation model based on any deep learning model;
[0034] (4) Learning and training
[0035] Prepare the annotation for the image data, including annotated data and unannotated data; then, use a part of the annotated data for preliminary training, compare the density map predicted by the model with the true total number of weeds, optimize the model through weak supervision loss, then iteratively generate pseudo-labels, and perform joint training in combination with the original annotated data; finally, adjust the pixel-level accuracy and consistency of the model by weighted summation of two losses, namely the density map loss function and the total number of weeds loss function;
[0036] (5) Model output
[0037] Output the weed density map, where the output pixel value size indicates the relative weed density; then calculate the total number of weeds in the grid area by accumulating all the pixel values of the cumulative weed density map.
[0038] As an improvement, when estimating the total number of weeds through the density map in (5), all the pixel values in the density map are accumulated, where D i is the value of the i-th pixel in the density map, and N is the total number of pixels in the density map.
[0039] Beneficial effects: The model and estimation method proposed in the present invention combine a global branch and a local branch, and propose a multi-scale feature extraction method, enabling the model to adapt to sparse regions and accurately detect weeds in high-density regions. Further, the introduction of the CBAM attention mechanism helps the model focus on weed features under uneven illumination and complex backgrounds, improving the robustness and accuracy of the model. Through the pseudo-label generation technology, unlabeled data is fully utilized, reducing the dependence on manual precise labeling.
[0040] The introduction of weak supervision learning enables the system to achieve high-precision weed density estimation with extremely low labeling costs. It can be used in farmlands with complex terrains (such as different soil colors and light intensities), effectively distinguish the boundaries between weeds and crops, reduce false detections and missed detections, and is also applicable to scenarios where it is difficult to label weeds in farmlands, especially in regions with complex terrains or few experts involved in labeling. Description of the Drawings
[0041] Figure 1 It is a schematic diagram of the CombinedUNet model architecture of the present invention.
[0042] Figure 2 It is a schematic diagram of the output density map in Embodiment 1 of the present invention. Detailed Embodiments
[0043] The technical solutions in the embodiments of the present invention will be clearly and completely described below, so that those skilled in the art can better understand the advantages and features of the present invention, and thus more clearly define the protection scope of the present invention. The embodiments described in the present invention are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0044] The CombinedUNet model architecture of the present invention includes a global branch, a local branch, and adopts the CBAM attention mechanism; feature fusion is performed on the features of the global branch and the local branch in the decoding stage, and skip connections are used to transfer high-resolution features in the encoder to the decoder; then upsampling is gradually performed to restore to the original resolution, that is, each decoder layer receives the skip connection features from the encoder and fuses them with the features of the current decoding layer.
[0045] In the present invention, the global branch is used to extract the global context features of a low-resolution image, including 4 layers of convolutional pooling and 1 bottleneck layer structure, with the input end being an image of 3×H×W (RGB 3 channels); the structure of the global branch includes an Encoder part, a Bottleneck part, and an output end; the Encoder part of the encoder includes 4 layers of encoders. After each layer extracts features through a convolutional block, a max pooling operation is performed to reduce the resolution. Each convolutional block includes 2 layers of convolution, batch normalization, and ReLU activation; the feature map of each layer is gradually reduced in resolution through the max pooling layer. Preferably, the size of the feature map is gradually reduced (pooling factor is 2), and the number of feature channels gradually increases: 3→64→128→256→512.
[0046] The input channel number of the Bottleneck part is 512, and the output channel number is 1024, which is used to extract the final global semantic features. The final bottleneck layer feature of the output end is b, with the size of Skip connections are made through the intermediate features [e1, e2, e3, e4].
[0047] In the present invention, the local branch is used to retain the high-resolution detail information of the image, paying special attention to local features, with the input end being an image of 3×H×W (RGB 3 channels). The structure of the local branch includes an Encoder part, a Bottleneck part, and an output end; the Encoder part of the local branch is similar to that of the global branch, but the local branch deliberately reduces the pooling operation to retain more spatial resolution information.
[0048] Furthermore, the decoder module of this model gradually restores the resolution by gradually upsampling and combining the intermediate features of the global and local branches, and outputs a density estimation map. The structure of this module includes 4 layers of decoding modules, and the input of each layer is the feature of the current decoding layer, the intermediate features of the local branch with skip connections, and the global branch.
[0049] The formula of the decoder module is:
[0050] ConvBlock(Concat(UpSample(d i+1 ),e local [i],e global [i]))
[0051] where UpSample: upsampling (bilinear interpolation, scale factor is 2);
[0052] e local [i],e global [i]: the intermediate features corresponding to the local branch and the global branch at the (i)-th layer, and i is an integer such as 1, 2, 3, 4, etc.
[0053] The number of decoded feature channels decreases layer by layer: d i = 1024 + 1024 → 512 → 256 → 128 → 64; The last output layer is a 1*1 convolution for generating a density estimation map, and the output feature size is 1×H×W.
[0054] In the present invention, the network process of the entire model is specifically as follows:
[0055] (1) Input processing
[0056] Two inputs: a global image and a local image, both with a size of 3×H×W. Expressed by the formula: global image X global ∈R 3×H×W , local image X local ∈R 3×H×W ;
[0057] (2) Global and local branch feature extraction:
[0058] The global branch extracts low-resolution global semantic features: b global ,[e global [1],e global [2],e global [3],e global [4],] = UNetGlobalBranch(X global );
[0059] The local branch extracts high-resolution local detail features: b local ,[e local [1],e local [2],e local [3],e local [4],] = UNetLocalBranch(X local );
[0060] e local [i],e global [i]: The intermediate feature corresponding to the (i)-th layer of the local branch and the global branch, where i is an integer from 1 to 4;.
[0061] (3) Feature fusion: Concatenate the global and local features and use CBAM to enhance the fused features:
[0062] b final = Conv1×1(CBAM(Concat(b local ,b global_aligned )));
[0063] b local : The final output feature of the local branch, high-resolution local detail feature; bglobal_aligned : The feature after the global branch feature is aligned by bilinear interpolation, b final : The fused multi-scale feature.
[0064] (4) Gradual decoding: Use skip connections to combine global and local features, and gradually restore the resolution:
[0065] Y output = Decoder(b final , e local , e global ).
[0066] Y output : The finally output density map; e local : The intermediate feature of the local branch; e global : The intermediate feature of the global branch.
[0067] (5) Output density map: The finally output Y output ∈R 1×H×W , with the size of 1×H×W, representing the density estimation value.
[0068] See Figure 1 The specific implementation manners of the present invention are shown as follows. Among them, the solid arrows represent the data flow direction, and the specific structure of the deep learning network is within the wireframe. The entire model architecture includes a global image input processing module Global input, a feature fusion module Concat, a local image input processing module Local input, a decoding module attention mechanism CBAM, and an output module Output; further, in the present invention, the global image input processing module Global input is used to receive a global image (low resolution, capturing macroscopic semantics), and then connect to a global branch. The global branch consists of multiple convolutional and pooling operations, gradually reducing the resolution to extract global context features (such as the distribution trend of weeds in a large area), and finally outputting a low-resolution feature map.
[0069] The local image input processing module Local input, (high resolution, retaining details), then connects to a local branch. The local branch retains the high-resolution spatial information and focuses on local detail features (such as the texture and leaf morphology of a single weed plant).
[0070] The attention mechanism CBAM is embedded in the feature fusion module and includes channel attention (ChannelAttention) and spatial attention (Spatial Attention). The channel attention dynamically assigns channel weights through global pooling to strengthen the features related to weeds; the spatial attention generates a weight map through spatial pooling and convolution to highlight the significant areas of weeds and suppress background interference.
[0071] Feature Fusion Module Concat: Align the global and local feature sizes through bilinear interpolation, enhance the concatenated features through CBAM, and then reduce the dimension through 1×1 convolution to output the fused multi-scale features.
[0072] Decoding Module: Upsample + Conv3x3, which combines skip connections to gradually upsample, fuse the low-level high-resolution features of the global branch and the deep semantic features of the local branch, and finally restore to the input resolution to generate a density map. Further, Conv3x3(64->128): 3×3 convolution, and the number of channels increases from 64 to 128.
[0073] Output Module Output: Output a density map with a size of H×W×1, and the pixel value represents the relative weed density of the corresponding area. In the present invention, ei: the feature extracted from the i-th layer. Resize: Change the size of the feature map output by the global branch to match the size of the feature map output by the local branch.
[0074] In addition, (1226x1979->613x990): Halve the resolution through pooling (such as max pooling, stride 2).
[0075] In the present invention, the following introduces the estimation method of the weed density estimation model based on a deep learning model through specific Example 1. Based on a deep learning model (CombinedUNet combined with CBAM), estimate the weed density of the segmented farmland grid image patches, and output the weed density map and the statistical result of the weed quantity of each grid, providing data support for precise spraying. The following are the detailed implementation steps:
[0076] 1. Input image patches
[0077] Data source: Image patches generated by the image grid division module, with the file format of.png. The size of each image patch is (813, 1283) pixels, corresponding to the actual farmland area of (1m×1.5m). Each image patch has GPS positioning information, which is convenient to associate the output result with the actual geographical location.
[0078] Loading and preprocessing:
[0079] (1) Normalization: Scale the pixel values from [0, 255] to [0, 1] to unify the input data distribution.
[0080] (2) Tensor conversion: Use a deep learning framework (such as PyTorch or TensorFlow) to convert the image into tensor format.
[0081] (3) Batch loading: Organize multiple image patches into a batch to improve the model inference efficiency.
[0082] 2. Design the CombinedUNet model architecture
[0083] (1) Global branch:
[0084] Design goal: Capture the macroscopic features of the image (such as the large - scale weed distribution trend).
[0085] Structure: It contains 4 layers of convolutional pooling, each layer consists of 2 convolutional blocks to extract large - scale features. The feature maps of each layer are gradually reduced in resolution through the max - pooling layer to encode the global context information.
[0086] (2) Local branch:
[0087] Design goal: Focus on the local details of the image (such as the texture or leaf features of individual weeds).
[0088] Structure: It contains 3 layers of shallow convolutions to extract small - scale features. No pooling operation is performed to maintain the original resolution.
[0089] (3) CBAM attention mechanism:
[0090] Channel Attention Module: Generate channel weights through global average pooling and max - pooling. Assign weights to different channels to highlight weed - related features.
[0091] Spatial Attention Module: Generate spatial weights based on average pooling and max - pooling among channels. Strengthen the saliency of weeds in spatial positions and ignore background interference.
[0092] CBAM is embedded after each layer of convolution in the encoder to dynamically adjust the feature weights.
[0093] (4) Feature fusion: The features of the global branch and the local branch are fused in the decoding stage. Use Skip Connections to transfer the high - resolution features in the encoder to the decoder to improve feature fidelity.
[0094] (5) Decoder: Gradually upsample to restore the original resolution. Each layer of the decoder receives the skip - connection features from the encoder and fuses them with the features of the current decoding layer.
[0095] 3. Weak - supervised learning training
[0096] (1) Data preparation:
[0097] Annotated data: A small number of grid blocks have real weed density labels for supervised training.
[0098] Unlabeled data: Most grid blocks only have the label of the total number of weeds, and pseudo-label generation technology is used to enhance the supervision signal.
[0099] (2) Training process:
[0100] Use part of the accurately labeled data for preliminary training.
[0101] Compare the density map predicted by the model with the true total number of weeds, and optimize the model through weak supervision loss (such as KL divergence).
[0102] Iteratively generate pseudo-labels and conduct joint training in combination with the original labeled data to improve the generalization ability of the model.
[0103] (3) Loss function:
[0104] Density map loss (Pixel-wise Loss): Measure the pixel-level difference between the predicted density map and the true density map (such as mean square error).
[0105] Total weed count loss (Global Count Loss): Constrain the sum of the predicted density map to be close to the true number of weeds.
[0106] Joint loss: Weighted sum of the above two losses to ensure that the model takes into account both pixel-level accuracy and global consistency.
[0107] 4. Model output
[0108] (1) Weed density map:
[0109] The output size is the same as the input image (813, 1283), and each pixel value represents the relative weed density in that area.
[0110] Pixels in high-density areas have higher values, and pixels in low-density or weed-free areas are close to 0.
[0111] (2) Weed count statistics:
[0112] Estimate the total number of weeds in the grid area by accumulating all pixel values of the weed density map.
[0113] 5. Deployment and inference
[0114] Deploy the trained model on a cloud server or a local high-performance workstation. Support batch inference, input multiple image blocks each time to accelerate the processing speed. Each image block is input into the model in turn to generate a density map and a weed count estimate. Return the inference results through an API or store them in a file for use by subsequent modules. See Figure 2 as shown, for an example of obtaining a weed density map.
[0115] See Figure 2Among them, the density map has the same resolution as the input UAV image. There are 16 image patches in the map, and the size of each image patch is (813, 1283) pixels. The color mapping (such as grayscale or pseudo-color) corresponding to the actual farmland area (1m × 1.5m) intuitively shows the weed distribution.
[0116] Among them, the color mapping rule (jet color mapping): blue (0.0): no weeds or extremely low density; cyan (0.25): low-density area; green (0.5): medium-density area; yellow (0.75): medium-high density area; red (1.0): highest density area (center of the area); color gradient: reflects the density attenuation trend of weeds spreading from the center to the outside.
[0117] Pixel value: The value of each pixel is the relative density, ranging from [0, 1]; value = 1: the maximum weed density in the corresponding area; value = 0: no weed area; total number of weeds: The area (integral) of the density map of each area is proportional to the number of weeds recorded. The total number of each image patch can be approximately estimated by accumulating all pixel values.
[0118] By accumulating all pixel values, the number of weeds from left to right and from top to bottom is as follows: 6, 39, 51, 28, 9, 52, 19, 41, 13, 24, 12, 22, 10, 27, 25, 28. Through the output density map, the weed distribution and density in the overall area can be intuitively seen, and the number of weeds in each image patch, that is, the actual area (1m × 1.5m), can be obtained by accumulating pixel values, providing specific data guidance for weed management.
[0119] The above-described embodiments merely represent several implementation manners of the present invention. Their descriptions are relatively specific and detailed, but they should not be construed as limiting the scope of the invention patent. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the present invention patent shall be subject to the appended claims.
Claims
1. A weed density estimation model based on a deep learning model, characterized in that: The model is the CombinedUNet model architecture, including an input processing module, a feature extraction module for the global branch and the local branch, a feature fusion module, a decoding module, and an output density map module; The model inputs the global image and the local image through the input processing module, then extracts features from the input image under the feature extraction modules of the global branch and the local branch, then performs splicing and feature fusion through the feature fusion module, and the decoding module performs step-by-step decoding, and finally outputs the density map.
2. The weed density estimation model based on the deep learning model according to claim 1, wherein: The input sizes of both the global image and the local image are 3×H×W, where the global image is denoted as X global ∈R 3×H×W , and the local image is denoted as X local ∈R 3×H×W , where H and W are the height and width of the image respectively.
3. The estimation method of the weed density estimation model based on the deep learning model according to claim 1 or 2, characterized in that: The global branch is used to extract low-resolution global semantic features, expressed as: b global, [e global [1],e global [2],e global [3],e global [4],] = UNetGlobalBranch(X global ); The local branch is used to extract high-resolution local detail features, expressed as: b local, [e local [1],e local [2],e local [3],e local [4],] = UNetLocalBranch(X local )。 4. The weed density estimation model based on the deep learning model according to claim 1, characterized in that: The specific method for the feature fusion module to perform feature fusion is as follows: (1) Feature alignment: Use bilinear interpolation to adjust the global branch feature size to be the same as the output size of the local branch: b global →F.interpolate(b global , size=b local.shape [2:]); (2) Feature concatenation: Concatenate the local feature and the global feature, expressed as b combined =Concat(b local , b global_resized , dim=1), and the number of channels after concatenation is 2048; (3) Use the channel attention mechanism and spatial attention mechanism of CBAM for enhancement and output the fused feature; (4) Use 1×1 convolution for dimensionality reduction processing and output the feature size.
5. The weed density estimation model based on the deep learning model according to claim 4, characterized in that: Among them, the channel attention mechanism is a processing method that generates channel weights through global average pooling and max pooling. The formula is: F CA = F · σ(MLP(GAP(F))), where MLP: fully connected layer, GAP(F): global average pooling; The spatial attention mechanism generates spatial weights through the combination of average pooling and max pooling with convolution. The formula is: F SA = F CA ·σ(Conv2D(Concat(Mean(F CA ), Max(F CA )))); Feature b after enhanced feature output enhanced , denoted as b enhanced = CBAM(b combined ).
6. The weed density estimation model based on a deep learning model according to claim 4 or 5, characterized in that: The specific steps of the channel attention mechanism are as follows. (1) Set the goal and generate channel weights according to the assigned weights of each channel. (2) Input feature, the feature is X ∈ R C×H×W , C: number of channels; H, W: The height and width of the feature map; (3) Perform global pooling operation to output two-channel description vectors, and the vectors are Output z avg , z max ∈R C ; (4) Shared fully-connected network, input z avg and z max into a shared two-layer fully-connected network respectively to obtain two weight results; (5) Weight the two weight results and sum them to obtain the final channel weight, s c = s avg + s max ; Then multiply each channel of the input feature XX by the corresponding weight to generate the enhanced feature map X CA = X ⊙ s c , where ⊙ is element-wise multiplication of channels.
7. The weed density estimation model based on a deep learning model according to claim 4 or 5, characterized in that: The specific steps of the spatial attention mechanism are as follows. (1) Set the goal, highlight the weed information in specific regions of the image in the spatial dimension, and generate the spatial weight formula. (2) Input feature, the feature is X CA ∈R C×H×W , C: number of channels; H, W: The height and width of the feature map; (3) Perform pooling processing in the spatial dimension, and output two spatial feature vectors. The formula is X avg ,X max ∈R 1×H×W ; (4) Concatenation and convolution processing, concatenating two spatial feature maps in the spatial dimension to form a 2-channel feature map: X S = Concat(X avg , X max ), X S ∈R 2×H×W ; (5) Input a convolutional layer with a convolution kernel of k×k to generate a spatial weight map W S ∈R 1×H×W , and then apply the spatial weights to the input feature map, X SA =X CA ⊙W S , where ⊙ is element-wise multiplication along the channels.
8. A method for estimating weed density based on a deep learning model, characterized in that: The specific steps of the method are as follows: (1) Input image patches Collect and obtain image patches in image format, and each image patch corresponds to the GPS positioning information of the target. (2) Image preprocessing Perform normalization processing on the image patches in (1), scale the pixel values from [0, 255] to [0, 1] to unify the input data distribution; then perform tensor conversion, and convert the image into tensor format using the deep learning framework; then perform batch loading, and organize multiple image patches into a batch. (3) Build the model Build the weed density estimation model based on the deep learning model described in any one of claims 1-7. (4) Learning and training Prepare the annotation for the image data, including annotated data and unannotated data; then, use a part of the annotated data for preliminary training, compare the density map predicted by the model with the true total weed quantity, optimize the model through the weakly supervised loss, then iteratively generate pseudo-labels, and combine with the original annotated data for joint training; finally, adjust the pixel-level accuracy and consistency of the model through the weighted sum of two losses, namely the density map loss function and the total weed quantity loss function. (5) Model output Output the weed density map, where the output pixel value size indicates the relative weed density; Then calculate the total weed quantity in the grid area by accumulating all the pixel values of the weed density map.
9. The method for estimating weed density based on a deep learning model according to claim 8, wherein: When estimating the total number of weeds through the density map in (5), all pixel values in the density map are accumulated. where D i is the value of the i-th pixel in the density map, and N is the total number of pixels in the density map.