Agricultural pest detection method and system based on adaptive sparse convolution and feature fusion

The adaptive sparse convolution and feature fusion method addresses the computational challenges of target detection in agriculture by optimizing sparse levels and integrating foreground and background features, ensuring efficient and accurate detection in resource-constrained environments.

CN120318679APending Publication Date: 2025-07-15JIANGSU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510380516.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-15

AI Technical Summary

Technical Problem

Existing target detection methods in agriculture face challenges with high computational complexity and resource demands, particularly in resource-constrained environments, due to deep convolutional networks and Transformer-based models, which hinder real-time detection and deployment in embedded systems.

Method used

An adaptive sparse convolution and feature fusion method is employed, utilizing a self-adaptive multi-layer mask ratio strategy and a difference-guided lightweight feature fusion approach to optimize feature extraction and integration, reducing computational load while maintaining detection accuracy.

Benefits of technology

The method achieves efficient and accurate target detection by dynamically adjusting sparse levels and integrating foreground and background features, reducing computational resources and enhancing detection performance in resource-limited agricultural settings.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318679A_ABST
    Figure CN120318679A_ABST
Patent Text Reader

Abstract

The invention discloses an adaptive sparse convolution and feature fusion-based agricultural pest detection method and system, and the method comprises the steps: carrying out the sparse convolution of features of different levels of a feature pyramid through employing an adaptive multi-layer mask ratio strategy, and extracting efficient foreground features; in particular, sparse convolution is performed on an image by generating a pixel-level binary sample mask, and the binary sample mask is supervised by an optimal mask ratio of foreground features in a true value image tag, so that compact foreground feature coverage is realized by the generated binary sample mask, and the calculation amount of a model is remarkably reduced. In addition, the invention further provides a lightweight feature fusion strategy guided by difference mapping, pixel-level difference features between sparse foreground features and background features are calculated, space weights and channel weights are generated according to the difference features, the sparse foreground features and the background features are subjected to weighted fusion by using the space weights and the channel weights, and the sparse foreground features and the background features are subjected to weighted fusion. And the detection performance of the model is further improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of object detection in artificial intelligence technology, and particularly relates to an agricultural pest and disease detection method and system based on adaptive sparse convolution and feature fusion. Background Art

[0002] Object detection is an important task in computer vision, whose purpose is to detect all objects of interest in an image or image sequence and determine their position and class information. This technology plays an important role in multiple practical application fields. Especially in the field of agricultural production, it has a wide range of applications, covering multiple aspects such as crop management, pest and disease identification, weed detection, automated picking, soil analysis, livestock monitoring, and precision agriculture. It helps farmers adjust management strategies in a timely manner, improve production efficiency and yield, and promote the intelligent development of agriculture by analyzing images to identify pests and diseases, monitor crop growth, predict yields, precisely remove weeds, automate picking, optimize irrigation and fertilization, monitor animal health, and farmland conditions.

[0003] Traditional object detection methods can generally be divided into methods based on convolutional neural networks (such as YOLO, SSD, and Faster R-CNN series) and methods based on Transformers (such as DETR series). With the continuous development of deep learning technology, object detection methods based on convolutional neural networks gradually tend to adopt deeper feature extraction strategies, by introducing deeper network structures to extract deep features with more semantic information. Although this method has made significant progress in improving object detection accuracy, it is accompanied by a substantial increase in computational overhead. Especially in agricultural application scenarios with limited hardware resources, the computational complexity and memory requirements of the network have become important challenges that need to be solved urgently. In addition, although methods based on Transformers have demonstrated excellent global modeling capabilities in object detection, the self-attention mechanism of Transformers has a high computational complexity and usually requires a large amount of labeled data and long training time, which leads to a significant increase in its computational and memory overhead, thus limiting its application in real-time detection in agricultural applications and embedded systems. Therefore, although Transformers have certain advantages in dealing with complex agricultural production scenarios, their computational efficiency and training cost are still bottlenecks at present.

[0004] In recent years, the emergence of sparse convolution has brought a feasible solution to the object detection task in low-resource scenarios. Sparse convolution only performs convolution operations on the non-zero elements in the input feature map, effectively reducing the computational amount of the model and thus improving the detection efficiency. However, existing object detection models based on sparse convolution all follow the traditional method of the original sparse convolution, that is, setting a fixed mask ratio or only focusing on the foreground. The former still introduces unnecessary redundant calculations, while the latter may lead to a decline in the detection performance of the model. Therefore, the detection accuracy and detection efficiency still cannot be effectively balanced. Summary of the Invention

[0005] Aiming at the deficiencies in the prior art, the present invention provides an agricultural pest and disease detection method and system based on adaptive sparse convolution and feature fusion. The method proposes an adaptive multi-layer mask ratio strategy to perform sparse convolution on features at different levels in the feature pyramid, so as to extract efficient foreground features. Specifically, a pixel-level binary sample mask is generated to perform sparse convolution on the image, and the binary sample mask is supervised by the optimal mask ratio of the foreground features in the ground truth image label, so that the generated binary sample mask can achieve a compact foreground feature coverage, thereby significantly reducing the computational amount of the model. In addition, a large amount of semantic information closely related to the foreground features is also contained in the background features. Therefore, the method also proposes a lightweight feature fusion method guided by differential mapping, aiming to effectively fuse the background information and sparse foreground features and further improve the detection performance of the model.

[0006] The present invention realizes the above technical objectives through the following technical means.

[0007] Agricultural pest and disease detection method based on adaptive sparse convolution and feature fusion:

[0008] Send the original image of crop pests and diseases into the feature pyramid network for feature extraction and feature fusion to generate a feature pyramid;

[0009] Add a binary sample mask generated under the supervision of the optimal mask ratio of the ground truth label to each layer of features in the feature pyramid, and screen out the corresponding pixel-level features through the binary sample mask to form sparse foreground features;

[0010] Calculate the pixel-level difference features between the sparse foreground features and the background features, generate spatial weights and channel weights according to the difference features, and use the spatial weights and channel weights to perform weighted fusion on the sparse foreground features and the background features;

[0011] Send the features after weighted fusion into the bounding box subnet and the classification subnet respectively to obtain the localization result and classification result of agricultural pests and diseases.

[0012] Furthermore, the acquisition of the sparse foreground features is specifically as follows:

[0013] Adopt the mask loss function L mask for constraint: wherein, represents the mask ratio of the binary sample mask, L represents the number of layers of the feature pyramid, and P i represents the optimal mask ratio of the feature of the i-th layer of the feature pyramid, and H i represents the binary mask matrix;

[0014] Minimize the said L mask , generate a binary sample mask closer to the optimal mask ratio of the ground truth label, so as to extract sparse foreground features.

[0015] Furthermore, the binary mask matrix is obtained in the following manner:

[0016] For the feature X of the i-th layer of the feature pyramid i ∈R B×C×H×W , adopt the shared convolution kernel W mask ∈R C×1×3×3 for processing: Based on the said W mask convolve the said X i to generate a grayscale sample feature map S i ∈R B×1×H×W , and then use the Gumbel-Softmax method to convert the grayscale sample feature map into a binary mask matrix H i :

[0017]

[0018] wherein, B represents the batch size of the feature of the i-th layer, C represents the number of channels of the feature of the i-th layer, H represents the height of the feature of the i-th layer, W represents the width of the feature of the i-th layer, and g1, g2 ∈ R B×1×H×W represent two random gumbel noises, σ is the sigmoid function, and τ is the corresponding temperature parameter in Gumbel-Softmax.

[0019] Furthermore, the optimal mask ratio of the feature of the i-th layer of the feature pyramid is obtained in the following manner:

[0020] Through the localization result of the ground truth label, obtain the ground truth detection result of the feature of the i-th layer of the feature pyramid and further obtain the optimal mask ratio of the feature of the i-th layer of the feature pyramid:

[0021]

[0022] wherein, h i represents the height of the feature of the i-th layer, and w idenotes the width of the feature of the i-th layer, and Pos(C i ) denotes the number of pixels belonging to the foreground instance, and Global(C i ) denotes the total number of pixels belonging to the foreground instance.

[0023] Furthermore, the extraction of the sparse foreground feature is specifically as follows: The feature of the i-th layer of the feature pyramid uses a binary sample mask to filter out valid pixel points, and the pixel points form a sparse feature vector. The sparse feature vector is convolved with a 3×3 convolutional kernel, and the output of the convolution operation is mapped back to the two-dimensional space of the i-th layer feature to form a sparse foreground feature after sparse convolution.

[0024] Further, Gaussian blur and normalization are performed on each pixel position of the difference feature to generate a spatial weight where x and y respectively represent the position coordinates of the current pixel in the difference feature map, x′ and y′ respectively represent the position coordinates of all pixels in the difference feature map, D is the difference feature, and D = |F s -F b |, F s is the sparse foreground feature, and F b is the background feature.

[0025] Furthermore, the difference feature D is summed in the spatial dimension to obtain a channel weight where c represents the channel index of the current pixel in the difference feature map.

[0026] Furthermore, the feature after weighted fusion is:

[0027] F f = W s ⊙ ((W c ⊙ 1 c ) · F s ) + (1 - W s ) ⊙ (((1 - W c ) ⊙ 1 c ) · F b )

[0028] where ⊙ represents element-wise multiplication, and 1 c represents a vector of all 1s.

[0029] Further, the classification sub-network first performs 4 times of 3×3 convolution operations with a convolution kernel of 256, and uses the Relu activation function once after each convolution, then performs a 3×3 convolution operation with a convolution kernel of K×A, and finally uses the Softmax function to obtain the confidence of each anchor box; the bounding box sub-network first performs 4 times of 3×3 convolution operations with a convolution kernel of 256, and uses the Relu activation function once after each convolution, then performs a 3×3 convolution operation with a convolution kernel of 4×A, and finally uses the Softmax function to obtain the localization coordinates of the target to be detected; and maps the confidence of each anchor box and the localization coordinates of the target to be detected back to the original image to obtain the classification result and localization result of agricultural pests and diseases; where K is the number of categories of the target to be detected, and A is the number of anchor boxes.

[0030] An agricultural pest and disease detection system based on adaptive sparse convolution and feature fusion, comprising:

[0031] An original image feature extraction module that extracts and fuses features from the original pictures of crop pests and diseases to generate a feature pyramid;

[0032] A foreground feature extraction module that adds a binary sample mask to each layer of features of the feature pyramid, and filters out the corresponding pixel-level features through the binary sample mask to form sparse foreground features;

[0033] A foreground and background feature fusion module that generates spatial weights and channel weights according to the differential features, and then performs weighted fusion on the sparse foreground features and the background features;

[0034] A model detection module that obtains the localization result and classification result of agricultural pests and diseases from the weighted fusion features.

[0035] The beneficial effects of the present invention are:

[0036] (1) The present invention proposes an innovative adaptive multi-layer mask ratio strategy for dynamically adjusting the sparsity of features at different levels in the feature pyramid. By introducing a learnable mask ratio parameter, this strategy enables the model to adaptively adjust the sparsity of each layer of feature maps according to the feature distribution of the input samples, thereby significantly improving the efficiency and accuracy of feature extraction. Specifically, mask generation based on the ground truth image is adopted, and by analyzing the information entropy of the feature map, the optimal mask ratio of each layer of features is dynamically determined. This mechanism not only enables the model to focus on the feature regions with the largest amount of information but also effectively reduces redundant calculations. In addition, by introducing a mask loss function, the model can continuously optimize the mask generation process through end-to-end training, making the generated binary sample mask closer to the optimal mask ratio of the ground truth label, achieving more accurate foreground feature extraction, and achieving a better balance between computational efficiency and feature representation ability.

[0037] (2) The present invention proposes a lightweight feature fusion method guided by differential mapping, aiming to effectively integrate foreground features and background features. By dynamically adjusting their fusion weights according to the significant differences between foreground features and background features, this method can accurately extract foreground features while enhancing the global context information of the background region. This dynamic weighted fusion strategy can not only improve the recognition accuracy of foreground objects, but also better capture the global structure and semantic information of the background region, thus achieving a comprehensive understanding of complex scenes. In this way, the method not only significantly improves the overall performance of object detection, but also effectively reduces the consumption of computing resources, ensuring the lightweight design of the model. This efficient feature fusion mechanism effectively balances detection accuracy and computational efficiency, providing good adaptability for the real-time application and hardware acceleration of object detection tasks in agricultural application scenarios.

[0038] (3) The foreground feature extraction module and the foreground and background feature fusion module of the present invention are designed to be lightweight, plug-and-play, and are easy to be efficiently applied in various models to complete object detection tasks. The present invention has achieved the best results on multiple widely used basic object detection models. While ensuring the improvement of model detection, it reduces the computational amount of the model and greatly improves the detection efficiency. It is particularly worth pointing out that in agricultural application scenarios with low hardware resources, the technical advantages of the present invention are fully demonstrated. This characteristic of high efficiency and low resource consumption makes the present invention particularly suitable for deployment on edge computing devices and mobile terminals, providing a reliable solution for real-time pest detection and identification in the agricultural field. In addition, the modular design makes this method easy to migrate to other agricultural application scenarios, such as crop growth monitoring, weed identification, etc., and has broad application prospects and promotion value. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 is the framework diagram of the agricultural pest detection system with adaptive sparse convolution and feature fusion according to the present invention;

[0040] Figure 2 is the diagram of the adaptive multi-layer mask ratio strategy supervised by the optimal mask ratio of the ground truth label according to the present invention;

[0041] FIG. 3(a) is the original image;

[0042] FIG. 3(b) is the heat map of the detection result of the base detector RetinaNet;

[0043] FIG. 3(c) is the heat map of the detection result of the model of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0044] The present invention will be further described below in conjunction with the accompanying drawings and specific embodiments, but the protection scope of the present invention is not limited thereto.

[0045] As Figure 1 shown, the agricultural pest and disease detection system of the self-adaptive sparse convolution and feature fusion of the present invention follows the process of image feature processing, which is successively: original image feature extraction → foreground feature extraction → foreground and background feature fusion → model detection. The original picture is put into the (Feature Pyramid Network) FPN for feature extraction. By constructing a bottom-up and top-down feature fusion path, feature maps of different scales are fused to generate a feature pyramid with rich multi-scale information. In the foreground feature extraction stage, the present invention innovatively proposes an adaptive multi-layer mask ratio strategy, adding a binary sample mask generated under the supervision of the optimal mask ratio of the ground truth label to each layer of features of the feature pyramid, and screening out the corresponding pixel-level features through this binary sample mask to form sparse foreground features, so as to achieve accurate foreground feature coverage. In the foreground and background feature fusion stage, an innovative lightweight feature fusion method is proposed, guiding the fusion of sparse foreground features and background features through difference mapping. Specifically, the pixel-level difference between the sparse foreground features and the background features is calculated, and spatial and channel weights are generated according to the difference mapping, so as to perform weighted fusion on the sparse foreground features and the background features. While retaining the high sensitivity of the sparse foreground features to the target area, the global context information of the background features is fully utilized, thus significantly improving the overall performance of the model. In the model detection stage, the fused features are respectively sent into the bounding box subnet and the classification subnet. The classification subnet first performs 3×3 convolution operations with a convolution kernel of 256 four times, and uses a Relu activation function once after each convolution, and then performs a 3×3 convolution operation with a convolution kernel of K×A, where K is the number of categories of the target to be detected and A is the number of anchor boxes. Finally, the Softmax function is used to obtain the confidence of each anchor box; correspondingly, the bounding box subnet first performs 3×3 convolution operations with a convolution kernel of 256 four times, and uses a Relu activation function once after each convolution, and then performs a 3×3 convolution operation with a convolution kernel of 4×A, and finally uses the Softmax function to obtain the localization coordinates of the target to be detected; and the confidence of each anchor box and the localization coordinates of the target to be detected are mapped back to the original image to obtain the classification result and localization result of agricultural pests and diseases.

[0046] With the above detection system, the pest and disease detection model of the present invention includes an original image feature extraction module, a foreground feature extraction module, a foreground and background feature fusion module, and a model detection module.

[0047] The present invention is carried out in a given base detector RetinaNet, using a large-scale benchmark dataset IP102 for crop pest and disease identification. The IP102 dataset is a crop pest and disease dataset for pest and disease classification and detection tasks, containing more than 75,000 images of 8 crops, including rice, corn, wheat, sugar beet, alfalfa, grape, citrus and mango, and 102 types of pests. These images present a natural long-tail distribution. Considering the difficulty of marking bounding boxes, IP102 randomly selects some images from each type of pest image to form a data subset for the target detection task, including a total of 18,983 images that can be used for crop pest target detection tasks, and divides these images into training sets and test sets in a ratio of 4:1.

[0048] The system architecture includes a foreground feature extraction branch and a foreground and background feature fusion branch. In the foreground feature extraction branch, a pixel-level binary sample mask generated under the supervision of the optimal mask ratio of the true value label is introduced to perform sparse convolution on the image to effectively extract the foreground features; the pixel-level difference between the sparse foreground features and the background features is calculated, and spatial and channel weights are generated based on the difference mapping, so as to perform weighted fusion of the sparse foreground features and the background features. The specific process is as follows:

[0049] 1) Extract and fuse the features of the original images in the training set to generate a feature pyramid; Figure 2 As shown, for the feature X of the i-th layer of the feature pyramid i ∈R B×C×H×W , using shared convolution kernel W mask ∈R C×1×3×3 Processing is performed, where B, C, H, and W represent the batch size, number of channels, height, and width of the features of the i-th layer respectively: First, based on W mask X i Convolution generates grayscale sample feature map S i ∈R B×1×H×W , and then the Gumbel-Softmax method is used to further transform the grayscale sample feature map into a binary mask matrix H i ∈{0,1} B×1×H×W , specifically expressed as:

[0050]

[0051] Among them, g1, g2∈R B×1×H×W represents two random gumbel noises, σ is the sigmoid function, and τ is the corresponding temperature parameter in Gumbel-Softmax.

[0052] 2) According to the above formula, during the training and detection phases of the pest and disease detection model, sparse convolution is only performed on the area with a mask value of 1, thereby significantly reducing the overall computational cost. However, in the absence of any additional constraints, common sparse convolution-based detectors tend to generate masks with larger mask ratios in order to obtain higher accuracy, which increases the overall computational cost to a certain extent and cannot flexibly adapt to the feature distribution of different scenes, limiting the optimization space of model performance. In order to solve this problem, the sparsity in the present invention is controlled by the optimal mask ratio of the true value label. Figure 2 As shown in Figure 1, the adaptive multi-layer mask ratio strategy first estimates the optimal mask ratio based on the true value label: through the positioning result of the true value label (provided by the IP102 dataset), the true value detection result of the feature pyramid layer i is obtained Among them, h i 、w i They represent the height and width of the i-th layer feature respectively, and the optimal mask ratio of the i-th layer feature of the feature pyramid is further obtained, which is specifically expressed as:

[0053]

[0054] Among them, Pos(C i ) and Global(C i ) represent the number of pixels belonging to the foreground instance and the total number of pixels, respectively.

[0055] 3) Then, in order to guide the model to generate a mask close to the optimal mask ratio of the true value label, the mask loss function L is used mask Constraints are specifically expressed as:

[0056]

[0057] in, represents the mask ratio of the binary sample mask, and L represents the number of layers of the feature pyramid.

[0058] By minimizing L mask , generating a binary sample mask that is closer to the optimal mask ratio of the true value label, thereby extracting better sparse foreground features and achieving better coverage of the foreground area.

[0059] Among them, sparse foreground features are extracted. Specifically: the i-th layer features of the feature pyramid use binary sample masks to filter out valid pixels, and these pixels form sparse feature vectors; then these sparse feature vectors are convolved with a 3×3 convolution kernel (the calculation results only correspond to the valid positions marked by the binary sample mask); finally, the output of the convolution operation is mapped back to the two-dimensional space of the i-th layer features to form sparse foreground features after sparse convolution.

[0060] 4) In the foreground and background feature fusion branch, the present invention proposes a lightweight feature fusion method guided by difference mapping. The guiding model dynamically adjusts the fusion weights of the foreground and background features according to the significant differences between them, ensuring the accurate extraction of foreground target features while enhancing the global context information of the background region, and ultimately improving the detection accuracy of the model. Specifically, first, the pixel-level difference feature D between the sparse foreground feature and the background feature is calculated, which is specifically expressed as:

[0061] D = |F s - F b | (4)

[0062] where F s is the sparse foreground feature and F b is the background feature.

[0063] 5) Then, Gaussian blur and normalization are performed on each pixel position of the difference feature D to generate the spatial weight map W s ∈ R C×H×W , which is specifically expressed as:

[0064]

[0065] where x and y respectively represent the position coordinates of the current pixel in the difference feature map, and x' and y' respectively represent the position coordinates of all pixels in the difference feature map.

[0066] The spatial weight map W s can highlight the regions with large differences between the sparse foreground feature and the global background feature, ensure smooth weight distribution, and suppress the influence of noise.

[0067] 6) The difference feature D is summed in the spatial dimension to obtain the channel-level weight W c ∈ R C , which is specifically expressed as:

[0068]

[0069] where c represents the channel index of the current pixel in the difference feature map.

[0070] The channel weight W c is used to capture the saliency of the foreground and background in different feature dimensions.

[0071] 7) The spatial weight W s and the channel weight W c are used to perform weighted fusion on the sparse foreground feature and the background feature to obtain the fused feature F f , which is specifically expressed as:

[0072] F f = Ws ⊙((W c ⊙1 c )·F s )+(1 - W s )⊙(((1 - W c )⊙1 c )·F b ) (7)

[0073] Among them, ⊙ represents element-wise multiplication, and 1 c represents a vector of all 1s, which is used to expand the weight dimension.

[0074] The globally weighted and fused feature F f ∈ R C×H×W contains both foreground and background features extracted by sparse convolution, which can improve the understanding ability of global context information. In addition, the feature fusion method guided by differential mapping does not rely on complex attention mechanisms, has simple and efficient calculations, and further realizes the lightweight design of the model.

[0075] 8) The loss function for pest and disease detection is defined as follows:

[0076] L′ = L det + α·L mask (8)

[0077] Among them, L det is a loss function composed of the classification loss and regression loss of RetinaNet, and α is a hyperparameter that balances the mask importance of L;

[0078] The above loss function is used to jointly optimize the class prediction and bounding box position prediction of the pest and disease detection model for the target.

[0079] The foreground feature extraction module and the foreground and background feature fusion module are two sets of independent and learnable parameterized modules. During the foreground feature extraction process, the optimal mask ratio is first estimated based on the ground truth label, so as to guide the network to generate a sample mask close to the optimal mask ratio of the ground truth label, and then sparse convolution is performed on the original image features, which greatly reduces the computational redundancy of the model and improves the detection efficiency of the model. In addition, the foreground feature extraction module and the foreground and background feature fusion module of the present invention are a plug-and-play feature extraction module, which can be embedded into any computer vision model to achieve efficient feature calculation to complete tasks such as classification and detection, while maintaining a low computational cost.

[0080] To verify the effectiveness of the present invention, a heatmap visualization analysis of the pest and disease detection performance was carried out on the IP102 test set. Specifically, the trained pest and disease detection model was used to perform a visualization analysis of the detection and localization results of randomly selected test samples. Figure 3(a) is the original image, Figure 3(b) is the heatmap of the detection results of the base detector RetinaNet, and Figure 3(c) is the heatmap of the detection results of the model of the present invention. By comparing the results of the heatmaps in Figure 3(b) and Figure 3(c), it can be seen that during the detection process of the model proposed by the present invention, the shape of the key attention area is closer to the characteristics of the target object itself. Combining with the analysis of the crop pest target detection task, when the pests to be detected have special backgrounds corresponding to their categories (such as tree trunks, leaves, etc.), the present invention can effectively distinguish the foreground from the background, capture more background difference features, and thus more accurately identify the pests and their species. This further shows that the present invention can learn the deep information in the features (such as pest shape, inhabiting background, etc.) and extract more useful features for the crop pest and disease target detection task.

[0081] The described embodiments are the preferred embodiments of the present invention, but the present invention is not limited to the above embodiments. Without departing from the essential content of the present invention, any obvious improvements, substitutions or modifications that those skilled in the art can make all fall within the protection scope of the present invention.

Claims

1. An agricultural pest and disease detection method based on adaptive sparse convolution and feature fusion, characterized in that: The original images of agricultural pests and diseases are sent into a Feature Pyramid Network (FPN) for feature extraction and feature fusion to generate a feature pyramid; A binary sample mask generated under the supervision of the optimal masking ratio of the ground truth label is added to each layer of features of the feature pyramid, and the corresponding pixel-level features are screened out through the binary sample mask to form sparse foreground features; The pixel-level difference features between the sparse foreground features and the background features are calculated, and spatial weights and channel weights are generated according to the difference features. The sparse foreground features and the background features are weighted and fused using the spatial weights and channel weights; The weighted and fused features are respectively sent into a bounding box subnet and a classification subnet to obtain the localization result and classification result of agricultural pests and diseases.

2. The agricultural pest and disease detection method according to claim 1, wherein The acquisition of the sparse foreground features is specifically as follows: Adopt the mask loss function L mask for constraint: wherein, represents the mask ratio of the binary sample mask, L represents the number of layers of the feature pyramid, and P i represents the optimal mask ratio of the feature at the i-th layer of the feature pyramid, and H i represents the binary mask matrix; Minimize the L mask , generate a binary sample mask that is closer to the optimal mask ratio of the true label, so as to extract sparse foreground features.

3. The agricultural pest and disease detection method according to claim 2, characterized in that The binary mask matrix is obtained through the following method: For the feature X of the i-th layer of the feature pyramid i ∈R B×C×H×W , a shared convolution kernel W mask ∈R C×1×3×3 is used for processing: Based on the W mask to convolve the X i to generate a grayscale sample feature map S i ∈R B×1×H×W , and then the Gumbel-Softmax method is used to convert the grayscale sample feature map into a binary mask matrix H i : Among them, B represents the batch size of the i-th layer features, C represents the number of channels of the i-th layer features, H represents the height of the i-th layer features, W represents the width of the i-th layer features, and g1, g2 ∈ R B×1×H×W represent two random Gumbel noises, σ is the sigmoid function, and τ is the corresponding temperature parameter in Gumbel-Softmax.

4. The agricultural pest and disease detection method according to claim 2, characterized in that, The optimal masking ratio of the i-th layer of features of the feature pyramid is obtained through the following method: Obtain the ground-truth detection result of the features at the $i$-th layer of the feature pyramid through the positioning result of the ground-truth label Further obtain the optimal mask ratio of the features at the $i$-th layer of the feature pyramid: Among them, h i represents the height of the feature of the i-th layer, w i represents the width of the feature of the i-th layer, Pos(C i ) represents the number of pixels belonging to the foreground instance, Global(C i ) represents the total number of pixels belonging to the foreground instance.

5. The agricultural pest and disease detection method according to claim 2, wherein, The extraction of the sparse foreground features is specifically as follows: The effective pixel points are screened out from the i-th layer of features of the feature pyramid using the binary sample mask, and the pixel points form a sparse feature vector. The sparse feature vector is convolved with a 3×3 convolutional kernel, and the output of the convolution operation is mapped back to the two-dimensional space of the i-th layer of features to form the sparse foreground features after sparse convolution.

6. The agricultural pest and disease detection method according to claim 1, wherein Perform Gaussian blur and normalization on each pixel position of the difference feature to generate spatial weights where x and y respectively represent the position coordinates of the current pixel in the difference feature map, x' and y' respectively represent the position coordinates of all pixels in the difference feature map, D is the difference feature, and D = |F s - F b |, F s is the sparse foreground feature, and F b is the background feature.

7. The agricultural pest and disease detection method according to claim 6, characterized in that, Sum the difference features D in the spatial dimension to obtain the channel weights where c represents the channel index of the current pixel in the difference feature map.

8. The agricultural pest and disease detection method according to claim 7, characterized in that, The weighted and fused features are: F f = W s ⊙ ((W c ⊙ 1 c ) · F s ) + (1 - W s ) ⊙ (((1 - W c ) ⊙ 1 c ) · F b ) where, ⊙ represents element-wise multiplication, and 1 c represents a vector of all 1s.

9. The agricultural pest and disease detection method according to claim 1, characterized in that The classification subnet first performs 4 3×3 convolution operations with a convolutional kernel of 256, and uses a Relu activation function after each convolution. Then it performs a 3×3 convolution operation with a convolutional kernel of K×A, and finally uses the Softmax function to obtain the confidence of each anchor box; the bounding box subnet first performs 4 3×3 convolution operations with a convolutional kernel of 256, and uses a Relu activation function after each convolution. Then it performs a 3×3 convolution operation with a convolutional kernel of 4×A, and finally uses the Softmax function to obtain the localization coordinates of the target to be detected; and the confidence of each anchor box and the localization coordinates of the target to be detected are mapped back to the original image to obtain the classification result and localization result of agricultural pests and diseases; where K is the number of categories of the target to be detected, and A is the number of anchor boxes.

10. A system for implementing the agricultural pest and disease detection method according to any one of claims 1-9, characterized in that, It includes: An original image feature extraction module that performs feature extraction and feature fusion on the original images of agricultural pests and diseases to generate a feature pyramid; A foreground feature extraction module that adds a binary sample mask to each layer of features of the feature pyramid, and screens out the corresponding pixel-level features through the binary sample mask to form sparse foreground features; A foreground and background feature fusion module that generates spatial weights and channel weights according to the difference features, and then performs weighted fusion on the sparse foreground features and the background features; A model detection module that obtains the localization result and classification result of agricultural pests and diseases from the weighted and fused features.