A pest detection method for a smart agriculture scene
Patent Information
- Application Number
- CN202610918297.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-24
- Publication Date
- 2026-09-18
AI Technical Summary
[0004]由于田间害虫普遍具有目标尺寸微小、种类繁多、背景环境复杂的特点,直接套用通用YOLOv5模型存在小目标边缘特征利用不足、多尺度特征融合适配性差、模型参数量偏高难以适配边缘设备部署的核心问题,难以在复杂田间环境下同时实现高精度、低时延的害虫检测,无法满足智慧农业场景下规模化虫情智能监测的实际需求
本发明通过设置HPFANet高频增强多尺度特征融合网络强化微小害虫边缘细节并自适应融合多尺度特征,配合轻量化LSDHead解耦检测头降低模型计算开销,再结合GIoU边界损失、置信度损失与Focal类别损失组成的联合损失函数动态优化模型训练过程,完整配套标准化图像采集预处理、虫情统计与远程预警输出流程,有效解决传统YOLOv5模型应用于农田场景时微小害虫漏检率高、复杂背景干扰强、多尺度目标适配性差、模型难以部署在边缘设备、样本不均衡导致识别精度不足等技术问题,能够在田间、温室等真实农业场景完成高精度、低延迟的自动化害虫监测与虫害预警,适配规模化智慧农业虫情管控需求,具有较好的实用价值与应用前景。
Smart Images

Figure CN122780955A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart agriculture pest detection technology, specifically to a pest detection method for smart agriculture scenarios. Background Technology
[0002] Pest infestations are a significant factor restricting agricultural production efficiency. Early and accurate monitoring and intelligent early warning of field pests are crucial for ensuring crop yields and achieving reduced pesticide use while increasing efficiency. Traditional pest monitoring mainly relies on manual field inspections and identification statistics, which suffers from low efficiency, large subjective errors, and high labor costs, making it difficult to meet the real-time and refined monitoring needs of large-scale smart agriculture. With the rapid development of computer vision and deep learning technologies, image-based pest identification methods based on target detection algorithms are gradually being applied to agricultural production scenarios, providing a feasible technical path for automated pest monitoring.
[0003] Among existing target detection algorithms, the YOLOv5 series combines the advantages of fast detection speed and balanced recognition accuracy, making it the mainstream algorithm framework for current industrial and agricultural applications. Several studies have already introduced the YOLOv5 algorithm into the field of farmland pest detection, achieving automated identification of common pests through fine-tuning of the basic model, thus improving the efficiency of pest monitoring to some extent. However, existing solutions are mostly based on direct adaptation to general target detection architectures, without fully integrating the morphological characteristics of farmland pests and the specific characteristics of the detection environment for in-depth optimization.
[0004] Because field pests are generally characterized by small target size, wide variety, and complex background environment, directly applying the general YOLOv5 model has core problems such as insufficient utilization of small target edge features, poor adaptability of multi-scale feature fusion, and high model parameter quantity, making it difficult to adapt to edge device deployment. It is difficult to achieve high-precision and low-latency pest detection in complex field environments at the same time, and cannot meet the actual needs of large-scale intelligent pest monitoring in smart agriculture scenarios. Summary of the Invention
[0005] To address the above technical problems, this invention provides a pest detection method for smart agriculture scenarios, the method comprising the following steps: S1. Collect images of pests in the field, preprocess and augment the images, and construct a standardized pest detection dataset; S2, Constructing an Improved Version Pest detection model, the model is based on Based on the infrastructure, retain Feature extraction network, which will integrate the original Structure replacement A high-frequency enhanced multi-scale feature fusion network replaces the original detection head with... Lightweight decoupling detection head; S3, Through High-frequency enhanced multi-scale feature fusion network The output multi-scale features are enhanced and adaptively fused, and the fused features are then input. A lightweight, decoupled detection head yields detection prediction results. S4. The pest detection model is iteratively trained using a joint loss function to optimize the model parameters. S5. Input the field image to be detected into the trained pest detection model, output the detection results of the pest target, and complete the pest situation statistics and early warning push.
[0006] As a further improvement of the present invention, step S1 specifically includes the following steps: Collect images of field pests covering various crop types, lighting conditions, shooting angles, and insect species, while simultaneously collecting images of pest-free healthy crops. The original image is processed sequentially by cropping invalid regions, scaling and normalizing, enhancing illumination and color, and suppressing noise to obtain a standardized image with a uniform input size; Multiple data augmentation strategies are employed to expand sample size and improve class imbalance. The pest targets are labeled with bounding boxes and categories, and then divided into training set, validation set and test set.
[0007] As a further improvement of the present invention, in step S3, the feature enhancement processing specifically includes the following steps: The basic convolutional features of the input feature map are extracted by cascaded convolution, and the edge response features of the feature map are extracted by edge operators. The two are weighted and fused to obtain high-frequency enhanced features, which strengthen the detailed outline of the pest target. A combined channel attention and spatial attention mechanism is adopted to assign weights to high-frequency enhanced features, thereby enhancing the feature response of the target area of pests and suppressing crop background interference. The formula for calculating the high-frequency enhancement feature is as follows: In the formula, For the input feature map, This is the feature map after high-frequency enhancement. This represents the edge response weighting coefficient.
[0008] As a further improvement of the present invention, the adaptive fusion processing in step S3 specifically includes the following steps: Align the shallow, medium, and deep multi-scale features output by Backbone to the same spatial size; Calculate the information entropy of features at each scale, by The function obtains the corresponding fusion weights; Adaptive fusion features are obtained by weighted summation of multi-scale features based on fusion weights. The formula for calculating the adaptive fusion feature is as follows: In the formula, These are the aligned multi-layer features, The fusion weights are for the corresponding scales.
[0009] As a further improvement of the present invention, step S3 further includes a reparameter optimization and detection head prediction step, specifically: A reparameterized design is adopted for the training-inference phase structure. In the training phase, a multi-branch convolutional structure is used to enhance the feature representation capability. In the inference phase, the multi-branch structure is equivalently fused into a single convolutional structure to improve the inference efficiency of edge devices. Input the fused features The lightweight decoupled detection head adopts a decoupled structure of classification and regression branches. The regression branch introduces a channel shuffling structure to enhance the interaction of localization information, while the classification branch uses depthwise separable convolution to reduce the amount of computation. Anchor frame parameters adapted to the size of farmland pests are generated based on clustering methods, and a lightweight activation function is used in the detection head to reduce computational complexity.
[0010] As a further improvement of the present invention, in step S4, the construction of the joint loss function specifically includes the following steps: respectively Loss optimization for bounding box localization stability, binary cross-entropy loss optimization for target confidence assessment. Loss optimization improves the recognition performance of hard-to-classify samples; The three types of loss terms are weighted and combined to form a joint loss function for iterative model training; The formula for calculating the joint loss function is as follows: In the formula, for Bounding box regression loss, For target confidence loss, for Category loss, These are the weighting coefficients for the corresponding loss terms.
[0011] As a further improvement of the present invention, a dynamic adjustment step for loss weights is also included: The weight coefficients of each loss term are dynamically adjusted based on the detection results of the validation set. In the early stage of training, the weight of the bounding box loss is increased to strengthen the localization learning of small targets. In the middle and later stages of training, the weights of the category loss and confidence loss are increased to improve classification accuracy and false detection suppression capability.
[0012] As a further improvement of the present invention, step S5 specifically includes the following steps: The model output results are filtered for confidence and deduplicated by nonmaximum suppression, and the output pest target category, location, quantity and confidence information are obtained. The detection results are visualized and overlaid onto the original image, and insect infestation statistics reports are generated at preset intervals. When the number of pests exceeds a preset threshold within a unit of time, an pest infestation warning is generated and pushed to the user terminal or agricultural IoT platform.
[0013] Based on this, a pest detection system for smart agriculture scenarios is provided. The system is used to implement the above-mentioned method and includes a data acquisition and preprocessing module for acquiring field pest images and completing image preprocessing, data augmentation, and detection dataset construction. An improved HIF-YOLOv5 model building module is used to build a pest detection model that includes a Backbone feature extraction network, an HPFANet high-frequency enhanced multi-scale feature fusion network, an LSDHead lightweight decoupled detection head, and a detection output layer. The model training and loss function optimization module is used to iteratively train the pest detection model using a joint loss function to achieve dynamic optimization of the model parameters. The pest detection and result output module is used to input field images to be detected and output detection results, completing the visualization and statistical analysis of pest conditions and remote early warning push.
[0014] Based on this, the present invention also provides a computer-readable storage medium storing a computer program that, when executed, implements the above-described method.
[0015] Compared with the prior art, the beneficial effects of the present invention are as follows: This invention enhances the edge details of tiny pests by setting up an HPFANet high-frequency enhanced multi-scale feature fusion network and adaptively fuses multi-scale features. It also reduces model computational overhead by using a lightweight LSDHead decoupled detection head. Furthermore, it dynamically optimizes the model training process using a joint loss function composed of GIoU boundary loss, confidence loss, and Focal category loss. The invention provides a complete and standardized workflow for image acquisition and preprocessing, pest statistics, and remote early warning output. This effectively solves the technical problems of traditional YOLOv5 models when applied to farmland scenarios, such as high false negative rates for tiny pests, strong interference from complex backgrounds, poor adaptability to multi-scale targets, difficulty in deploying the model on edge devices, and insufficient recognition accuracy due to imbalanced samples. It can achieve high-precision, low-latency automated pest monitoring and early warning in real agricultural scenarios such as fields and greenhouses, meeting the needs of large-scale smart agriculture pest management and possessing good practical value and application prospects. Attached Figure Description
[0016] To more clearly illustrate the technical solution of the present invention, the drawings used in the embodiments are briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0017] Figure 1 This is a flowchart of the method in Embodiment 1 of the present invention; Figure 2 This is a system composition block diagram in Embodiment 2 of the present invention. Detailed Implementation
[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0019] It should be noted that, unless otherwise defined, the technical or scientific terms used in the embodiments of this disclosure should have the ordinary meaning understood by one of ordinary skill in the art to which this disclosure pertains. The terms "first," "second," and similar terms used in the embodiments of this disclosure do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Terms such as "comprising" or "including" mean that the element or object preceding the word encompasses the elements or objects listed following the word and their equivalents, without excluding other elements or objects. Terms such as "connected" or "linked" are not limited to physical or mechanical connections, but can include electrical connections, whether direct or indirect. Terms such as "upper," "lower," "left," and "right" are used only to indicate relative positional relationships; when the absolute position of the described object changes, the relative positional relationship may also change accordingly.
[0020] First, some technical terms used in this invention will be explained: YOLOv5: A single-stage real-time object detection algorithm that balances detection speed and accuracy. It is a general-purpose visual detection network that can perform object category recognition and bounding box localization.
[0021] High-frequency enhancement features: The Sobel operator is used to extract high-frequency details of image contours and textures, and then fused with convolutional features to enhance the contour features of small pests and weaken the interference of messy farmland background.
[0022] HPFANet: This invention presents an improved high-frequency enhanced multi-scale feature fusion network to replace the original YOLOv5 Neck and adapt it for multi-scale feature extraction and fusion of small pests.
[0023] Information entropy: Quantifies the richness of effective information in a feature map, used to dynamically calculate the fusion weights of shallow, medium, and deep features, and adaptively adapt to pests of different sizes.
[0024] Reparameterization of structure: Multi-branch convolutions enhance feature representation during training and are merged into single convolutions during inference, accelerating inference speed on edge devices without sacrificing accuracy.
[0025] LSDHead Lightweight Decoupled Detection Head: Separates classification and localization task branches, and combines depthwise separable convolution and channel shuffling to achieve lightweight detection head, making it suitable for embedded terminal deployment.
[0026] GIoU loss: Bounding box regression loss, which solves the problem of training instability when the overlap between the predicted bounding box and the ground truth box is low for small targets, and improves the localization accuracy of tiny pests.
[0027] FocalLoss: A class loss function that reduces the weight of simple samples, focuses on difficult-to-identify small pests and rare insect samples, and alleviates the problem of imbalanced samples.
[0028] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0029] Example 1 As the background technology shows, field pests are generally characterized by small target size, wide variety, and complex background environment. Directly applying the general YOLOv5 model has core problems such as insufficient utilization of small target edge features, poor adaptability of multi-scale feature fusion, and high model parameter quantity, which makes it difficult to adapt to edge device deployment. It is difficult to achieve high-precision and low-latency pest detection in complex field environments at the same time, and cannot meet the actual needs of large-scale intelligent pest monitoring in smart agriculture scenarios.
[0030] Please see Figure 1 This invention provides a pest detection method for smart agriculture scenarios, which includes the following steps: S1. Collect images of pests in the field, preprocess and augment the images, and construct a standardized pest detection dataset; First, images of field pests covering various crop types, lighting conditions, shooting angles, and insect species are collected. Simultaneously, images of healthy crops without pests are collected to enhance the model's ability to distinguish normal crop backgrounds and reduce the false detection rate against complex backgrounds.
[0031] Preferably, the crop types include wheat, corn, vegetables, and soybeans; preferably, the lighting conditions cover sunny days, cloudy days, early morning, and evening; preferably, the shooting angles include frontal, side, and oblique angles; preferably, the collected pest targets include common small pests in the field such as aphids, spider mites, thrips, armyworms, and cabbage caterpillars, and the pixel size of the pest targets is mainly distributed in the range of 5 to 50 pixels.
[0032] Based on this, the original image is processed sequentially with invalid region cropping, size scaling and normalization, illumination and color enhancement and noise suppression to obtain a standardized image with a uniform input size, thereby reducing the negative impact of environmental interference on model training.
[0033] Preferably, invalid region cropping is used to remove non-core monitoring areas such as soil, sky, and large areas of weeds, while retaining core monitoring areas such as crop leaves, stems, and areas prone to pests; preferably, size scaling and normalization uniformly adjusts the image size to 640×640 pixels and normalizes the pixel values to the [0,1] range to adapt to the model input size requirements; preferably, illumination and color enhancement adopts adaptive histogram equalization and... The approach combines corrections, with adaptive histogram equalization used to enhance local contrast. Correction is used to adjust the grayscale distribution of images under different lighting conditions and improve the contrast between pest targets and crop backgrounds; preferably, noise suppression adopts a Gaussian bilateral filtering algorithm, which suppresses sensor noise and lighting noise while preserving the edge contour information of pest targets to the greatest extent.
[0034] To further expand the sample size and improve the imbalance of class samples, a multi-strategy data augmentation method was adopted to transform the preprocessed images, thereby improving the model's generalization ability to complex field environments.
[0035] Preferably, data augmentation methods include random rotation, horizontal mirroring, vertical mirroring, brightness perturbation, contrast perturbation, random cropping, Gaussian noise addition, and scale perturbation. These multi-dimensional transformations enrich the sample pose, illumination, and scale diversity, thereby improving the model's adaptability to changes in illumination, pose, and background.
[0036] After image preprocessing and enhancement, bounding boxes and category labels are generated for the pest targets. The labeled dataset is then divided into training, validation, and test sets for model training, parameter tuning, and performance verification, respectively.
[0037] Preferably, using The annotation tool is used for annotation, with the smallest bounding rectangle of the pest target as the annotation box, and the edge of the annotation box should fit the outline of the pest as closely as possible; preferably, after the annotation is completed, manual review is performed to remove invalid samples that are missing, mislabeled, or severely occluded; preferably, the dataset is divided into training set, validation set and test set in a ratio of 7:2:1.
[0038] S2, Constructing an Improved Version Pest detection model, the model is based on Based on the infrastructure, retain Feature extraction network, which will integrate the original Structure replacement A high-frequency enhanced multi-scale feature fusion network replaces the original detection head with... Lightweight decoupling detection head; After completing the dataset construction, further improvements will be made. Pest detection models, targeting native pests Customized architecture improvements are implemented to address shortcomings in the detection of minute pests, such as insufficient feature fusion and limited edge deployment. Preferably, The feature extraction network is used to perform hierarchical feature extraction on a field image with an input size of 640×640×3, and outputs three types of feature maps at different scales, namely shallow detail features. Mid-level semantic features With deep global features ;in It includes information on the edges, texture, and local morphology of pests. It includes information distinguishing the local structure of pests from the crop background. It includes target category and global context information.
[0039] To address the optimization needs of the feature fusion process, the original... Structure replacement A high-frequency enhanced multi-scale feature fusion network. Preferably, The high-frequency enhanced multi-scale feature fusion network consists of a high-frequency enhanced feature module, a local adaptive enhancement module, an automatic feature injection module, and a reparameter optimization module connected sequentially. It is used to enhance the edge detail features of tiny pests, suppress interference from complex backgrounds, and achieve adaptive multi-scale feature fusion. To address the need for lightweight deployment of the detection head, the original detection head is replaced with... Lightweight decoupling detection head. Preferably, The lightweight decoupled detection head includes a classification branch, a regression branch, a confidence prediction branch, an anchor box optimization unit, and a lightweight activation unit. By decoupling the classification and regression tasks and introducing a lightweight structure, the number of model parameters and computational complexity are reduced while ensuring detection accuracy, thereby improving the adaptability of edge device deployment.
[0040] S3, Through High-frequency enhanced multi-scale feature fusion network The output multi-scale features are enhanced and adaptively fused, and the fused features are then input. A lightweight, decoupled detection head yields detection prediction results. After the model is built, firstly through High-frequency enhanced multi-scale feature fusion network The output multi-scale features are progressively enhanced and adaptively fused, and the fused features are then input. A lightweight, decoupled detection head completes detection and prediction.
[0041] First, the high-frequency enhancement feature module is used to enhance the high-frequency detail features of the pest target, such as the outline, antennae, and wing texture, thereby improving the distinguishability of small pest targets in complex backgrounds.
[0042] Preferably, the high-frequency enhancement feature module obtains the high-frequency enhancement feature by fusing convolutional features and edge response features, and its expression is as follows: in, For the input feature map, This is the feature map after high-frequency enhancement. These are the edge response weight coefficients. After completing the high-frequency enhancement features, a local adaptive enhancement module is further used to suppress background interference such as crop leaf veins, dew, and weed textures, thereby adaptively enhancing the feature response intensity of the target area for pests.
[0043] Preferably, the local adaptive enhancement module employs a joint mechanism of channel attention and spatial attention to assign weights to the high-frequency enhanced feature map, the expression of which is: in, For channel attention weights, Spatial attention weights, This is the feature map after local adaptive enhancement; this module assigns higher weights to the pest target region and lower weights to the crop background region. Based on this, an automatic feature injection module completes the adaptive fusion processing of multi-scale features. The adaptive fusion processing specifically includes the following steps: Will The output shallow, medium, and deep multi-scale features are aligned to the same spatial size through upsampling or downsampling. Calculate the information entropy of features at each scale, by The function obtains the corresponding fusion weights; The multi-scale features are weighted and summed based on the fusion weights to obtain adaptive fusion features. Preferably, the information entropy calculation formula for each scale feature is as follows: in, Indicates the first The normalized feature response distribution in the feature map at each scale. Further through... The function obtains the fusion weights of features at each scale: in, For the first The fusion weights correspond to each scale feature. The formula for calculating the adaptive fusion feature is: in, These represent the aligned shallow, middle, and deep layer features, respectively. This refers to the fusion weights for features at corresponding scales. After completing the adaptive fusion of multi-scale features, the feature extraction path is optimized at the inference end through the reparameter optimization module. A reparameterized design of the training-inference stage structure is adopted. In the training stage, a multi-branch convolutional structure is used to enhance the feature representation capability, and in the inference stage, the multi-branch structure is equivalently fused into a single convolutional structure to improve the inference efficiency of edge devices.
[0044] Preferably, the multi-branch convolutional structure in the training phase includes a 1×1 convolutional branch, a 3×3 convolutional branch, a 5×5 convolutional branch, and an identity mapping branch; preferably, in the inference phase, the convolutional kernel parameters and bias parameters of each branch are equivalently fused to obtain a single 3×3 convolutional structure, reducing the inference computation path without sacrificing detection accuracy. The fused features obtained from the processing are further input A lightweight, decoupled detection head performs detection and prediction. The detection head employs a decoupled structure for classification and regression branches. The regression branch introduces a channel shuffling structure to enhance the interaction of localization information, while the classification branch uses depthwise separable convolutions to reduce computational load. Simultaneously, anchor box parameters adapted to the size of farmland pests are generated based on clustering methods, and the detection head uses a lightweight activation function to further reduce computational complexity. Preferably, the classification and regression branches independently transform the input fused features, reducing mutual interference between the two types of task features and improving the localization accuracy and category recognition accuracy of small pests. The feature transformation expression is as follows: in, These are classification branch features used for pest category prediction; These are regression branch features used for pest bounding box location regression; and These represent the feature transformation functions for the classification branch and the regression branch, respectively. Preferably, the regression branch employs a channel shuffling structure to enhance the spatial localization information interaction between different channels. First, the input features of the regression branch are divided into several groups according to channels, then the channels are rearranged, and finally input into the convolutional layer for bounding box prediction. The process is as follows: in, This is a characteristic of the channel after mixed washing. To predict bounding box parameters, preferably, the classification branch employs a depthwise separable convolutional structure, which includes depthwise convolution and pointwise convolution, the process of which is as follows: in, Represents depthwise convolution. This represents a 1×1 pointwise convolution. Output features for the classification branch.
[0045] Preferably, using The clustering method clusters the bounding boxes based on the width and height data of the training set, generating specialized anchor box parameters adapted to the size distribution of small agricultural pests; let the set of width and height of the pest target bounding boxes be denoted as . After clustering, the anchor frame set is obtained. ,in The number of anchor frames. Preferably, the detection head uses ReLU6 as a lightweight activation function, whose expression is: in, The input value for the activation function is used to reduce the nonlinear computational complexity by limiting the upper limit of the output, thereby improving the inference friendliness of the model on edge computing devices such as mobile devices and embedded platforms.
[0046] S4. Use a joint loss function to iteratively train the pest detection model and optimize the model parameters. After the model structure is built, the pest detection model is iteratively trained using a joint loss function. This optimizes the model parameters to address issues such as unstable localization of small targets, imbalanced class samples, and insufficient learning of difficult-to-classify samples in farmland pest detection. The construction of the joint loss function includes the following steps: GIoU loss is used to optimize the stability of bounding box localization, binary cross-entropy loss is used to optimize the target confidence judgment, and Focal loss is used to optimize the recognition effect of hard-to-classify samples. The three types of loss terms are weighted and combined to form a joint loss function for iterative model training. The formula for calculating the joint loss function is: in, For bounding box regression loss, For target confidence loss, For category loss, These are the weight coefficients for the corresponding loss terms. Preferably, to address the issues of insufficient overlap between predicted and ground truth bounding boxes and unstable localization regression in the early stages of training for small pest targets, the following is introduced in the bounding box regression part: Loss. Let the prediction box be... The real frame is The minimum bounding rectangle of the two is ,but Defined as: The corresponding bounding box loss is: By introducing a minimum bounding rectangle constraint, even if the predicted bounding box has little overlap with the ground truth bounding box, it can still provide an effective direction for localization optimization of the model. Preferably, the target confidence loss adopts a binary cross-entropy function to optimize the model's ability to distinguish the existence of pest targets and reduce the false detection rate in complex backgrounds. Its expression is: in, To predict confidence levels, For the true target to have a label, BCE represents the binary cross-entropy loss. Preferably, to address the problem of imbalanced sample sizes among different pest categories and the difficulty in classifying some small pests, a method is introduced... As a category loss, it reduces the impact of easily classified samples on the total loss and increases the model's attention to difficult-to-classify pest samples and small target samples. Its expression is: in, As a class balance factor, Adjustment factor for easy and difficult samples, This represents the model's predicted probability for the true class. To adapt to the optimization needs of different training stages, a dynamic adjustment step for loss weights is also included: dynamically adjusting the weight coefficients of each loss term based on the detection results of the validation set; increasing the bounding box loss weight in the early stage of training to strengthen small target localization learning; and increasing the class loss and confidence loss weights in the middle and later stages of training to improve classification accuracy and false detection suppression capabilities.
[0047] Preferably, the bounding box loss weights are increased in the early stages of training. To enhance the model's ability to learn the location of small target pests and accelerate localization convergence; to increase the weight of the category loss in the later stages of training. Weighted by confidence loss This improves the model's classification accuracy and background false detection suppression capabilities.
[0048] S5. Input the field images to be detected into the trained pest detection model, output the detection results of pest targets, and complete pest statistics and early warning push. After the model training and optimization are completed, the field images to be detected are input into the trained pest detection model to complete online inference and result output, and the pest situation statistics and early warning push functions are also implemented.
[0049] Specifically, the following steps are included: First, the model output results are filtered for confidence and deduplicated by nonmaximum suppression, and the output pest target category, location, quantity and confidence information are obtained.
[0050] Preferably, a confidence threshold of 0.5 is set, and valid detection results with a confidence level not lower than this threshold are retained. Preferably, a non-maximum suppression algorithm is used to deduplicate candidate boxes with high overlap to obtain the final pest detection results. Preferably, the output results include pest category, pest quantity, target box pixel coordinates, detection confidence level, detection time, and image acquisition location or device number.
[0051] Based on this, the detection results are visualized and overlaid onto the original image, and pest statistics reports are generated according to a preset cycle.
[0052] Preferably, the target bounding box, pest category, and corresponding confidence information are overlaid on the original detection image; preferably, the types, quantities, occurrence areas, and quantity change trends of field pests are statistically analyzed according to a preset time period to generate a standardized pest statistics report, providing a basis for decision-making in farmland pest control.
[0053] At the same time, the system is equipped with a pest infestation early warning mechanism. When the number of pests exceeds a preset threshold within a unit of time, pest infestation early warning information is generated and pushed to the user terminal or agricultural Internet of Things platform.
[0054] Preferably, when the number of a single pest species detected within a unit of time exceeds a preset alarm threshold, the system automatically generates pest infestation warning information; preferably, the detection results, pest statistics and abnormal warning information are pushed to the user terminal or agricultural IoT management platform via 4G, 5G or WiFi wireless communication, so as to realize remote monitoring and timely handling of pest occurrence.
[0055] This invention enhances the edge details of tiny pests by setting up an HPFANet high-frequency enhanced multi-scale feature fusion network and adaptively fuses multi-scale features. It also reduces model computational overhead by using a lightweight LSDHead decoupled detection head. Furthermore, it dynamically optimizes the model training process using a joint loss function composed of GIoU boundary loss, confidence loss, and Focal category loss. The invention provides a complete and standardized workflow for image acquisition and preprocessing, pest statistics, and remote early warning output. This effectively solves the technical problems of traditional YOLOv5 models when applied to farmland scenarios, such as high false negative rates for tiny pests, strong interference from complex backgrounds, poor adaptability to multi-scale targets, difficulty in deploying the model on edge devices, and insufficient recognition accuracy due to imbalanced samples. It can achieve high-precision, low-latency automated pest monitoring and early warning in real agricultural scenarios such as fields and greenhouses, meeting the needs of large-scale smart agriculture pest management and possessing good practical value and application prospects.
[0056] It should be noted that the method of this disclosure embodiment can be executed by a single device, such as a computer or server. The method of this embodiment can also be applied to a distributed scenario, where multiple devices cooperate to complete the task. In such a distributed scenario, one of these devices may execute only one or more steps of the method of this disclosure embodiment, and the multiple devices will interact with each other to complete the method described.
[0057] It should be noted that the above description describes some embodiments of this disclosure. Other embodiments are within the scope of the appended claims. In some cases, it should be understood that the sequence number of each step in the above embodiments does not imply the order of execution; the execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention. The actions or steps recorded in the claims can be performed in a different order than that in the above embodiments and can still achieve the desired result. In addition, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0058] Example 2 Please see Figure 2 Based on the same inventive concept, corresponding to any of the above embodiments, this disclosure also provides a pest detection system for smart agriculture scenarios, which includes a data acquisition and preprocessing module for acquiring field pest images and completing image preprocessing, data enhancement and detection dataset construction. An improved HIF-YOLOv5 model building module is used to build a pest detection model that includes a Backbone feature extraction network, an HPFANet high-frequency enhanced multi-scale feature fusion network, an LSDHead lightweight decoupled detection head, and a detection output layer. The model training and loss function optimization module is used to iteratively train the pest detection model using a joint loss function to achieve dynamic optimization of the model parameters. The pest detection and result output module is used to input field images to be detected and output detection results, completing the visualization and statistical analysis of pest conditions and remote early warning push.
[0059] The system described in the above embodiments is used to implement the corresponding pest detection method for smart agriculture scenarios in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0060] It should be noted that the pest detection system described above for smart agriculture scenarios is presented in the form of functional units. The term "module" here can be implemented in software and / or hardware, without specific limitations.
[0061] For example, a "module" can be a software program, a hardware circuit, or a combination of both that implements the above functions. The hardware circuit may include an application-specific integrated circuit (ASIC), electronic circuitry, a processor (e.g., a shared processor, a proprietary processor, or a group processor) and memory for executing one or more software or firmware programs, integrated logic circuitry, and / or other suitable components that support the described functions.
[0062] Example 3 Based on the same inventive concept, corresponding to the methods of any of the above embodiments, this disclosure also provides a non-transitory computer-readable storage medium storing computer instructions for causing the computer to execute the pest detection method for smart agriculture scenarios as described in any of the above embodiments.
[0063] The computer-readable medium of this embodiment includes permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, magnetic magnetic disk storage or other magnetic storage devices, or any other non-transfer medium that can be used to store information accessible by a computing device.
[0064] The computer instructions stored in the storage medium of the above embodiments are used to cause the computer to execute the pest detection method for smart agriculture scenarios as described in any of the above embodiments, and have the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0065] Those skilled in the art should understand that the discussion of any of the above embodiments is merely exemplary and is not intended to imply that the scope of this disclosure (including the claims) is limited to these examples; within the framework of this disclosure, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of the embodiments of this disclosure as described above, which are not provided in detail for the sake of brevity.
[0066] Additionally, to simplify the description and discussion, and to avoid obscuring the embodiments of this disclosure, the provided drawings may or may not show well-known power / ground connections to integrated circuit (IC) chips and other components. Furthermore, the apparatus may be shown in block diagram form to avoid obscuring the embodiments of this disclosure, and this also takes into account the fact that the details of implementation of these block diagram apparatuses are highly dependent on the platform on which the embodiments of this disclosure will be implemented (i.e., these details should be fully understood by those skilled in the art). While specific details (e.g., circuitry) have been set forth to describe exemplary embodiments of this disclosure, it will be apparent to those skilled in the art that the embodiments of this disclosure may be implemented without these specific details or with variations thereof. Therefore, these descriptions should be considered illustrative rather than restrictive.
[0067] Although this disclosure has been described in conjunction with specific embodiments thereof, many substitutions, modifications, and variations of these embodiments will be apparent to those skilled in the art from the foregoing description. For example, other memory architectures (e.g., dynamic RAM (DRAM)) may be used with the embodiments discussed.
[0068] Therefore, the units of the various examples described in the embodiments of this application can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0069] This disclosure is intended to cover all such substitutions, modifications, and variations that fall within the broad scope of the appended claims. Therefore, any omissions, modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.
Claims
1. A pest detection method for smart agriculture scenarios, characterized in that, The method includes the following steps: S1. Collect images of pests in the field, preprocess and augment the images, and construct a standardized pest detection dataset; S2, Constructing an Improved Version Pest detection model, the model is based on Based on the infrastructure, retain Feature extraction network, which will integrate the original Structure replacement A high-frequency enhanced multi-scale feature fusion network replaces the original detection head with... Lightweight decoupling detection head; S3, Through High-frequency enhanced multi-scale feature fusion network The output multi-scale features are enhanced and adaptively fused, and the fused features are then input. A lightweight, decoupled detection head yields detection prediction results. S4. The pest detection model is iteratively trained using a joint loss function to optimize the model parameters. S5. Input the field image to be detected into the trained pest detection model, output the detection results of the pest target, and complete the pest situation statistics and early warning push.
2. The method according to claim 1, characterized in that, Step S1 specifically includes the following steps: Collect images of field pests covering various crop types, lighting conditions, shooting angles, and insect species, while simultaneously collecting images of pest-free healthy crops. The original image is processed sequentially by cropping invalid regions, scaling and normalizing, enhancing illumination and color, and suppressing noise to obtain a standardized image with a uniform input size; Multiple data augmentation strategies are employed to expand sample size and improve class imbalance. The pest targets are labeled with bounding boxes and categories, and then divided into training set, validation set and test set.
3. The method according to claim 1, characterized in that, In step S3, the feature enhancement process specifically includes the following steps: The basic convolutional features of the input feature map are extracted by cascaded convolution, and the edge response features of the feature map are extracted by edge operators. The two are weighted and fused to obtain high-frequency enhanced features, which strengthen the detailed outline of the pest target. A combined channel attention and spatial attention mechanism is adopted to assign weights to high-frequency enhanced features, thereby enhancing the feature response of the target area of pests and suppressing crop background interference. The formula for calculating the high-frequency enhancement feature is as follows: In the formula, For the input feature map, This is the feature map after high-frequency enhancement. This represents the edge response weighting coefficient.
4. The method according to claim 1, characterized in that, In step S3, the adaptive fusion processing specifically includes the following steps: Align the shallow, medium, and deep multi-scale features output by Backbone to the same spatial size; Calculate the information entropy of features at each scale, by The function obtains the corresponding fusion weights; Adaptive fusion features are obtained by weighted summation of multi-scale features based on fusion weights. The formula for calculating the adaptive fusion feature is as follows: In the formula, These are the aligned multi-layer features, The fusion weights are for the corresponding scales.
5. The method according to claim 1, characterized in that, Step S3 also includes a reparameter optimization and detection head prediction step, specifically: A reparameterized design is adopted for the training-inference phase structure. In the training phase, a multi-branch convolutional structure is used to enhance the feature representation capability. In the inference phase, the multi-branch structure is equivalently fused into a single convolutional structure to improve the inference efficiency of edge devices. Input the fused features The lightweight decoupled detection head adopts a decoupled structure of classification and regression branches. The regression branch introduces a channel shuffling structure to enhance the interaction of localization information, while the classification branch uses depthwise separable convolution to reduce the amount of computation. Anchor frame parameters adapted to the size of farmland pests are generated based on clustering methods, and a lightweight activation function is used in the detection head to reduce computational complexity.
6. The method according to claim 1, characterized in that, In step S4, the construction of the joint loss function specifically includes the following steps: respectively Loss optimization for bounding box localization stability, binary cross-entropy loss optimization for target confidence assessment. Loss optimization improves the recognition performance of hard-to-classify samples; The three types of loss terms are weighted and combined to form a joint loss function for iterative model training; The formula for calculating the joint loss function is as follows: In the formula, for Bounding box regression loss, For target confidence loss, for Category loss, These are the weighting coefficients for the corresponding loss terms.
7. The method according to claim 6, characterized in that, It also includes a step for dynamically adjusting the loss weights: The weight coefficients of each loss term are dynamically adjusted based on the detection results of the validation set. In the early stage of training, the weight of the bounding box loss is increased to strengthen the localization learning of small targets. In the middle and later stages of training, the weights of the category loss and confidence loss are increased to improve classification accuracy and false detection suppression capability.
8. The method according to claim 1, characterized in that, Step S5 specifically includes the following steps: The model output results are filtered for confidence and deduplicated by nonmaximum suppression, and the output pest target category, location, quantity and confidence information are obtained. The detection results are visualized and overlaid onto the original image, and insect infestation statistics reports are generated at preset intervals. When the number of pests exceeds a preset threshold within a unit of time, an pest infestation warning is generated and pushed to the user terminal or agricultural IoT platform.
9. A pest detection system for smart agriculture scenarios, the system being used to implement the method described in any one of claims 1 to 8, characterized in that, It includes: The data acquisition and preprocessing module is used to acquire images of field pests and complete image preprocessing, data augmentation, and detection dataset construction. An improved HIF-YOLOv5 model building module is used to build a pest detection model that includes a Backbone feature extraction network, an HPFANet high-frequency enhanced multi-scale feature fusion network, an LSDHead lightweight decoupled detection head, and a detection output layer. The model training and loss function optimization module is used to iteratively train the pest detection model using a joint loss function to achieve dynamic optimization of the model parameters. The pest detection and result output module is used to input field images to be detected and output detection results, completing the visualization and statistical analysis of pest conditions and remote early warning push.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed, implements the method as described in any one of claims 1 to 8.